Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
arXiv cs.AI Security & Safety
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Languag…
arXiv cs.AI Security & Safety
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
arXiv cs.AI Security & Safety
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
arXiv cs.AI Security & Safety
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
arXiv cs.AI Security & Safety
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
arXiv cs.AI Security & Safety
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
arXiv cs.AI Security & Safety
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box R…
arXiv cs.AI Security & Safety
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-A…
arXiv cs.AI Security & Safety
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
arXiv cs.AI Security & Safety
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
arXiv cs.AI Security & Safety
CEAR: Certified Ensemble Adversarial Robustness in DNNs
arXiv cs.AI Security & Safety
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
arXiv cs.AI Security & Safety
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
arXiv cs.AI Security & Safety
Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
arXiv cs.AI Security & Safety
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
arXiv cs.AI Security & Safety
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Mode…
arXiv cs.AI Security & Safety
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
arXiv cs.AI Security & Safety
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
Ars Technica AI Security & Safety
Hackers duped Meta AI support chatbot to steal celebrity Instagram accounts
Engadget Security & Safety
Meta's AI support chatbot made it ridiculously easy for hackers to take over Instagram ac…
The Verge AI Security & Safety
Meta’s own AI was exploited to hijack Instagram accounts
Mashable Security & Safety
Hackers say that Meta AI helped them compromise big Instagram accounts
Ars Technica AI Security & Safety
Allegedly trashing Airbnbs to test robots puts startup in legal trouble
Hugging Face Daily Papers Security & Safety
Monitoring Agentic Systems Before They're Reliable
Hugging Face Daily Papers Security & Safety
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
arXiv cs.AI Security & Safety
AI Loss of Control Incident Management: Response & Resilience
arXiv cs.AI Security & Safety
The Surface You Test Is Not the Surface That Breaks
arXiv cs.AI Security & Safety
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Ag…
arXiv cs.AI Security & Safety
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack