Jailbreaking Multimodal Large Language Models using Multi-Clip Video
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Languag…
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box R…
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-A…
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
CEAR: Certified Ensemble Adversarial Robustness in DNNs
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Mode…
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
Hackers duped Meta AI support chatbot to steal celebrity Instagram accounts
Meta's AI support chatbot made it ridiculously easy for hackers to take over Instagram ac…
Meta’s own AI was exploited to hijack Instagram accounts
Hackers say that Meta AI helped them compromise big Instagram accounts
Allegedly trashing Airbnbs to test robots puts startup in legal trouble
Monitoring Agentic Systems Before They're Reliable
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
AI Loss of Control Incident Management: Response & Resilience
The Surface You Test Is Not the Surface That Breaks
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Ag…
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack