Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
CEAR: Certified Ensemble Adversarial Robustness in DNNs
arXiv cs.AI Security & Safety
Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
arXiv cs.AI Security & Safety
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
arXiv cs.AI Security & Safety
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
arXiv cs.AI Security & Safety
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Mode…
arXiv cs.AI Security & Safety
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
arXiv cs.AI Security & Safety
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
arXiv cs.AI Security & Safety
Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem
arXiv cs.AI Security & Safety
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
arXiv cs.AI Security & Safety
The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer
arXiv cs.AI Security & Safety
SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-…
arXiv cs.AI Security & Safety
Ethical Hyper-Velocity (EHV): A Hardware-Rooted Zero-Trust Runtime Enforcement Architectu…
arXiv cs.AI Security & Safety
Safety Must Precede the Deployment of Open-Ended AI
arXiv cs.AI Security & Safety
SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
arXiv cs.AI Security & Safety
Cross-modal linkage risk in clinical vision-language models
arXiv cs.AI Security & Safety
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS …
arXiv cs.AI Security & Safety
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
arXiv cs.AI Security & Safety
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Languag…
arXiv cs.AI Security & Safety
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
arXiv cs.AI Security & Safety
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box R…
arXiv cs.AI Security & Safety
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
Ars Technica AI Security & Safety
Hackers duped Meta AI support chatbot to steal celebrity Instagram accounts
Engadget Security & Safety
Meta's AI support chatbot made it ridiculously easy for hackers to take over Instagram ac…
The Verge AI Security & Safety
Meta’s own AI was exploited to hijack Instagram accounts
Mashable Security & Safety
Hackers say that Meta AI helped them compromise big Instagram accounts
Ars Technica AI Security & Safety
Allegedly trashing Airbnbs to test robots puts startup in legal trouble
Hugging Face Daily Papers Security & Safety
Monitoring Agentic Systems Before They're Reliable
Hugging Face Daily Papers Security & Safety
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
arXiv cs.AI Security & Safety
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations o…
arXiv cs.AI Security & Safety
The Surface You Test Is Not the Surface That Breaks