DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical C…
AI Safety: Not Optional, Not Later
Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
Membership Inference Attacks on Recommender System: A Survey
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
Rogue AI Breakouts Raise Pressure for New Rules
OpenAI’s rogue AI tried to hack another company in May
OpenAI agents hacked a software service before the Hugging Face incident
The Worst Spam Emails: Inside iLands' AI Agent Hustle
From Hacks to Bioweapons, Claude Misuse Is Now Everywhere
No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vu…
DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injectio…
Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against …
Architecting the Secure AI-SOC: A Neurosymbolic Framework for Pipeline Integrity and Thre…
Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitiv…
A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems
Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks
Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems
An Anthropic researcher’s doomsday warning comes at a very interesting time
Anthropic spent this week in hot water over cybersecurity
Claude users found ways around safeguards for bioweapons research
Anthropic Says US Adversaries Aimed Claude at Weapons Research
Anthropic Says Yemeni Cell Used Claude in Missile Development
Anthropic caught scientists using Claude to further biological weapon research
Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek
Former OpenAI, Anthropic Employee Post, Hugging Face Hack Sound Alarm on AI
Avian flu, drone swarms, and mass surveillance included in Anthropics safety report
Moonshot Secretly Routed User Requests Through Claude, Anthropic Says
AI agents are flooding public services with new requests