VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Att…
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated…
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Red-Teaming the Agentic Red-Team
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
Cryptographic certificates of validity for trustworthy AI
One Year Later...The Harms Persist, But So Do We!
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Ba…
AI Hiring Tools Yield Racial Bias and Systemic Rejection; 26% Black & 15% Asian
Red-Teaming the Agentic Red-Team
The Growing Crisis for America's Child Abuse Investigators
OpenAI's new Daybreak initiative will help open-source projects fend off bugs
Top spy agencies say AI cyber threats will impact you within months. Here’s why
LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity
Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents
GIF: Locally Sound Geometric Information Flow Control for LLMs
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports
MedFedPure: A Medical Federated Framework with MAE-based Detection and Diffusion Purifica…
When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Ag…
Detecting Malicious Agent Skills in the Wild using Attention
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Associati…
From CVE to CWE: Syscall-Based HIDS Generalisation
Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer
Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents
Signals in the Noise: Open Source Intelligence (OSINT) for AI Loss of Control Detection