How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated…
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Cryptographic certificates of validity for trustworthy AI
One Year Later...The Harms Persist, But So Do We!
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
Red-Teaming the Agentic Red-Team
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Ba…
AI Hiring Tools Yield Racial Bias and Systemic Rejection; 26% Black & 15% Asian
Red-Teaming the Agentic Red-Team
The Growing Crisis for America's Child Abuse Investigators
OpenAI's new Daybreak initiative will help open-source projects fend off bugs
Top spy agencies say AI cyber threats will impact you within months. Here’s why
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent…
Confidently Wrong: Severity-Aware Calibration of Prompt-Injection Detectors under Attack …
CLIP-guided Diffusion Model for Backdoor Generation in Sensor-based Human Activity Recogn…
Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long C…
TIF: Learning Temporal Invariance in Android Malware Detectors
Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents
ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents
Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning
AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports
How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
Detecting Malicious Agent Skills in the Wild using Attention
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
From CVE to CWE: Syscall-Based HIDS Generalisation
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Associati…
Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents