Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

Hugging Face Daily Papers Security & Safety
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated…
Hugging Face Daily Papers Security & Safety
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
arXiv cs.AI Security & Safety
Cryptographic certificates of validity for trustworthy AI
arXiv cs.AI Security & Safety
One Year Later...The Harms Persist, But So Do We!
arXiv cs.AI Security & Safety
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
arXiv cs.AI Security & Safety
Red-Teaming the Agentic Red-Team
Hugging Face Daily Papers Security & Safety
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Ba…
Hacker News (AI filter) Security & Safety
AI Hiring Tools Yield Racial Bias and Systemic Rejection; 26% Black & 15% Asian
Hugging Face Daily Papers Security & Safety
Red-Teaming the Agentic Red-Team
Bloomberg Technology Security & Safety
The Growing Crisis for America's Child Abuse Investigators
Engadget Security & Safety
OpenAI's new Daybreak⁠ initiative will help open-source projects fend off bugs
AI News Security & Safety
Top spy agencies say AI cyber threats will impact you within months. Here’s why
arXiv cs.AI Security & Safety
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent…
arXiv cs.AI Security & Safety
Confidently Wrong: Severity-Aware Calibration of Prompt-Injection Detectors under Attack …
arXiv cs.AI Security & Safety
CLIP-guided Diffusion Model for Backdoor Generation in Sensor-based Human Activity Recogn…
arXiv cs.AI Security & Safety
Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long C…
arXiv cs.AI Security & Safety
TIF: Learning Temporal Invariance in Android Malware Detectors
arXiv cs.AI Security & Safety
Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents
arXiv cs.AI Security & Safety
ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
arXiv cs.AI Security & Safety
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
arXiv cs.AI Security & Safety
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents
arXiv cs.AI Security & Safety
Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning
arXiv cs.AI Security & Safety
AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports
arXiv cs.AI Security & Safety
How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
arXiv cs.AI Security & Safety
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
arXiv cs.AI Security & Safety
Detecting Malicious Agent Skills in the Wild using Attention
arXiv cs.AI Security & Safety
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
arXiv cs.AI Security & Safety
From CVE to CWE: Syscall-Based HIDS Generalisation
arXiv cs.AI Security & Safety
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Associati…
arXiv cs.AI Security & Safety
Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents