Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

Hugging Face Daily Papers Security & Safety
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Att…
Hugging Face Daily Papers Security & Safety
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated…
Hugging Face Daily Papers Security & Safety
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
arXiv cs.AI Security & Safety
Red-Teaming the Agentic Red-Team
arXiv cs.AI Security & Safety
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
arXiv cs.AI Security & Safety
Cryptographic certificates of validity for trustworthy AI
arXiv cs.AI Security & Safety
One Year Later...The Harms Persist, But So Do We!
Hugging Face Daily Papers Security & Safety
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Ba…
Hacker News (AI filter) Security & Safety
AI Hiring Tools Yield Racial Bias and Systemic Rejection; 26% Black & 15% Asian
Hugging Face Daily Papers Security & Safety
Red-Teaming the Agentic Red-Team
Bloomberg Technology Security & Safety
The Growing Crisis for America's Child Abuse Investigators
Engadget Security & Safety
OpenAI's new Daybreak⁠ initiative will help open-source projects fend off bugs
AI News Security & Safety
Top spy agencies say AI cyber threats will impact you within months. Here’s why
arXiv cs.AI Security & Safety
LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity
arXiv cs.AI Security & Safety
Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning
arXiv cs.AI Security & Safety
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
arXiv cs.AI Security & Safety
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents
arXiv cs.AI Security & Safety
GIF: Locally Sound Geometric Information Flow Control for LLMs
arXiv cs.AI Security & Safety
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
arXiv cs.AI Security & Safety
How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
arXiv cs.AI Security & Safety
AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports
arXiv cs.AI Security & Safety
MedFedPure: A Medical Federated Framework with MAE-based Detection and Diffusion Purifica…
arXiv cs.AI Security & Safety
When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Ag…
arXiv cs.AI Security & Safety
Detecting Malicious Agent Skills in the Wild using Attention
arXiv cs.AI Security & Safety
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
arXiv cs.AI Security & Safety
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Associati…
arXiv cs.AI Security & Safety
From CVE to CWE: Syscall-Based HIDS Generalisation
arXiv cs.AI Security & Safety
Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer
arXiv cs.AI Security & Safety
Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents
arXiv cs.AI Security & Safety
Signals in the Noise: Open Source Intelligence (OSINT) for AI Loss of Control Detection