Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
ToolGuardian: Declarative Security for AI Agent-Tool Interactions
arXiv cs.AI Security & Safety
Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense
Bloomberg Technology Security & Safety
Instagram, Facebook Ran AI ‘Nudify’ Ads from China, Report Says
Hugging Face Blog Security & Safety
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
TechCrunch AI Security & Safety
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack
Engadget Security & Safety
OpenAI's rogue agent went on a hacking spree that lasted days, Reuters says
Wired AI Security & Safety
The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days
Wired AI Security & Safety
Did Chinese AI Steal From Anthropic, and OpenAI Loses Control of Two Models
arXiv cs.AI Security & Safety
Making Open-Source Text LLM Watermarks Durable Against Merging
arXiv cs.AI Security & Safety
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Mode…
arXiv cs.AI Security & Safety
Geometric Configurations of Perturbed Jailbreak Prompts
arXiv cs.AI Security & Safety
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-…
arXiv cs.AI Security & Safety
Incomplete Prompt Jailbreaks in Large Language Models
arXiv cs.AI Security & Safety
Code Monitor Red Teaming for Public-Test-Passing Code
arXiv cs.AI Security & Safety
AI Security Policy Should Assess Systems, Not Only Models
arXiv cs.AI Security & Safety
Robust Critics: Defending LLMs Against Multi-Turn Attacks
arXiv cs.AI Security & Safety
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
TechCrunch AI Security & Safety
How AI guardrails are impeding the work of offensive cybersecurity researchers
Ars Technica AI Security & Safety
AI arms race in line for a reckoning after OpenAI hacking incident
arXiv cs.AI Security & Safety
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
arXiv cs.AI Security & Safety
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception …
arXiv cs.AI Security & Safety
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
arXiv cs.AI Security & Safety
HijackKV: New Threat in Position-Independent KV Cache Reuse
arXiv cs.AI Security & Safety
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
arXiv cs.AI Security & Safety
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
arXiv cs.AI Security & Safety
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
arXiv cs.AI Security & Safety
Integrity of peer-to-peer distributed LLM inference under malicious nodes
arXiv cs.AI Security & Safety
Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
arXiv cs.AI Security & Safety
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language…
arXiv cs.AI Security & Safety
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety