ToolGuardian: Declarative Security for AI Agent-Tool Interactions
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense
Instagram, Facebook Ran AI ‘Nudify’ Ads from China, Report Says
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack
OpenAI's rogue agent went on a hacking spree that lasted days, Reuters says
The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days
Did Chinese AI Steal From Anthropic, and OpenAI Loses Control of Two Models
Making Open-Source Text LLM Watermarks Durable Against Merging
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Mode…
Geometric Configurations of Perturbed Jailbreak Prompts
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-…
Incomplete Prompt Jailbreaks in Large Language Models
Code Monitor Red Teaming for Public-Test-Passing Code
AI Security Policy Should Assess Systems, Not Only Models
Robust Critics: Defending LLMs Against Multi-Turn Attacks
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
How AI guardrails are impeding the work of offensive cybersecurity researchers
AI arms race in line for a reckoning after OpenAI hacking incident
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception …
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
HijackKV: New Threat in Position-Independent KV Cache Reuse
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
Integrity of peer-to-peer distributed LLM inference under malicious nodes
Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language…
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety