Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Privacy-Preserving AI Verification via Minimal Information Disclosure
A Security-Oriented Lifecycle Model for Large Language Model Systems
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Ag…
$S^3$: Improving Agent Safety through Multi-Stage Defense
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
OK, Well, Rogue AI Agents Are Hacking Again
AI fuels more than half of cybercrime in Africa as scams surge – Interpol
OpenAI, Anthropic AI Models Involved in More Security Incidents
Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already show…
Data Centers Exposed US Telecoms to China Hacks, US Panel to Say
Hackers Steal Bitcoin, Visa Outlines Stablecoin Plans | Bloomberg Crypto 8/4/2026
Third-party cyber evaluations involving OpenAI models
Apple caps bug bounty program due to deluge of AI submissions
Apple Asks Judge to Bar OpenAI From Using Alleged Trade Secrets
AI Now Fuels Over Half of Africa’s Cybercrime, Study Finds
Apple says more ex-employees may have taken confidential data to OpenAI
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in…
Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based L…
Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable…
From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners …
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
Securing Agentic AI: From Per-Action Checks to Trajectory Assurance
Exposed by Design: A Dynamic Security Assessment of Internet-Facing MCP Servers at Scale
Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Steal…
Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw
MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning