CmdNeedle: Measuring the Incompleteness of Command Denylists for AI Agents
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query …
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
MASCOT-Android: A Curated Dataset and Automated Collection Pipeline for Android Malware S…
InstantForget: Update-Free Backdoor Unlearning with Inference-Time Feature Reset
Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?
Honeypot Protocol
GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
Automated jailbreak attack targeting multiple defense strategies
Discrete optimal transport is a strong audio adversarial attack
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference P…
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Microsoft shares new data on email security performance
What Is Anthropic’s Mythos AI and Why Was It Blocked?
Chainguard, Cyber Firms Use AI to Hunt for Open-Source Flaws
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOM…
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security …
Same-Origin Policy for Agentic Browsers
Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in…
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automo…
Securing the Future of IoMT in the Post-Quantum Era: An Edge-Native Federated Learning Ap…
Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
Crypto Token’s 50% Wipeout Shows Magnitude of AI-Hacking Threat