Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
CmdNeedle: Measuring the Incompleteness of Command Denylists for AI Agents
arXiv cs.AI Security & Safety
FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query …
arXiv cs.AI Security & Safety
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
arXiv cs.AI Security & Safety
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
arXiv cs.AI Security & Safety
MASCOT-Android: A Curated Dataset and Automated Collection Pipeline for Android Malware S…
arXiv cs.AI Security & Safety
InstantForget: Update-Free Backdoor Unlearning with Inference-Time Feature Reset
arXiv cs.AI Security & Safety
Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?
arXiv cs.AI Security & Safety
Honeypot Protocol
arXiv cs.AI Security & Safety
GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
arXiv cs.AI Security & Safety
Automated jailbreak attack targeting multiple defense strategies
arXiv cs.AI Security & Safety
Discrete optimal transport is a strong audio adversarial attack
arXiv cs.AI Security & Safety
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
arXiv cs.AI Security & Safety
Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
arXiv cs.AI Security & Safety
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference P…
arXiv cs.AI Security & Safety
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Microsoft Source (AI + Cloud) Security & Safety
Microsoft shares new data on email security performance
Bloomberg Technology Security & Safety
What Is Anthropic’s Mythos AI and Why Was It Blocked?
Bloomberg Technology Security & Safety
Chainguard, Cyber Firms Use AI to Hunt for Open-Source Flaws
arXiv cs.AI Security & Safety
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOM…
arXiv cs.AI Security & Safety
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
arXiv cs.AI Security & Safety
Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security …
arXiv cs.AI Security & Safety
Same-Origin Policy for Agentic Browsers
arXiv cs.AI Security & Safety
Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in…
arXiv cs.AI Security & Safety
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
arXiv cs.AI Security & Safety
AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
arXiv cs.AI Security & Safety
I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automo…
arXiv cs.AI Security & Safety
Securing the Future of IoMT in the Post-Quantum Era: An Edge-Native Federated Learning Ap…
arXiv cs.AI Security & Safety
Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications
arXiv cs.AI Security & Safety
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
Bloomberg Technology Security & Safety
Crypto Token’s 50% Wipeout Shows Magnitude of AI-Hacking Threat