BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Revi…
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Adversarial Attacks Leverage Interference Between Features in Superposition
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Ad…
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
Detecting and Mitigating DDoS Attacks with AI: A Survey
Timestamp-Aware Spatio-Temporal Graph Contrastive Learning for Network Intrusion Detection
Graph neural networks at war: integrating cybersecurity and drone intelligence in the Isr…
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking …
An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and …
This Copilot vulnerability could expose emails, 2FA codes, and other sensitive data
Securing the future of AI agents
Critical Copilot vulnerability allowed hackers to seal 2FA code from users
AI Red Teaming Explained: What It Is and Why You Need It
A Survey on Agentic Security: Applications, Threats and Defenses
Phishing Email Detection Using Large Language Models
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Atta…
AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
Computational Safety for Generative AI: A Hypothesis Testing Perspective
A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framewor…
Robust Spoofed Speech Detection via Temporal Pyramid Modeling
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
The Proxy Knows Too Much: Sealing LLM API Routers with Attested TEEs
Communication-Efficient Verifiable Attention for LLM Inference
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits…
Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source…