Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Revi…
arXiv cs.AI Security & Safety
Adversarial Attacks Leverage Interference Between Features in Superposition
arXiv cs.AI Security & Safety
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
arXiv cs.AI Security & Safety
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Ad…
arXiv cs.AI Security & Safety
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
arXiv cs.AI Security & Safety
TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations
arXiv cs.AI Security & Safety
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
arXiv cs.AI Security & Safety
Detecting and Mitigating DDoS Attacks with AI: A Survey
arXiv cs.AI Security & Safety
Timestamp-Aware Spatio-Temporal Graph Contrastive Learning for Network Intrusion Detection
arXiv cs.AI Security & Safety
Graph neural networks at war: integrating cybersecurity and drone intelligence in the Isr…
arXiv cs.AI Security & Safety
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
arXiv cs.AI Security & Safety
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking …
arXiv cs.AI Security & Safety
An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and …
Mashable Security & Safety
This Copilot vulnerability could expose emails, 2FA codes, and other sensitive data
Google DeepMind Security & Safety
Securing the future of AI agents
Ars Technica AI Security & Safety
Critical Copilot vulnerability allowed hackers to seal 2FA code from users
AI News Security & Safety
AI Red Teaming Explained: What It Is and Why You Need It
arXiv cs.AI Security & Safety
A Survey on Agentic Security: Applications, Threats and Defenses
arXiv cs.AI Security & Safety
Phishing Email Detection Using Large Language Models
arXiv cs.AI Security & Safety
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Atta…
arXiv cs.AI Security & Safety
AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
arXiv cs.AI Security & Safety
Computational Safety for Generative AI: A Hypothesis Testing Perspective
arXiv cs.AI Security & Safety
A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framewor…
arXiv cs.AI Security & Safety
Robust Spoofed Speech Detection via Temporal Pyramid Modeling
arXiv cs.AI Security & Safety
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
arXiv cs.AI Security & Safety
The Proxy Knows Too Much: Sealing LLM API Routers with Attested TEEs
arXiv cs.AI Security & Safety
Communication-Efficient Verifiable Attention for LLM Inference
arXiv cs.AI Security & Safety
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits…
arXiv cs.AI Security & Safety
Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems
arXiv cs.AI Security & Safety
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source…