Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reve…
arXiv cs.AI Security & Safety
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
arXiv cs.AI Security & Safety
CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability
arXiv cs.AI Security & Safety
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Ag…
arXiv cs.AI Security & Safety
Stateful Online Monitoring Catches Distributed Agent Attacks
arXiv cs.AI Security & Safety
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, …
arXiv cs.AI Security & Safety
Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation
arXiv cs.AI Security & Safety
AI Loss of Control Incident Management: Response & Resilience
arXiv cs.AI Security & Safety
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
arXiv cs.AI Security & Safety
Prompt Injection as Role Confusion
arXiv cs.AI Security & Safety
Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Cre…
arXiv cs.AI Security & Safety
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Bac…
OpenAI News Security & Safety
“Tech and Tariffs” Campaign: Influence activity targeting US tech policy
OpenAI News Security & Safety
“Data Center Bandwagon” Campaign: US-targeted influence activity
The Verge AI Security & Safety
AI grifters are creating fake Black people to sell Shein junk
Bloomberg Technology Security & Safety
AI Dangers Eclipse Nuclear Weapons at Singapore Defense Forum
Hacker News (AI filter) Security & Safety
CAPTCHAs can still detect AI agents
arXiv cs.AI Security & Safety
The Distillation Game: Adaptive Attacks & Efficient Defenses
arXiv cs.AI Security & Safety
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and C…
arXiv cs.AI Security & Safety
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
arXiv cs.AI Security & Safety
Provably Secure Agent Guardrail
arXiv cs.AI Security & Safety
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
arXiv cs.AI Security & Safety
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
arXiv cs.AI Security & Safety
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
arXiv cs.AI Security & Safety
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
arXiv cs.AI Security & Safety
SelfGrader: LLM Jailbreak Detection via Anchored Token-Level Logits
arXiv cs.AI Security & Safety
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
arXiv cs.AI Security & Safety
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
arXiv cs.AI Security & Safety
AIRGuard: Guarding Agent Actions with Runtime Authority Control
arXiv cs.AI Security & Safety
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening