Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reve…
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Ag…
Stateful Online Monitoring Catches Distributed Agent Attacks
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, …
Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation
AI Loss of Control Incident Management: Response & Resilience
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
Prompt Injection as Role Confusion
Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Cre…
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Bac…
“Tech and Tariffs” Campaign: Influence activity targeting US tech policy
“Data Center Bandwagon” Campaign: US-targeted influence activity
AI grifters are creating fake Black people to sell Shein junk
AI Dangers Eclipse Nuclear Weapons at Singapore Defense Forum
CAPTCHAs can still detect AI agents
The Distillation Game: Adaptive Attacks & Efficient Defenses
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and C…
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
Provably Secure Agent Guardrail
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
SelfGrader: LLM Jailbreak Detection via Anchored Token-Level Logits
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
AIRGuard: Guarding Agent Actions with Runtime Authority Control
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening