Prompt Injection as Role Confusion
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability
Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations o…
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Cre…
Stateful Online Monitoring Catches Distributed Agent Attacks
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, …
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Bac…
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reve…
“Tech and Tariffs” Campaign: Influence activity targeting US tech policy
“Data Center Bandwagon” Campaign: US-targeted influence activity
AI grifters are creating fake Black people to sell Shein junk
AI Dangers Eclipse Nuclear Weapons at Singapore Defense Forum
CAPTCHAs can still detect AI agents
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned S…
SelfGrader: LLM Jailbreak Detection via Anchored Token-Level Logits
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
Provably Secure Agent Guardrail
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
The Distillation Game: Adaptive Attacks & Efficient Defenses
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and C…
AIRGuard: Guarding Agent Actions with Runtime Authority Control
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
Fed up with vibe coders, dev sneaks data-nuking prompt injection into their code