Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Prompt Injection as Role Confusion
arXiv cs.AI Security & Safety
CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability
arXiv cs.AI Security & Safety
Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation
arXiv cs.AI Security & Safety
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations o…
arXiv cs.AI Security & Safety
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
arXiv cs.AI Security & Safety
Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Cre…
arXiv cs.AI Security & Safety
Stateful Online Monitoring Catches Distributed Agent Attacks
arXiv cs.AI Security & Safety
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, …
arXiv cs.AI Security & Safety
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Bac…
arXiv cs.AI Security & Safety
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reve…
OpenAI News Security & Safety
“Tech and Tariffs” Campaign: Influence activity targeting US tech policy
OpenAI News Security & Safety
“Data Center Bandwagon” Campaign: US-targeted influence activity
The Verge AI Security & Safety
AI grifters are creating fake Black people to sell Shein junk
Bloomberg Technology Security & Safety
AI Dangers Eclipse Nuclear Weapons at Singapore Defense Forum
Hacker News (AI filter) Security & Safety
CAPTCHAs can still detect AI agents
arXiv cs.AI Security & Safety
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
arXiv cs.AI Security & Safety
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
arXiv cs.AI Security & Safety
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
arXiv cs.AI Security & Safety
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
arXiv cs.AI Security & Safety
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned S…
arXiv cs.AI Security & Safety
SelfGrader: LLM Jailbreak Detection via Anchored Token-Level Logits
arXiv cs.AI Security & Safety
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
arXiv cs.AI Security & Safety
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
arXiv cs.AI Security & Safety
Provably Secure Agent Guardrail
arXiv cs.AI Security & Safety
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
arXiv cs.AI Security & Safety
The Distillation Game: Adaptive Attacks & Efficient Defenses
arXiv cs.AI Security & Safety
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and C…
arXiv cs.AI Security & Safety
AIRGuard: Guarding Agent Actions with Runtime Authority Control
arXiv cs.AI Security & Safety
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
Ars Technica AI Security & Safety
Fed up with vibe coders, dev sneaks data-nuking prompt injection into their code