Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

Hugging Face Daily Papers Security & Safety
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
Ars Technica AI Security & Safety
Ukraine's one-time test used fully autonomous drones to kill Russian soldiers
TechCrunch AI Security & Safety
Google sues alleged Chinese cybercrime operation that used AI to send scam texts
Ars Technica AI Security & Safety
Google sues Chinese cybercrime network that used Gemini to automate scams
Engadget Security & Safety
Google sues Chinese scammers using Gemini AI for fraud
Bloomberg Technology Security & Safety
Scammers Used Gemini AI to Help Build Spam Messages, Google Says
Hacker News (AI filter) Security & Safety
AI agent bankrupted their operator while trying to scan DN42
arXiv cs.AI Security & Safety
SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems
arXiv cs.AI Security & Safety
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using A…
arXiv cs.AI Security & Safety
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
arXiv cs.AI Security & Safety
Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models
arXiv cs.AI Security & Safety
SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems
arXiv cs.AI Security & Safety
A Survey on Long-Term Memory Security in LLM Agents: Attacks, Defenses, and Governance Ac…
arXiv cs.AI Security & Safety
The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI S…
arXiv cs.AI Security & Safety
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web …
arXiv cs.AI Security & Safety
MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems
arXiv cs.AI Security & Safety
The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Require…
arXiv cs.AI Security & Safety
Reframing AI Loss of Control: What It Is, How to Have It, How to Lose It
arXiv cs.AI Security & Safety
PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learnin…
arXiv cs.AI Security & Safety
PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections
Bloomberg Technology Security & Safety
ChatGPT Pushed Back on LA Fire Suspect’s Vision of Burning City
Wired AI Security & Safety
Grok Is Still Hosting Sexualized Deepfakes of Famous Women
Bloomberg Technology Security & Safety
Former xAI Staffer Says He Was Fired for Questioning Grok Safety
Engadget Security & Safety
OpenAI says fake accounts from China tried to turn Americans against data centers
arXiv cs.AI Security & Safety
Robust Privacy: Inference-Stage Privacy through Certified Robustness
arXiv cs.AI Security & Safety
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
arXiv cs.AI Security & Safety
MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation
arXiv cs.AI Security & Safety
T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking
arXiv cs.AI Security & Safety
Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirica…
arXiv cs.AI Security & Safety
Runtime Skill Audit: Targeted Runtime Probing for Agent Skill Security