Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
arXiv cs.AI Security & Safety
Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulne…
arXiv cs.AI Security & Safety
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
arXiv cs.AI Security & Safety
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
arXiv cs.AI Security & Safety
An AI Agent Execution Environment to Safeguard User Data
arXiv cs.AI Security & Safety
ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents
arXiv cs.AI Security & Safety
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
arXiv cs.AI Security & Safety
Data Security in Large Language Models: Risks, Defense, and Directions
arXiv cs.AI Security & Safety
Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landsc…
arXiv cs.AI Security & Safety
When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based …
arXiv cs.AI Security & Safety
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository S…
arXiv cs.AI Security & Safety
DSS: Dynamic Semantic Steering for Robust Concept Erasure in Diffusion Models
arXiv cs.AI Security & Safety
Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language …
arXiv cs.AI Security & Safety
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
arXiv cs.AI Security & Safety
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
arXiv cs.AI Security & Safety
IntraGuard: Committee-Side Defenses Against Review Outsourcing to Commercial Chatbots
arXiv cs.AI Security & Safety
PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Di…
arXiv cs.AI Security & Safety
SENTINEL: A Multi-Pathway Architecture for Detecting Living-Off-the-Land APT Attacks on W…
arXiv cs.AI Security & Safety
When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
Bloomberg Technology Security & Safety
Hugging Face Scientist on Safety Issues With Agentic AI
Bloomberg Technology Security & Safety
Google DeepMind Staffer Says AI May ‘Kill Us All’ in Exit Post
Bloomberg Technology Security & Safety
Gebru: AI Security & Safety Is About Human Control
Bloomberg Technology Security & Safety
Cloudflare CEO: Good Guys Have More Tools Than Bad Guys
Ars Technica AI Security & Safety
AI bots "Timmy," "Ren," and "Jackie" are flooding social media with slop
Wired AI Security & Safety
New York Seizes a Dozen Celebrity Deepfake Websites
Bloomberg Technology Security & Safety
Ex-Google DeepMind Insider: Why We MUST Slow Down AI Now
Hacker News (AI filter) Security & Safety
Adversarial Fashion Makes a Statement on AI Panopticon
Hacker News (AI filter) Security & Safety
What a time to be alive – rouge AI agents attack RubyGems.org
Wired AI Security & Safety
Sexually Explicit Deepfake Sites Target 100-Plus Politicians in Europe
Bloomberg Technology Security & Safety
OpenAI President on Doing Business in the Wake of Hugging Face