Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Dete…
arXiv cs.AI Security & Safety
Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers…
arXiv cs.AI Security & Safety
ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
Bloomberg Technology Security & Safety
Instagram, Facebook Ran AI ‘Nudify’ Ads from China, Report Says
Hugging Face Blog Security & Safety
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
TechCrunch AI Security & Safety
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack
Engadget Security & Safety
OpenAI's rogue agent went on a hacking spree that lasted days, Reuters says
Wired AI Security & Safety
The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days
Wired AI Security & Safety
Did Chinese AI Steal From Anthropic, and OpenAI Loses Control of Two Models
arXiv cs.AI Security & Safety
Incomplete Prompt Jailbreaks in Large Language Models
arXiv cs.AI Security & Safety
Geometric Configurations of Perturbed Jailbreak Prompts
arXiv cs.AI Security & Safety
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-…
arXiv cs.AI Security & Safety
AI Security Policy Should Assess Systems, Not Only Models
arXiv cs.AI Security & Safety
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
arXiv cs.AI Security & Safety
Making Open-Source Text LLM Watermarks Durable Against Merging
arXiv cs.AI Security & Safety
Code Monitor Red Teaming for Public-Test-Passing Code
arXiv cs.AI Security & Safety
Robust Critics: Defending LLMs Against Multi-Turn Attacks
arXiv cs.AI Security & Safety
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Mode…
TechCrunch AI Security & Safety
How AI guardrails are impeding the work of offensive cybersecurity researchers
Ars Technica AI Security & Safety
AI arms race in line for a reckoning after OpenAI hacking incident
arXiv cs.AI Security & Safety
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
arXiv cs.AI Security & Safety
Integrity of peer-to-peer distributed LLM inference under malicious nodes
arXiv cs.AI Security & Safety
HijackKV: New Threat in Position-Independent KV Cache Reuse
arXiv cs.AI Security & Safety
An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intellige…
arXiv cs.AI Security & Safety
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception …
arXiv cs.AI Security & Safety
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
arXiv cs.AI Security & Safety
ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in…
arXiv cs.AI Security & Safety
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
arXiv cs.AI Security & Safety
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language…
arXiv cs.AI Security & Safety
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems