ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Dete…
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers…
ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
Instagram, Facebook Ran AI ‘Nudify’ Ads from China, Report Says
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack
OpenAI's rogue agent went on a hacking spree that lasted days, Reuters says
The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days
Did Chinese AI Steal From Anthropic, and OpenAI Loses Control of Two Models
Incomplete Prompt Jailbreaks in Large Language Models
Geometric Configurations of Perturbed Jailbreak Prompts
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-…
AI Security Policy Should Assess Systems, Not Only Models
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
Making Open-Source Text LLM Watermarks Durable Against Merging
Code Monitor Red Teaming for Public-Test-Passing Code
Robust Critics: Defending LLMs Against Multi-Turn Attacks
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Mode…
How AI guardrails are impeding the work of offensive cybersecurity researchers
AI arms race in line for a reckoning after OpenAI hacking incident
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
Integrity of peer-to-peer distributed LLM inference under malicious nodes
HijackKV: New Threat in Position-Independent KV Cache Reuse
An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intellige…
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception …
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in…
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language…
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems