Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks
Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based L…
AI Security Leaderboard: Methodology, Results and Minimal Standard
OpenAI Hack Could Have Been 'Way Worse,' Hugging Face CEO Says
Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems
An AI-supervised remote exam went so badly that 58,000 students must retake it
Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spe…
SQLite Critical CVEs or LLM Slop?
Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone
Here’s why AI agents lie and cheat to reach their goals
The Cyber Alchemist's Ghosh on AI Cyber-security risks
DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models
Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent M…
Cyberattacks hit water facilities in seven states across the US
7 States’ Water Systems Hit by Cyberattacks Likely Tied to Iran
Nobody Knows if OpenAI’s and Anthropic’s AI Hacking Sprees Are Illegal
OpenAI reportedly finds evidence that more of its agents ran amok
Claude published malicious code to the Internet and attacked 3 real companies
Anthropic, OpenAI Cyber Failures Point to US Security Risks
High school defends staying silent while boys made AI nudes of 59 classmates
Anthropic Hack Adds To Fears Over AI Safety
It’s time to panic about AI safety
AI scammers outperform humans when it comes to building trust
AI labs want to pump the brakes, but Amazon and SpaceX are still blasting off
Anthropic says Claude accidentally hacked real companies too
A $2 sticker let me bypass the Meta Glasses' anti-creep feature
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Anthropic says its AI models also hacked three organizations on their own
What We Know So Far About Hacking by Anthropic AI Models