Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

TechCrunch AI Security & Safety
OpenAI’s rogue agents keep escaping, with no formal process to investigate them
Ars Technica AI Security & Safety
OpenAI agents discussed ways to escape their sandbox on public wiki
Ars Technica AI Security & Safety
Once popular for attacking AI, ASCII smuggling is embraced by spammers
TechCrunch AI Security & Safety
Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowl…
Engadget Security & Safety
Rogue OpenAI agents took over a German coding forum in a previously undisclosed hijacking
The Verge AI Security & Safety
Oh good, looks like yet another swarm of rogue AI agents from OpenAI
The Verge AI Security & Safety
Instagram’s AI detection is a mess (again)
Hacker News (AI filter) Security & Safety
OpenAI agents hijacked German website in previously undisclosed AI breakout
AINews / smol.ai Security & Safety
collusion.wiki
arXiv cs.AI Security & Safety
PatchBench: Evaluating AI Agents for Vulnerability Patching
arXiv cs.AI Security & Safety
A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Ha…
arXiv cs.AI Security & Safety
SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations …
arXiv cs.AI Security & Safety
Measuring Harmfulness of Computer-Using Agents
Wired AI Security & Safety
Nobody Is Saying Why OpenAI and Anthropic Had Outages Today
Engadget Security & Safety
SpaceXAI apologizes for outage that affected Grok and other 'compute partners'
TechCrunch AI Security & Safety
Abliteration.ai is making a business out of removing AI guardrails
Engadget Security & Safety
Anthropic automatically signs out Claude users to protect them from hackers
Engadget Security & Safety
PSA: Don't rely on AI to plan anything that could put your life at risk... like a mountai…
OpenAI News Security & Safety
Safety overview: GPT-6 Astra
The Verge AI Security & Safety
Researchers fear safety disaster ahead of OpenAI’s Astra release
Bloomberg Technology Security & Safety
NYSE Used Anthropic’s Project Glasswing to Find Cyber Flaws
Bloomberg Technology Security & Safety
OpenAI Faces New Lawsuits Linked to Shooting at Canadian School
TechCrunch AI Security & Safety
OpenAI faces 30 more lawsuits tied to Tumbler Ridge shooting
arXiv cs.AI Security & Safety
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderati…
arXiv cs.AI Security & Safety
Auditing Harness Tampering in Self-Improving Agents
arXiv cs.AI Security & Safety
The Safeguard Worked. Is the LLM System Safer?
arXiv cs.AI Security & Safety
Optimizing Byzantine Node Placement in Decentralized Federated Learning
arXiv cs.AI Security & Safety
Capability-Gated Language Models: Security Composes, Utility Does Not
arXiv cs.AI Security & Safety
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
arXiv cs.AI Security & Safety
SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems