OpenAI Makes AI Safety Changes in Wake of Hugging Face Breach
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
French Tax Office to Use AI to Probe Vulnerabilities After Hack
OpenAI president urges enterprises to hasten AI security defences
Microsoft Copilot reveals secret input that allowed it to be hacked
Pacing model development in an era of cyber-critical capabilities
not much happened today
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs
Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campai…
SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
Towards Risk-free AI Agent Deployment
Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Age…
Workspace Topology as an Attack Vector in Agentic Coding Assistants
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endog…
Synchronized Logit Steering: Real-world Steganography
Israel creates fake think tank in likely attempt to dupe AI chatbots
Digital Twin-Based Intrusion Detection for Vehicle Powertrain CAN Bus Systems
Grok CSAM lawsuit expands as more step forward
Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
Odd Lots: Is There An AI Kill Switch If Things Go Wrong?
AI-Generated GitHub Copilot "Autofix" Allowed Compromise of Snowflake's Jira
DSPrompt: Dynamic Soft Prompt Defense Against M-RAG Corruption
What the OpenAI/Hugging Face Hack Really Tells Us About AI Danger
Another woman joins lawsuit accusing Grok of generating CSAM
The Defender’s Window
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with…
Woman claims her stepfather used Grok to transform childhood photo into explicit imagery
How to tell if your AI platforms’ accounts have been hacked