Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base Invariants
arXiv cs.AI Security & Safety
Towards a Resilience-Theoretic Foundation for Adversarial Robustness in Industrial Contro…
arXiv cs.AI Security & Safety
Unveiling Hidden Threats: Using Fractal Triggers to Boost Stealthiness of Distributed Bac…
arXiv cs.AI Security & Safety
Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Edi…
arXiv cs.AI Security & Safety
Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt O…
arXiv cs.AI Security & Safety
Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts
arXiv cs.AI Security & Safety
Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Me…
arXiv cs.AI Security & Safety
Versioned Transitive Dependency-Closure Binding and Operation-Time Effect Governance for …
arXiv cs.AI Security & Safety
A TTP by TTP Approach: Precise Malware Detection via Malicious TTP Recognition
arXiv cs.AI Security & Safety
How to Backdoor Image Knowledge Distillation
Bloomberg Technology Security & Safety
Cisco President on Defending Against AI Attacks
Ars Technica AI Security & Safety
Why this month's Microsoft patch release is a doozy
TechCrunch AI Security & Safety
Hackers are stealing Claude tokens from subscribers
Ars Technica AI Security & Safety
“This is the AI men actually use”: Meta ads pushed apps nudifying real teens
TechCrunch AI Security & Safety
Chrome is now shipping updates every 2 weeks as AI changes the security landscape
Bloomberg Technology Security & Safety
Meta Ran Hundreds of Ads Showing AI Child Sexual Abuse, NGO Says
Bloomberg Technology Security & Safety
Boston Scientific Warns of Financial Hit from Cyber Attack
Mashable Security & Safety
OpenAI is figuring out how to tell people when its agents go rogue
arXiv cs.AI Security & Safety
AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
arXiv cs.AI Security & Safety
CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
arXiv cs.AI Security & Safety
Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection
arXiv cs.AI Security & Safety
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalize…
arXiv cs.AI Security & Safety
Uncensored Open-weight Models: Redistribution as the Persistence Layer
arXiv cs.AI Security & Safety
Rethinking Indirect Prompt Injection as a Test-Time Search Problem
arXiv cs.AI Security & Safety
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation…
Engadget Security & Safety
OpenAI responds after report exposed another incident in which its AI agents went rogue
Mashable Security & Safety
Rogue AI agents commandeered German website and used it as a messaging board
The Verge AI Security & Safety
OpenAI admits to German wiki ‘incident’
Wired AI Security & Safety
OpenAI Agents Hacked Another Website
Latent Space (swyx) Security & Safety
[AINews] Collusion.wiki: A second undisclosed OpenAI agent swarm incident...