Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

AWS Machine Learning Blog Security & Safety
Govern AI agent tool access with Amazon Bedrock AgentCore Gateway
Hacker News (AI filter) Security & Safety
How a Texas student blew the whistle on a rogue AI hacking attempt
Ars Technica AI Security & Safety
As demand for Meta AI glasses explodes, it’s harder to avoid creepy recordings
Mashable Security & Safety
AI audio deepfakes are leading new ai-impersonation scams
Hacker News (AI filter) Security & Safety
Guess which of these LLM outputs is watermarked
Ars Technica AI Security & Safety
Grok exfiltrates user data when malicious instructions are encrypted
Hugging Face Daily Papers Security & Safety
Inadvertent Context Leakage in Language Models
Hugging Face Daily Papers Security & Safety
TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Sch…
arXiv cs.AI Security & Safety
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio …
arXiv cs.AI Security & Safety
Breaking the weakest link to evade vision language models
Hugging Face Daily Papers Security & Safety
Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynami…
TechCrunch AI Security & Safety
Researchers say OpenAI revoked their access to limited cyber program
The Verge AI Security & Safety
OpenAI hit the brakes. Now what?
Wired AI Security & Safety
Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks
Ars Technica AI Security & Safety
Meta ran ads for an app promising to nudify female politicians
Hugging Face Daily Papers Security & Safety
Breaking the weakest link to evade vision language models
Hugging Face Daily Papers Security & Safety
When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alig…
arXiv cs.AI Security & Safety
Future-Back Threat Modeling: A Foresight-Driven Security Framework
arXiv cs.AI Security & Safety
Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Me…
arXiv cs.AI Security & Safety
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
arXiv cs.AI Security & Safety
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on …
arXiv cs.AI Security & Safety
Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
arXiv cs.AI Security & Safety
COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models
arXiv cs.AI Security & Safety
The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evi…
arXiv cs.AI Security & Safety
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
The Verge AI Security & Safety
Robin Williams’ Instagram account brought back to fight ‘AI abuse’
The Verge AI Security & Safety
OpenAI lays out new security changes after its AI hacked Hugging Face
Bloomberg Technology Security & Safety
AI Stress Tester: Models Have Crossed a ‘Threshold of Competency’
Wired AI Security & Safety
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
TechCrunch AI Security & Safety
OpenAI institutes new safeguards after Hugging Face breach