Govern AI agent tool access with Amazon Bedrock AgentCore Gateway
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
How a Texas student blew the whistle on a rogue AI hacking attempt
As demand for Meta AI glasses explodes, it’s harder to avoid creepy recordings
AI audio deepfakes are leading new ai-impersonation scams
Guess which of these LLM outputs is watermarked
Grok exfiltrates user data when malicious instructions are encrypted
Inadvertent Context Leakage in Language Models
TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Sch…
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio …
Breaking the weakest link to evade vision language models
Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynami…
Researchers say OpenAI revoked their access to limited cyber program
OpenAI hit the brakes. Now what?
Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks
Meta ran ads for an app promising to nudify female politicians
Breaking the weakest link to evade vision language models
When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alig…
Future-Back Threat Modeling: A Foresight-Driven Security Framework
Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Me…
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on …
Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models
The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evi…
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
Robin Williams’ Instagram account brought back to fight ‘AI abuse’
OpenAI lays out new security changes after its AI hacked Hugging Face
AI Stress Tester: Models Have Crossed a ‘Threshold of Competency’
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
OpenAI institutes new safeguards after Hugging Face breach