AI Firms Debate Putting Cyber Tests Online After Model Hacks
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
AI Firms Debate Putting Cyber Tests Online After Model Hacks
OpenAI subpoenaed by Alabama AG over Hugging Face hack
Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scamm…
Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds
On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal N…
Measuring Activation Control in Large Language Models
Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition
Disrupting a new covert influence campaign from Russia
Instinct’s powerful AI assistant is raising privacy and security concerns
Chinese Hackers Use DeepSeek to Boost Attacks, Researchers Say
China’s Hackers Use AI Tech to Lift Attacks, Researchers Say
Nvidia senior manager linked to Supermicro scheme smuggling AI servers to China
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
Taiwan Indicts Nvidia Manager Following Chip Smuggling Probe
They Dedicated Their Lives to Teaching. Then the Deepfakes Started
Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Prov…
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning…
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
Frontier AI labs still won’t say how they’d contain a rogue model
How Claude Watermarks AI-Generated Text
Anthropic’s Opus 4.6 is a smut-machine
Can a vehicle wrap hide your car from Flock cameras?