The Safeguard Worked. Is the LLM System Safer?
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderati…
OpenAI confirms Astra has reached critical cyber threshold, but will be available soon
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier
Open AI’s Astra model is on the way—and very good at breaking into computer systems
Dropbox User Accounts Breached by Hackers Who Accessed Data
OpenAI delayed its new model’s development after the Hugging Face hack
Palo Alto Networks Tops Profit Outlook on AI Security Demand
OpenAI Will Limit Access to New Astra Model’s Cybersecurity Features
Why MCP servers are becoming AI’s newest attack surface
Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolvin…
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agen…
AgenTRIM: Tool Risk Mitigation for Agentic AI
Apple shares ‘shocking evidence’ against former employee accused of stealing company data…
Hugging Face hack could indicate cultural issues at OpenAI
Polish Intelligence Probes Fire at Top Drone Firm WB Electronics
FISGuard: Defending Against Membership Inference via Fixed Input Subspaces
ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refus…
OpenStamp: A Watermark for Open-Source Language Models
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and …
When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI
Author Explores Cybersecurity Risks of Driverless Cars
Smartphone LED detects hidden cameras with AI
The Cybersecurity Apocalypse Is Coming in ‘Months,’ AI Giants Warn
SentinelOne CEO on Earnings, AI's Cybersecurity Impact
He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them