TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
Detecting and Mitigating DDoS Attacks with AI: A Survey
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking …
Adversarial Attacks Leverage Interference Between Features in Superposition
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Revi…
Membership Inference Attacks against Large Audio Language Models
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Ad…
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
Graph neural networks at war: integrating cybersecurity and drone intelligence in the Isr…
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and …
This Copilot vulnerability could expose emails, 2FA codes, and other sensitive data
Securing the future of AI agents
Critical Copilot vulnerability allowed hackers to seal 2FA code from users
AI Red Teaming Explained: What It Is and Why You Need It
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference P…
Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
A Survey on Agentic Security: Applications, Threats and Defenses
Discrete optimal transport is a strong audio adversarial attack
Automated jailbreak attack targeting multiple defense strategies
GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?
InstantForget: Update-Free Backdoor Unlearning with Inference-Time Feature Reset
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query …
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source…
MASCOT-Android: A Curated Dataset and Automated Collection Pipeline for Android Malware S…
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection