Phantom Transfer: Data Poisoning can Survive Data-Level Defences
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Narrow Secret Loyalty Dodges Black-Box Audits
High-Precision APT Malware Attribution with Out-of-Scope Resilience
MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for M…
Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing …
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on L…
AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses
NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in …
A New Framework for Cybersecurity Refusals in AI Agents
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
FlowGuard: Flow Matching for Identity-Independent Detection of Data-Free Model Stealing A…
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittlen…
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
Inference Cost Attacks for Retrieval-Augmented Large Language Models
D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
U of T researchers demonstrate AI worm could target any online device
Google rolls out fake call detection to protect against AI deepfake impersonation scams
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
Anthropic scales Claude Mythos to critical infrastructure in 15+ countries
RogueMerge: Robust and Unified Attacks against LLM Model Merging
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box R…
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-A…
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills