AI Agents Enable Adaptive Computer Worms
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You …
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for M…
D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittlen…
FORGE: Multi-Agent Graduated Exploitation and Detection Engineering
FlowGuard: Flow Matching for Identity-Independent Detection of Data-Free Model Stealing A…
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in …
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on L…
A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
Inference Cost Attacks for Retrieval-Augmented Large Language Models
A New Framework for Cybersecurity Refusals in AI Agents
Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing …
Narrow Secret Loyalty Dodges Black-Box Audits
A Hybrid Approach For Malware Classification Using Secondary Features Fusion
DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair
U of T researchers demonstrate AI worm could target any online device
Google rolls out fake call detection to protect against AI deepfake impersonation scams
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
Anthropic scales Claude Mythos to critical infrastructure in 15+ countries
RogueMerge: Robust and Unified Attacks against LLM Model Merging
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
CEAR: Certified Ensemble Adversarial Robustness in DNNs
Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Mode…