Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
AI Agents Enable Adaptive Computer Worms
arXiv cs.AI Security & Safety
Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You …
arXiv cs.AI Security & Safety
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
arXiv cs.AI Security & Safety
MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for M…
arXiv cs.AI Security & Safety
D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
arXiv cs.AI Security & Safety
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittlen…
arXiv cs.AI Security & Safety
FORGE: Multi-Agent Graduated Exploitation and Detection Engineering
arXiv cs.AI Security & Safety
FlowGuard: Flow Matching for Identity-Independent Detection of Data-Free Model Stealing A…
arXiv cs.AI Security & Safety
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
arXiv cs.AI Security & Safety
NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in …
arXiv cs.AI Security & Safety
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on L…
arXiv cs.AI Security & Safety
A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
arXiv cs.AI Security & Safety
Inference Cost Attacks for Retrieval-Augmented Large Language Models
arXiv cs.AI Security & Safety
A New Framework for Cybersecurity Refusals in AI Agents
arXiv cs.AI Security & Safety
Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing …
arXiv cs.AI Security & Safety
Narrow Secret Loyalty Dodges Black-Box Audits
arXiv cs.AI Security & Safety
A Hybrid Approach For Malware Classification Using Secondary Features Fusion
arXiv cs.AI Security & Safety
DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair
Hacker News (AI filter) Security & Safety
U of T researchers demonstrate AI worm could target any online device
TechCrunch AI Security & Safety
Google rolls out fake call detection to protect against AI deepfake impersonation scams
Hugging Face Daily Papers Security & Safety
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
TechCrunch AI Security & Safety
Anthropic scales Claude Mythos to critical infrastructure in 15+ countries
Hugging Face Daily Papers Security & Safety
RogueMerge: Robust and Unified Attacks against LLM Model Merging
arXiv cs.AI Security & Safety
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
arXiv cs.AI Security & Safety
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
arXiv cs.AI Security & Safety
CEAR: Certified Ensemble Adversarial Robustness in DNNs
arXiv cs.AI Security & Safety
Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
arXiv cs.AI Security & Safety
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
arXiv cs.AI Security & Safety
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
arXiv cs.AI Security & Safety
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Mode…