Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Phantom Transfer: Data Poisoning can Survive Data-Level Defences
arXiv cs.AI Security & Safety
Narrow Secret Loyalty Dodges Black-Box Audits
arXiv cs.AI Security & Safety
High-Precision APT Malware Attribution with Out-of-Scope Resilience
arXiv cs.AI Security & Safety
MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for M…
arXiv cs.AI Security & Safety
Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing …
arXiv cs.AI Security & Safety
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on L…
arXiv cs.AI Security & Safety
AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses
arXiv cs.AI Security & Safety
NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in …
arXiv cs.AI Security & Safety
A New Framework for Cybersecurity Refusals in AI Agents
arXiv cs.AI Security & Safety
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
arXiv cs.AI Security & Safety
A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
arXiv cs.AI Security & Safety
FlowGuard: Flow Matching for Identity-Independent Detection of Data-Free Model Stealing A…
arXiv cs.AI Security & Safety
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittlen…
arXiv cs.AI Security & Safety
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
arXiv cs.AI Security & Safety
Inference Cost Attacks for Retrieval-Augmented Large Language Models
arXiv cs.AI Security & Safety
D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
Hacker News (AI filter) Security & Safety
U of T researchers demonstrate AI worm could target any online device
TechCrunch AI Security & Safety
Google rolls out fake call detection to protect against AI deepfake impersonation scams
Hugging Face Daily Papers Security & Safety
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
TechCrunch AI Security & Safety
Anthropic scales Claude Mythos to critical infrastructure in 15+ countries
Hugging Face Daily Papers Security & Safety
RogueMerge: Robust and Unified Attacks against LLM Model Merging
arXiv cs.AI Security & Safety
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
arXiv cs.AI Security & Safety
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
arXiv cs.AI Security & Safety
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box R…
arXiv cs.AI Security & Safety
SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems
arXiv cs.AI Security & Safety
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
arXiv cs.AI Security & Safety
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-A…
arXiv cs.AI Security & Safety
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
arXiv cs.AI Security & Safety
Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem
arXiv cs.AI Security & Safety
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills