Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Symmetry Defeats Auditing
arXiv cs.AI Security & Safety
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
arXiv cs.AI Security & Safety
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
arXiv cs.AI Security & Safety
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Conte…
arXiv cs.AI Security & Safety
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
arXiv cs.AI Security & Safety
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
arXiv cs.AI Security & Safety
Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
arXiv cs.AI Security & Safety
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
arXiv cs.AI Security & Safety
Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Dire…
arXiv cs.AI Security & Safety
Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
arXiv cs.AI Security & Safety
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
arXiv cs.AI Security & Safety
Backdoor Attacks on Fault Detection and Localization in Cyber-Physical Systems
arXiv cs.AI Security & Safety
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
arXiv cs.AI Security & Safety
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversit…
arXiv cs.AI Security & Safety
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
arXiv cs.AI Security & Safety
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level R…
arXiv cs.AI Security & Safety
Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels
Hugging Face Daily Papers Security & Safety
A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG
Hugging Face Daily Papers Security & Safety
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level R…
Hugging Face Daily Papers Security & Safety
SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adv…
Bloomberg Technology Security & Safety
Indian Government, Tech Firms Running Tests for Mythos Threat
arXiv cs.AI Security & Safety
Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
arXiv cs.AI Security & Safety
Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models
arXiv cs.AI Security & Safety
Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Ima…
arXiv cs.AI Security & Safety
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
arXiv cs.AI Security & Safety
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
arXiv cs.AI Security & Safety
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbr…
arXiv cs.AI Security & Safety
Cryptographic Registry Provenance: Structural Defense Against Dependency Confusion in AI …
arXiv cs.AI Security & Safety
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
arXiv cs.AI Security & Safety
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack