Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned S…
Ars Technica AI Security & Safety
Fed up with vibe coders, dev sneaks data-nuking prompt injection into their code
arXiv cs.AI Security & Safety
Symmetry Defeats Auditing
arXiv cs.AI Security & Safety
Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Dire…
arXiv cs.AI Security & Safety
Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels
arXiv cs.AI Security & Safety
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
arXiv cs.AI Security & Safety
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
arXiv cs.AI Security & Safety
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Conte…
arXiv cs.AI Security & Safety
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
arXiv cs.AI Security & Safety
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversit…
arXiv cs.AI Security & Safety
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
arXiv cs.AI Security & Safety
Backdoor Attacks on Fault Detection and Localization in Cyber-Physical Systems
arXiv cs.AI Security & Safety
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
arXiv cs.AI Security & Safety
Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
arXiv cs.AI Security & Safety
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
arXiv cs.AI Security & Safety
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level R…
arXiv cs.AI Security & Safety
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
arXiv cs.AI Security & Safety
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
arXiv cs.AI Security & Safety
Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
Hugging Face Daily Papers Security & Safety
A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG
Hugging Face Daily Papers Security & Safety
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level R…
Hugging Face Daily Papers Security & Safety
SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adv…
Bloomberg Technology Security & Safety
Indian Government, Tech Firms Running Tests for Mythos Threat
arXiv cs.AI Security & Safety
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
arXiv cs.AI Security & Safety
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
arXiv cs.AI Security & Safety
Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
arXiv cs.AI Security & Safety
GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It?
arXiv cs.AI Security & Safety
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
arXiv cs.AI Security & Safety
Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
arXiv cs.AI Security & Safety
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models