BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned S…
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Fed up with vibe coders, dev sneaks data-nuking prompt injection into their code
Symmetry Defeats Auditing
Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Dire…
Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Conte…
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversit…
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
Backdoor Attacks on Fault Detection and Localization in Cyber-Physical Systems
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level R…
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level R…
SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adv…
Indian Government, Tech Firms Running Tests for Mythos Threat
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It?
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models