VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerab…
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks
POISE: Position-Aware Undetectable Skill Injection on LLM Agents
AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 -- Engine-Level Properties (…
Hiding in Plain Floats: Steganographic Carriers for Indirect Prompt and Content Injection
An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Re…
Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Expla…
Revisiting the shutdown problem
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Pro…
Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructu…
Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems
Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents
SHIELD-IDS: Structurally Heterogeneous Ensemble with Integrated Layered Defense for Intru…
Model Poisoning Against Federated Model Adaptation with Chain of Bit-Flips
FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing
Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configu…
Targeting World Models to Compromise Robot Learning Pipelines
SecureClaw: Clawing Back Control of LLM Agents
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Finger…
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attac…
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits
The Confidence Trap: Calibration Attacks for Graph Neural Networks
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Compute…
Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
For the 2nd time in weeks, Microsoft packages laced with credential stealer