PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Instrumental convergence and power-seeking
VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation
Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
Revisiting the shutdown problem
POISE: Position-Aware Undetectable Skill Injection on LLM Agents
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerab…
RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configu…
An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Re…
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attac…
SecureClaw: Clawing Back Control of LLM Agents
Hiding in Plain Floats: Steganographic Carriers for Indirect Prompt and Content Injection
Targeting World Models to Compromise Robot Learning Pipelines
AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 -- Engine-Level Properties (…
The Confidence Trap: Calibration Attacks for Graph Neural Networks
SHIELD-IDS: Structurally Heterogeneous Ensemble with Integrated Layered Defense for Intru…
Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Expla…
FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing
RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-di…
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Finger…
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Compute…
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Pro…
Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems
Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructu…
Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents