Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation
arXiv cs.AI Security & Safety
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerab…
arXiv cs.AI Security & Safety
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
arXiv cs.AI Security & Safety
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
arXiv cs.AI Security & Safety
RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks
arXiv cs.AI Security & Safety
POISE: Position-Aware Undetectable Skill Injection on LLM Agents
arXiv cs.AI Security & Safety
AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 -- Engine-Level Properties (…
arXiv cs.AI Security & Safety
Hiding in Plain Floats: Steganographic Carriers for Indirect Prompt and Content Injection
arXiv cs.AI Security & Safety
An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Re…
arXiv cs.AI Security & Safety
Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
arXiv cs.AI Security & Safety
Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Expla…
arXiv cs.AI Security & Safety
Revisiting the shutdown problem
arXiv cs.AI Security & Safety
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Pro…
arXiv cs.AI Security & Safety
Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructu…
arXiv cs.AI Security & Safety
Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems
arXiv cs.AI Security & Safety
Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents
arXiv cs.AI Security & Safety
SHIELD-IDS: Structurally Heterogeneous Ensemble with Integrated Layered Defense for Intru…
arXiv cs.AI Security & Safety
Model Poisoning Against Federated Model Adaptation with Chain of Bit-Flips
arXiv cs.AI Security & Safety
FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing
arXiv cs.AI Security & Safety
Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configu…
arXiv cs.AI Security & Safety
Targeting World Models to Compromise Robot Learning Pipelines
arXiv cs.AI Security & Safety
SecureClaw: Clawing Back Control of LLM Agents
arXiv cs.AI Security & Safety
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Finger…
arXiv cs.AI Security & Safety
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attac…
arXiv cs.AI Security & Safety
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
arXiv cs.AI Security & Safety
PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits
arXiv cs.AI Security & Safety
The Confidence Trap: Calibration Attacks for Graph Neural Networks
arXiv cs.AI Security & Safety
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Compute…
arXiv cs.AI Security & Safety
Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
Ars Technica AI Security & Safety
For the 2nd time in weeks, Microsoft packages laced with credential stealer