Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models
arXiv cs.AI Security & Safety
Cryptographic Registry Provenance: Structural Defense Against Dependency Confusion in AI …
arXiv cs.AI Security & Safety
Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
arXiv cs.AI Security & Safety
Lessons from Penetration Tests on Large-Scale Agent Systems
arXiv cs.AI Security & Safety
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
arXiv cs.AI Security & Safety
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbr…
Hugging Face Daily Papers Security & Safety
MRMMIA: Membership Inference Attacks on Memory in Chat Agents
Hugging Face Daily Papers Security & Safety
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
Ars Technica AI Security & Safety
Millions of AI agents imperiled by critical vulnerability in open source package
Ars Technica AI Security & Safety
FBI agent explains how easy it is to ID people posting AI porn without consent
Hugging Face Daily Papers Security & Safety
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
Hugging Face Daily Papers Security & Safety
Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models
Bloomberg Technology Security & Safety
BNP Paribas Works With Mistral to Prep for Mythos-Like AI Models
arXiv cs.AI Security & Safety
How does Bayesian Sampling help Membership Inference Attacks?
arXiv cs.AI Security & Safety
Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malw…
arXiv cs.AI Security & Safety
Batch Normalization Amplifies Memorization and Privacy Risks
arXiv cs.AI Security & Safety
AI-Driven Adaptive Adversaries and the Erosion of Cryptographic Trust in Public Key Syste…
arXiv cs.AI Security & Safety
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
arXiv cs.AI Security & Safety
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
arXiv cs.AI Security & Safety
Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LL…
arXiv cs.AI Security & Safety
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
arXiv cs.AI Security & Safety
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
arXiv cs.AI Security & Safety
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluat…
arXiv cs.AI Security & Safety
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
arXiv cs.AI Security & Safety
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures
arXiv cs.AI Security & Safety
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Age…
arXiv cs.AI Security & Safety
Demystifying the Mythos or Disrupting Bugonomics? From Zero-Day Asymmetry to Defender Rem…
arXiv cs.AI Security & Safety
Attested Tool-Server Admission: A Security Extension to the Model Context Protocol
arXiv cs.AI Security & Safety
Enhancing Reliability in LLM-Based Secure Code Generation
arXiv cs.AI Security & Safety
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods