Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Cryptographic Registry Provenance: Structural Defense Against Dependency Confusion in AI …
Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
Lessons from Penetration Tests on Large-Scale Agent Systems
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbr…
MRMMIA: Membership Inference Attacks on Memory in Chat Agents
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
Millions of AI agents imperiled by critical vulnerability in open source package
FBI agent explains how easy it is to ID people posting AI porn without consent
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models
BNP Paribas Works With Mistral to Prep for Mythos-Like AI Models
How does Bayesian Sampling help Membership Inference Attacks?
Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malw…
Batch Normalization Amplifies Memorization and Privacy Risks
AI-Driven Adaptive Adversaries and the Erosion of Cryptographic Trust in Public Key Syste…
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LL…
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluat…
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Age…
Demystifying the Mythos or Disrupting Bugonomics? From Zero-Day Asymmetry to Defender Rem…
Attested Tool-Server Admission: A Security Extension to the Model Context Protocol
Enhancing Reliability in LLM-Based Secure Code Generation
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods