Advance Zero Trust for AI: New tools and guidance to secure AI agents and DevSecOps
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery
Researchers watched OpenAI, Anthropic models take extreme measures in hacking test
Cybersecurity Concerns After OpenAI, Anthropic Tests
Rogue AI agents created fake online identities in another hacking attempt
OpenAI, Anthropic Model Tests Reveal More ‘Unsanctioned’ Actions
OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research i…
A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Determinist…
Privacy-Preserving AI Verification via Minimal Information Disclosure
DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial
Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, C…
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Sy…
LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
AI Security Leaderboard: Methodology, Results and Minimal Standard
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC…
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack C…
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Ag…
Risky Business: Measuring The Faithfulness-Safety Tension
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
Attribute-based Undetectable Watermarking for Generative AI Models
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detecti…