Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Researchers watched OpenAI, Anthropic models take extreme measures in hacking test
Cybersecurity Concerns After OpenAI, Anthropic Tests
Rogue AI agents created fake online identities in another hacking attempt
OpenAI, Anthropic Model Tests Reveal More ‘Unsanctioned’ Actions
OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research i…
Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images
Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Ag…
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Ind…
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Determinist…
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detecti…
$S^3$: Improving Agent Safety through Multi-Stage Defense
A Security-Oriented Lifecycle Model for Large Language Model Systems
Attribute-based Undetectable Watermarking for Generative AI Models
AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Age…
Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC…
Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, C…
Privacy-Preserving AI Verification via Minimal Information Disclosure
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Sy…
Risky Business: Measuring The Faithfulness-Safety Tension
LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
AI Security Leaderboard: Methodology, Results and Minimal Standard
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack C…
PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity