Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

Wired AI Security & Safety
Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery
Mashable Security & Safety
Researchers watched OpenAI, Anthropic models take extreme measures in hacking test
Bloomberg Technology Security & Safety
Cybersecurity Concerns After OpenAI, Anthropic Tests
The Verge AI Security & Safety
Rogue AI agents created fake online identities in another hacking attempt
Bloomberg Technology Security & Safety
OpenAI, Anthropic Model Tests Reveal More ‘Unsanctioned’ Actions
Engadget Security & Safety
OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research i…
arXiv cs.AI Security & Safety
Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images
arXiv cs.AI Security & Safety
Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
arXiv cs.AI Security & Safety
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Ag…
arXiv cs.AI Security & Safety
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
arXiv cs.AI Security & Safety
ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Ind…
arXiv cs.AI Security & Safety
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Determinist…
arXiv cs.AI Security & Safety
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detecti…
arXiv cs.AI Security & Safety
$S^3$: Improving Agent Safety through Multi-Stage Defense
arXiv cs.AI Security & Safety
A Security-Oriented Lifecycle Model for Large Language Model Systems
arXiv cs.AI Security & Safety
Attribute-based Undetectable Watermarking for Generative AI Models
arXiv cs.AI Security & Safety
AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Age…
arXiv cs.AI Security & Safety
Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC…
arXiv cs.AI Security & Safety
Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, C…
arXiv cs.AI Security & Safety
Privacy-Preserving AI Verification via Minimal Information Disclosure
arXiv cs.AI Security & Safety
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Sy…
arXiv cs.AI Security & Safety
Risky Business: Measuring The Faithfulness-Safety Tension
arXiv cs.AI Security & Safety
LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
arXiv cs.AI Security & Safety
AI Security Leaderboard: Methodology, Results and Minimal Standard
arXiv cs.AI Security & Safety
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
arXiv cs.AI Security & Safety
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack C…
arXiv cs.AI Security & Safety
PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks
arXiv cs.AI Security & Safety
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
arXiv cs.AI Security & Safety
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
arXiv cs.AI Security & Safety
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity