Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Learning Intrusion Response Strategies for OT Systems
arXiv cs.AI Security & Safety
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
arXiv cs.AI Security & Safety
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
arXiv cs.AI Security & Safety
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks
arXiv cs.AI Security & Safety
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents
arXiv cs.AI Security & Safety
Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through th…
arXiv cs.AI Security & Safety
CS-Guard: Benchmarking LLM Guardrails for Code Generation Security
arXiv cs.AI Security & Safety
Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models
arXiv cs.AI Security & Safety
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Disco…
Ars Technica AI Security & Safety
Six Chinese AI firms accused of aggressively copying US frontier models
Wired AI Security & Safety
I Let an AI Agent Hack All My Gadgets—and I’d Do It Again
Bloomberg Technology Security & Safety
Anthropic Worker Resigns, Warns of AI Risks to Humanity
Bloomberg Technology Security & Safety
Ledger Hires New Security Chief as Crypto Hacks Hit $1.4 Billion
Ars Technica AI Security & Safety
Man told ChatGPT he was feeling delusional. ChatGPT insisted he was Jesus.
The Verge AI Security & Safety
Worried Anthropic researchers warn that AI ‘could kill all humans’
Mashable Security & Safety
Anthropic researcher quits, says AI could kill us all by the end of the decade
AINews / smol.ai Security & Safety
not much happened today
arXiv cs.AI Security & Safety
Versioned Transitive Dependency-Closure Binding and Operation-Time Effect Governance for …
arXiv cs.AI Security & Safety
Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning
arXiv cs.AI Security & Safety
Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts
arXiv cs.AI Security & Safety
Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Edi…
arXiv cs.AI Security & Safety
SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction
arXiv cs.AI Security & Safety
Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces
arXiv cs.AI Security & Safety
Hardware Trojan Threats to Multi-Chiplet Photonic Neural Network Accelerators
arXiv cs.AI Security & Safety
From Review to Authorization: Key-Isolated Threshold Signing for LLM Agents
arXiv cs.AI Security & Safety
PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base Invariants
arXiv cs.AI Security & Safety
Towards a Resilience-Theoretic Foundation for Adversarial Robustness in Industrial Contro…
arXiv cs.AI Security & Safety
How to Backdoor Image Knowledge Distillation
arXiv cs.AI Security & Safety
FATS: A Prompt Injection Attack Utilizing Feign Security Agents with Deceptive Few-shots …
arXiv cs.AI Security & Safety
Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Me…