Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
HauntAttack: When Attack Follows Reasoning as a Shadow
arXiv cs.AI Security & Safety
Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and…
arXiv cs.AI Security & Safety
From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Commu…
arXiv cs.AI Security & Safety
Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk Go…
arXiv cs.AI Security & Safety
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critic…
arXiv cs.AI Security & Safety
Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities
arXiv cs.AI Security & Safety
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities In…
Hacker News (AI filter) Security & Safety
What happened after 2k people tried to hack my AI assistant
Ars Technica AI Security & Safety
Anthropic says Alibaba must be punished for largest Claude cloning attack
Engadget Security & Safety
Russia allegedly used a forensics platform to hack an activist's phone, despite having it…
arXiv cs.AI Security & Safety
A Marketplace for AI-Generated Adult Content and Deepfakes
arXiv cs.AI Security & Safety
What Does It Mean to Break a Distillation Defense?
arXiv cs.AI Security & Safety
The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapab…
arXiv cs.AI Security & Safety
SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward
arXiv cs.AI Security & Safety
Verifiable Manifest Signing and Transparency Enforcement for Secure MCP-Based LLM Pipelin…
arXiv cs.AI Security & Safety
Epistemic Bias Injection: Manipulating LLM Opinion via Selective Context Retrieval
arXiv cs.AI Security & Safety
Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection o…
arXiv cs.AI Security & Safety
Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study
arXiv cs.AI Security & Safety
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
arXiv cs.AI Security & Safety
PVF:Understanding AI Vulnerability Against SDCs
arXiv cs.AI Security & Safety
A Hybrid CNN-LSTM Intrusion Detection Framework for Cybersecurity in Smart Renewable Ener…
arXiv cs.AI Security & Safety
Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks
Microsoft Source (AI + Cloud) Security & Safety
Scaling cybercrime disruption through innovation and AI
Hacker News (AI filter) Security & Safety
Anthropic says Alibaba illicitly extracted Claude AI model capabilities
Bloomberg Technology Security & Safety
Anthropic Accuses Alibaba of ‘Illicitly’ Accessing AI Models
Hugging Face Daily Papers Security & Safety
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
Bloomberg Technology Security & Safety
Microsoft Says Copilot AI Helped Knock Down Cybercrime Tools
Bloomberg Technology Security & Safety
Fight Against Child Predators Gets Short Shrift While Cases Explode in AI Era
Bloomberg Technology Security & Safety
Scammers Are Using AI to Create Fake Auto Loan Documents
Hugging Face Daily Papers Security & Safety
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Att…