Explorar

Noticias de IA

1359 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Explainable AI-Driven Cyber Risk Analytics and Model Reliability Assessment for Intellige…
arXiv cs.AI Security & Safety
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
arXiv cs.AI Security & Safety
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
arXiv cs.AI Security & Safety
Beyond Rewards in Reinforcement Learning for Cyber Defence
arXiv cs.AI Security & Safety
Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchm…
Bloomberg Technology Security & Safety
AI Scientist Bengio: Building Systems We Don't Know How to Control
Bloomberg Technology Security & Safety
AI Scientist Bengio on Engineering Safer Agents
Engadget Security & Safety
Police have yet to catch a thief who used a Waymo to steal yoga clothes
Hugging Face Daily Papers Security & Safety
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monit…
The Verge AI Security & Safety
AI leaders call for tougher protections against AI-aided bioweapons
arXiv cs.AI Security & Safety
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
arXiv cs.AI Security & Safety
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
arXiv cs.AI Security & Safety
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety
arXiv cs.AI Security & Safety
Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfi…
arXiv cs.AI Security & Safety
From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in…
arXiv cs.AI Security & Safety
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
arXiv cs.AI Security & Safety
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Mod…
arXiv cs.AI Security & Safety
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
arXiv cs.AI Security & Safety
Notarized Agents: Receiver-Attested Confidential Receipts for AI Agent Actions
arXiv cs.AI Security & Safety
What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in A…
arXiv cs.AI Security & Safety
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight…
arXiv cs.AI Security & Safety
A Systematic Investigation of RL-Jailbreaking in LLMs
Hugging Face Daily Papers Security & Safety
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
Bloomberg Technology Security & Safety
CrowdStrike Hits Projections, Signals Resilient Cyber Demand
Mashable Security & Safety
AI is fueling Reddits spam problem
Mashable Security & Safety
Survey: Teens regularly see harmful content, messages on Snapchat
Engadget Security & Safety
Researchers show how AI-powered worms could wreak havoc on the internet
arXiv cs.AI Security & Safety
High-Precision APT Malware Attribution with Out-of-Scope Resilience
arXiv cs.AI Security & Safety
AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses
arXiv cs.AI Security & Safety
Phantom Transfer: Data Poisoning can Survive Data-Level Defences