Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monit…
arXiv cs.AI Security & Safety
Cognitive Threat Intelligence and Explainable Federated Security Analytics for distribute…
arXiv cs.AI Security & Safety
Data Flow Control: Data Safety Policies for AI Agents
Bloomberg Technology Security & Safety
AI Scientist Bengio: Building Systems We Don't Know How to Control
Bloomberg Technology Security & Safety
AI Scientist Bengio on Engineering Safer Agents
Engadget Security & Safety
Police have yet to catch a thief who used a Waymo to steal yoga clothes
Hugging Face Daily Papers Security & Safety
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monit…
The Verge AI Security & Safety
AI leaders call for tougher protections against AI-aided bioweapons
arXiv cs.AI Security & Safety
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
arXiv cs.AI Security & Safety
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Mod…
arXiv cs.AI Security & Safety
From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in…
arXiv cs.AI Security & Safety
What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in A…
arXiv cs.AI Security & Safety
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
arXiv cs.AI Security & Safety
A Systematic Investigation of RL-Jailbreaking in LLMs
arXiv cs.AI Security & Safety
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
arXiv cs.AI Security & Safety
Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfi…
arXiv cs.AI Security & Safety
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety
arXiv cs.AI Security & Safety
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
arXiv cs.AI Security & Safety
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight…
arXiv cs.AI Security & Safety
Notarized Agents: Receiver-Attested Confidential Receipts for AI Agent Actions
Hugging Face Daily Papers Security & Safety
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
Bloomberg Technology Security & Safety
CrowdStrike Hits Projections, Signals Resilient Cyber Demand
Mashable Security & Safety
AI is fueling Reddits spam problem
Mashable Security & Safety
Survey: Teens regularly see harmful content, messages on Snapchat
Engadget Security & Safety
Researchers show how AI-powered worms could wreak havoc on the internet
arXiv cs.AI Security & Safety
FORGE: Multi-Agent Graduated Exploitation and Detection Engineering
arXiv cs.AI Security & Safety
DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair
arXiv cs.AI Security & Safety
A Hybrid Approach For Malware Classification Using Secondary Features Fusion
arXiv cs.AI Security & Safety
AI Agents Enable Adaptive Computer Worms
arXiv cs.AI Security & Safety
Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You …