From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monit…
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
Cognitive Threat Intelligence and Explainable Federated Security Analytics for distribute…
Data Flow Control: Data Safety Policies for AI Agents
AI Scientist Bengio: Building Systems We Don't Know How to Control
AI Scientist Bengio on Engineering Safer Agents
Police have yet to catch a thief who used a Waymo to steal yoga clothes
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monit…
AI leaders call for tougher protections against AI-aided bioweapons
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Mod…
From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in…
What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in A…
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
A Systematic Investigation of RL-Jailbreaking in LLMs
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfi…
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight…
Notarized Agents: Receiver-Attested Confidential Receipts for AI Agent Actions
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
CrowdStrike Hits Projections, Signals Resilient Cyber Demand
AI is fueling Reddits spam problem
Survey: Teens regularly see harmful content, messages on Snapchat
Researchers show how AI-powered worms could wreak havoc on the internet
FORGE: Multi-Agent Graduated Exploitation and Detection Engineering
DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair
A Hybrid Approach For Malware Classification Using Secondary Features Fusion
AI Agents Enable Adaptive Computer Worms
Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You …