Explainable AI-Driven Cyber Risk Analytics and Model Reliability Assessment for Intellige…
Explorar
Noticias de IA
1359 elementos — filtrados, clasificados y sin duplicados
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
Beyond Rewards in Reinforcement Learning for Cyber Defence
Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchm…
AI Scientist Bengio: Building Systems We Don't Know How to Control
AI Scientist Bengio on Engineering Safer Agents
Police have yet to catch a thief who used a Waymo to steal yoga clothes
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monit…
AI leaders call for tougher protections against AI-aided bioweapons
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety
Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfi…
From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in…
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Mod…
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
Notarized Agents: Receiver-Attested Confidential Receipts for AI Agent Actions
What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in A…
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight…
A Systematic Investigation of RL-Jailbreaking in LLMs
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
CrowdStrike Hits Projections, Signals Resilient Cyber Demand
AI is fueling Reddits spam problem
Survey: Teens regularly see harmful content, messages on Snapchat
Researchers show how AI-powered worms could wreak havoc on the internet
High-Precision APT Malware Attribution with Out-of-Scope Resilience
AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses
Phantom Transfer: Data Poisoning can Survive Data-Level Defences