DrugBench: Evaluating AI Control Protocols for Medication Harm Mitigation
Explorar
Noticias de IA
30590 elementos — filtrados, clasificados y sin duplicados
Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Foundation Models for Epileptogenic Zone Identification in Drug-Resistant Epilepsy
Only Ask What You Don't Know: Grounded Delta Planning for Efficient Multi-step RAG
Data Evolution by Wittgenstein's Rule Following
Confidently Wrong: Severity-Aware Calibration of Prompt-Injection Detectors under Attack …
Skill Coverage: A Test Adequacy Metric for Agent Skills
Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention
Mind the Noise: Sensitivity of Transformer-based Interaction-Aware Trajectory Prediction …
Social World Model for Lifelong Social Intelligence
Expected Free Energy-based Planning as Variational Inference
A-Evolve-Training: Autonomous Post-Training of a 30B Model
Learning Splitting Heuristics for Parallel String Solvers
Bridging Multi-Valued Heuristics and Dimensionality Reduction in Multi-Objective Search
Hypothesis-Disciplined Multi-Agent Automated Formalization of Asymptotic Statistical Theo…
An LLM-Explainable DRL Framework for Passenger-Directed Autonomous Driving
SkillHarness: Harnessing Safe Skills for Computer-Use Agents
Harnessing Agent Skills: Architectural Patterns and a Reference Architecture for Skill-Me…
Human Decision-Making with AI Assistance under Correlated Features
Latent Goal Prediction from Language for Model-Based Planning
In LLM Reasoning, there is Irrationality on top of Value Misalignment
Path-dependent program induction under resource constraints explains human sequence learn…
Darwin Mobile Agent: A Roadmap for Self-Evolution
PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate
The New Associationism: Lessons from Deep Learning
Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought …
On the Identifiability of User Adaptation in Co-Adaptive Neural Interfaces
A Hybrid TGN-SEAL Model for Dynamic Graph Link Prediction
ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots