HSRM: Hidden-State Reward Models for Test-Time Verification
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustw…
Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization…
Which Rules Matter Now? Policy-Centroid Routing Before an Intelligent System Acts
HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Ta…
MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reason…
Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends
ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents
CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models
SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self…
MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for C…
TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI
Rate-Coding Bundle Memory: A Unified Model of Memory and Control for Symbolic Computation…
SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential G…
GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning
Reward-Oracle MCTS for Formal Theorem Proving: Sample-Efficient Search and the Need for K…
Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management
AI Can Be Easily Persuaded in Clinical Decision Making
DiffPDE: Masked Diffusion Language Models as PDE Solver
CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target
Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis
Learning-Assisted Congestion-Aware Route Scheduling for Semiconductor Fab Material Contro…
Evaluating Tiny Recursive Models Across Training for Code Generation
CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework
DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Ben…
Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evide…
Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Hori…
EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolvin…
Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-R…