Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for P…
Explorar
Noticias de IA
21863 elementos — filtrados, clasificados y sin duplicados
R-APS: Compositional Reasoning and In-Context Meta-Learning for Constrained Design via Re…
Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1)
Dual Advantage Fields
Description-Code Inconsistency in Real-world MCP Servers: Measurement, Detection, and Sec…
Abduction Prover in Isabelle/HOL
Large Language Models Hack Rewards, and Society
Constraint-Enhanced Physical Search through Correlation Matching
LCSHBench: A Multilingual, Consensus-Grounded Benchmark for Library of Congress Subject H…
Strabo: Declarative Specification and Implementation of Agentic Interaction Protocols
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats
What Type of Inference is Active Inference?
Semantic Constraint Synthesis for Adaptive Trajectory Optimization via Large Language Mod…
HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite
AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
BRAINCELL-AID: An Agentic AI Created Brain Cell Type Resource for Community Annotation
EvalStop: Using World Feedback to Detect and Correct Reward Overoptimization in Multi-Ten…
Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinfor…
MimeLens: Position-Agnostic Content-Type Detection for Binary Fragments
Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learni…
LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment
Exact Unlearning in Reinforcement Learning
Metric-Aware Hybrid Forecasting for the CTF4Science Lorenz Challenge
SSSD: Simply-Scalable Speculative Decoding
AIP: A Graph Representation for Learning and Governing Agent Skills
PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification
Incremental Sheaf Cohomology on Cellular Complexes: O(1)-in-n Lazy Edit Processing under …
MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enter…
Aligning Deep Implicit Preferences by Learning to Reason Defensively
Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC…