Reward Modeling for Multi-Agent Orchestration
Explorar
Noticias de IA
22116 elementos — filtrados, clasificados y sin duplicados
Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight…
Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibili…
GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with V…
ReCal: Reward Calibration for RL-based LLM Routing
Speculative Rollback Correction for Quality-Diverse Web Agent Imitation
Improving Crash Frequency Prediction from Simulated Traffic Conflicts Using Machine Learn…
A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constr…
Graph Reduction in Multirelational Networks: A Spreading-Oriented Reduction Benchmark
Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs
HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection
Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
Token Complexity Theory for AI-Augmented Computing
Two-Layer Linear Auto-Regressive Models Estimate Latent States
LLM-Powered Personalized Glycemic Assessment in Type 2 Diabetes with Wearable Sensor Data
AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages
Agentic MPC for Semantic Control System Resynthesis
Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning
ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling
Localizing Anchoring Pathways in Language Models
Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: …
Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning
OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Respon…
TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for M…
JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications
Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in …
Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier
Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via St…
Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning