MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clini…
Explorar
Noticias de IA
30329 elementos — filtrados, clasificados y sin duplicados
Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion
In-Context Reward Adaptation for Robust Preference Modeling
On Language Generation in the Limit with Bounded Memory
Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback
PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game?
LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Langua…
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Da…
Benchmarking at the Edge of Comprehension
Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangle…
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approxi…
Recurrent Structural Policy Gradient for Partially Observable Mean Field Games
DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization
When Models Learn to Ask Why: Adaptive Causal Reasoning for Trustworthy Medical Vision-La…
MediHive: A Decentralized Agent Collective for Medical Reasoning
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Po…
The Planetary Cost of AI Acceleration, Part II: The 10th Planetary Boundary and the 6.5-Y…
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With N…
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodi…
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
GPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activit…
ParaTool: Shifting Tool Representations from Context to Parameters
A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search