Reasoning with Sampling: Cutting at Decision Points
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Real-rootedness of the Poincar\'e polynomials of $\overline{\mathcal M}_{0,n}$: an AI-ass…
Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language…
CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimi…
Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review
LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models
AutoSizer: Automatic Sizing of Analog and Mixed-Signal Circuits via Large Language Model …
Online Fair Division with Additional Information
Thinking Before Constraining: A Unified Decoding Framework for Large Language Models
EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics
Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Polic…
MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs
On the Optimizer Dependence of Neural Scaling Laws
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smoot…
Evolutionary Rule Extraction from Corporate Default Prediction Models
PRO-CUA: Process-Reward Optimization for Computer Use Agents
Battery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter Estimation
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Divers…
Pushing the Limits of Block Rotations in Post-Training Quantization
A Composable Multimodal Framework for cine CMR-Text-Driven Prediction of Heart Failure Ou…
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consum…
DenseSteer: Steering Small Language Models towards Dense Math Reasoning
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Ag…
ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Mo…