When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodi…
Explorar
Noticias de IA
30334 elementos — filtrados, clasificados y sin duplicados
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
GPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activit…
ParaTool: Shifting Tool Representations from Context to Parameters
A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search
A Survey on Recent Advances in Conversational Data Generation
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Gener…
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Ev…
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
Dataset-Driven Channel Masks in Transformers for Multivariate Time Series
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
Steering Language Models Before They Speak: Logit-Level Interventions
GPIC: A Giant Permissive Image Corpus for Visual Generation
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
GiPL: Generative augmented iterative Pseudo-Labeling for Cross-Domain Few-Shot Object Det…
PhoneWorld: Scaling Phone-Use Agent Environments
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generati…
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Mul…
Composing Non-Conjugate Factor Graphs with Closed-Form Variational Inference
CORE-T: COherent REtrieval of Tables for Text-to-SQL
How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignmen…
TRACER: Persistent Regularization for Robust Multimodal Finetuning
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio