The Planetary Cost of AI Acceleration, Part II: The 10th Planetary Boundary and the 6.5-Y…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Po…
MediHive: A Decentralized Agent Collective for Medical Reasoning
When Models Learn to Ask Why: Adaptive Causal Reasoning for Trustworthy Medical Vision-La…
FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization
DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
Recurrent Structural Policy Gradient for Partially Observable Mean Field Games
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approxi…
Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangle…
Benchmarking at the Edge of Comprehension
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Da…
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Langua…
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning
PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game?
Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback
On Language Generation in the Limit with Bounded Memory
In-Context Reward Adaptation for Robust Preference Modeling
Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion
MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clini…
NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs
Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context…
BitTP: The Lightweight Trajectory Prediction Model with BitLLM for Edge-Devices
Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction
Review Arcade: On the Human Alignment and Gameability of LLM Reviews
Frontier LLM-based agents can overcome the ontology curation bottleneck for natural pheno…
VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
The Hamilton-Jacobi Theory of Deep Learning