MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangl…
Explorar
Noticias de IA
21270 elementos — filtrados, clasificados y sin duplicados
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoni…
Quantization Degradation in Large Language Models: A Signal-Noise Perspective
DeepMind’s hurricane breakthrough has surprised weather scientists
Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation…
Classical $\mathrm{SU}(2)$ Models Match or Exceed Shallow Variational Quantum Circuits on…
TutorMoments: Do AI tutors know when to help and when to hold back?
Determining playoff clinching scenarios in the NHL using constraint programming
When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series
Stanford Evo 2 AI model generates phages against E. coli
Scientists Used AI to Create 16 New Viruses
Flow-Corrected Shape Optimization: Taming Manifold Drift in High-Dimensional 3D Models
MAUPITI: On-Device Prototype-Based Learning on a Smart Infrared Sensor
Recent advances in weakly supervised learning: New supervision paradigms, assumption rela…
HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
GSBF: Gaussian Splatting for Environment-Aware Beamforming
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and V…
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect…
VLMs for Videogame Data Annotation
CourseGraph: Finding overlaps and differences in Computer Science courses across universi…
Training a Conditioned Video Game Agent on a VLM Annotated Dataset
Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Pr…
Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bid…
OPERA: Operator-residual feedback for reliable autonomous optical experiments with langua…
Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Fo…