UR$^2$: Unify RAG and Reasoning through Reinforcement Learning
Explorar
Noticias de IA
21863 elementos — filtrados, clasificados y sin duplicados
Lean-GAP: A Dataset of Formalized Graduate Algebra Problems
ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models
Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretr…
The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs
When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning
Solipsistic Superintelligence is Unlikely to be Cooperative
Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning
Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Pers…
ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Model…
Evaluating Transformer and LSTM Frameworks for Prediction in Ungauged Basins
Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability
SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural …
Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Hu…
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitor…
Physics-Guided Policy Optimization with Self-Distillation
FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences
Relational Linearity is a Predictor of Hallucinations
Plan, Verify and Fill: A Structured Parallel Decoding Approach for Diffusion Language Mod…
RobotValues: Evaluating Household Robots When Human Values Conflict
AI-Generated Traces for Novice Programmers: Learning Effects and Learner Differences in a…
Coupled Local and Global World Models for Efficient First Order RL
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classifi…
InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning
WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making
Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path…
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapte…