MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adver…
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Dr…
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Mo…
Benchmarking Agentic Review Systems
Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning
Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelli…
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
Optimal Scheduling in a Question-Answering Forum of Knowledge Workers
cAPM: Continual AI-Assisted Pace-Mapping with Active Learning
Multi-Agent Transactive Memory
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns
Deontic Policies for Runtime Governance of Agentic AI Systems
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifa…
How LLMs Fail and Generalize in RTL Coding for Hardware Design?
Augmenting Game AI with Deep Reinforcement Learning
Hidden Anchors in Multi-Agent LLM Deliberation
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Ann…
Thermodynamic Measure of Intelligence
Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerab…
AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, K…
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Boards (PCB) Schematic…
Efficient and Sound Probabilistic Verification for AI Agents
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Pla…
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
Context-Aware Hierarchical Bayesian Modeling of IVF Laboratory Environmental Conditions
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Mod…