Explorar

Noticias de IA

22318 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adver…
arXiv cs.AI Research & Papers
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Dr…
arXiv cs.AI Research & Papers
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Mo…
arXiv cs.AI Research & Papers
Benchmarking Agentic Review Systems
arXiv cs.AI Research & Papers
Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning
arXiv cs.AI Research & Papers
Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelli…
arXiv cs.AI Research & Papers
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
arXiv cs.AI Research & Papers
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
arXiv cs.AI Research & Papers
Optimal Scheduling in a Question-Answering Forum of Knowledge Workers
arXiv cs.AI Research & Papers
cAPM: Continual AI-Assisted Pace-Mapping with Active Learning
arXiv cs.AI Research & Papers
Multi-Agent Transactive Memory
arXiv cs.AI Research & Papers
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns
arXiv cs.AI Research & Papers
Deontic Policies for Runtime Governance of Agentic AI Systems
arXiv cs.AI Research & Papers
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifa…
arXiv cs.AI Research & Papers
How LLMs Fail and Generalize in RTL Coding for Hardware Design?
arXiv cs.AI Research & Papers
Augmenting Game AI with Deep Reinforcement Learning
arXiv cs.AI Research & Papers
Hidden Anchors in Multi-Agent LLM Deliberation
arXiv cs.AI Research & Papers
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Ann…
arXiv cs.AI Research & Papers
Thermodynamic Measure of Intelligence
arXiv cs.AI Research & Papers
Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerab…
arXiv cs.AI Research & Papers
AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, K…
arXiv cs.AI Research & Papers
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Boards (PCB) Schematic…
arXiv cs.AI Research & Papers
Efficient and Sound Probabilistic Verification for AI Agents
arXiv cs.AI Research & Papers
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Pla…
arXiv cs.AI Research & Papers
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
arXiv cs.AI Research & Papers
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
arXiv cs.AI Research & Papers
Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution
arXiv cs.AI Research & Papers
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
arXiv cs.AI Research & Papers
Context-Aware Hierarchical Bayesian Modeling of IVF Laboratory Environmental Conditions
arXiv cs.AI Research & Papers
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Mod…