Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidan…
RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
Constraint-Anchored Reasoning Traces
Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool…
RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Pl…
From Overload to Insights: How AI Agents Can Support Scientists in Analyzing Complex Data
Environment-free Synthetic Data Generation for API-Calling Agents
Lomekwi: Resource-Bounded Tool Discovery in LLM Agents
Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs
Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Aut…
When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Desc…
Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-…
Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction
A Diagnostic Framework for AI Agent Behavior
Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase…
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, …
Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinfo…
Coordinated Disentanglement with Iterative Mode Discovery Under Hidden Correlations
An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation…
Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
Intermittent Control Is Not Diluted Control: A Switching Effect in Artificial Agency
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During D…
Panache: One-Pass Motif Discovery at Every Window Length
Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA …
The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a R…