Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question A…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Co…
Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selectio…
The Scaling Paradox in Human-AI Collaboration
AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language …
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their I…
Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucina…
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process…
Role Steering of Language Models for Social Simulations
Role-Decoupled Attention Residuals: Separating Matching and Content Retrieval Across Depth
HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive …
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimizat…
Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture S…
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Mode…
When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task …
Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations
CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding
BayesSeg: A Bayesian Optimization Framework for State Segmentation of Electricity Consump…
The Bayesian Reflex: A Predictive Coding Engine for Artificial Intelligence
F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Lang…
Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LL…
TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LL…
SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology …
Bayesian and Motivated Reasoning in AI Agents
RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learn…
Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates