Explorar

Noticias de IA

37834 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question A…
arXiv cs.AI Research & Papers
SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Co…
arXiv cs.AI Research & Papers
Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selectio…
arXiv cs.AI Research & Papers
The Scaling Paradox in Human-AI Collaboration
arXiv cs.AI Research & Papers
AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
arXiv cs.AI Research & Papers
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
arXiv cs.AI Research & Papers
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language …
arXiv cs.AI Research & Papers
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their I…
arXiv cs.AI Research & Papers
Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucina…
arXiv cs.AI Research & Papers
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process…
arXiv cs.AI Research & Papers
Role Steering of Language Models for Social Simulations
arXiv cs.AI Research & Papers
Role-Decoupled Attention Residuals: Separating Matching and Content Retrieval Across Depth
arXiv cs.AI Research & Papers
HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive …
arXiv cs.AI Research & Papers
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimizat…
arXiv cs.AI Research & Papers
Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture S…
arXiv cs.AI Research & Papers
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Mode…
arXiv cs.AI Research & Papers
When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task …
arXiv cs.AI Research & Papers
Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations
arXiv cs.AI Research & Papers
CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding
arXiv cs.AI Research & Papers
BayesSeg: A Bayesian Optimization Framework for State Segmentation of Electricity Consump…
arXiv cs.AI Research & Papers
The Bayesian Reflex: A Predictive Coding Engine for Artificial Intelligence
arXiv cs.AI Research & Papers
F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Lang…
arXiv cs.AI Research & Papers
Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LL…
arXiv cs.AI Research & Papers
TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LL…
arXiv cs.AI Research & Papers
SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs
arXiv cs.AI Research & Papers
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
arXiv cs.AI Research & Papers
Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology …
arXiv cs.AI Research & Papers
Bayesian and Motivated Reasoning in AI Agents
arXiv cs.AI Research & Papers
RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learn…
arXiv cs.AI Research & Papers
Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates