Explorar

Noticias de IA

27413 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
Agentic self-driving microscopy benchmarks support qualification but do not necessarily g…
arXiv cs.AI Research & Papers
OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decompo…
arXiv cs.AI Research & Papers
WorldClaw: Agentic 3D Open-World Generation at Scale
arXiv cs.AI Research & Papers
Coherence-Oriented Dream Scene Visualisation
arXiv cs.AI Research & Papers
MACRO: Markov Chain Routing of Transformer Layers
arXiv cs.AI Research & Papers
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
arXiv cs.AI Research & Papers
Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Basel…
arXiv cs.AI Research & Papers
A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decompo…
arXiv cs.AI Research & Papers
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
arXiv cs.AI Research & Papers
Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Ma…
arXiv cs.AI Research & Papers
Measuring and Detecting Harmful AI Sycophancy
arXiv cs.AI Research & Papers
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance …
arXiv cs.AI Research & Papers
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
arXiv cs.AI Research & Papers
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
arXiv cs.AI Research & Papers
Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
arXiv cs.AI Research & Papers
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
arXiv cs.AI Research & Papers
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limit…
arXiv cs.AI Research & Papers
Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing …
arXiv cs.AI Research & Papers
The Impossibility Triangle of Long-Context Modeling
arXiv cs.AI Research & Papers
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large La…
arXiv cs.AI Research & Papers
SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering
arXiv cs.AI Research & Papers
Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answe…
arXiv cs.AI Research & Papers
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reas…
arXiv cs.AI Research & Papers
Layer-wise Positional Bias in Short-Context Language Modeling
arXiv cs.AI Research & Papers
All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Gener…
arXiv cs.AI Research & Papers
Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centri…
arXiv cs.AI Research & Papers
Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with N…
arXiv cs.AI Research & Papers
ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distributio…
arXiv cs.AI Research & Papers
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
arXiv cs.AI Research & Papers
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Mode…