Agentic self-driving microscopy benchmarks support qualification but do not necessarily g…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decompo…
WorldClaw: Agentic 3D Open-World Generation at Scale
Coherence-Oriented Dream Scene Visualisation
MACRO: Markov Chain Routing of Transformer Layers
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Basel…
A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decompo…
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Ma…
Measuring and Detecting Harmful AI Sycophancy
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance …
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limit…
Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing …
The Impossibility Triangle of Long-Context Modeling
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large La…
SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering
Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answe…
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reas…
Layer-wise Positional Bias in Short-Context Language Modeling
All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Gener…
Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centri…
Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with N…
ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distributio…
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Mode…