TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System
Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brai…
Nexus: Structured Synergy for Efficient Text-to-Image Generation using Rectified Flow Mod…
Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and …
OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Re…
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise …
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert …
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
AI Evaluation Should Work With Humans
Measuring Cross-Task Behavioral Consistency in Language Model Agents
Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, a…
Active Perception for Embodied Disambiguation
No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation o…
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
Reward Machines for Signal Temporal Logic
How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generatin…
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
Exploring ESC Winners with Nested Diagrams
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Coverage Aware Active Evaluation for Failure Discovery with Paired Systems
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defens…
FLARE MCMC: Fidelity-based Layer-Adaptive REcursive proposals for MCMC
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small La…
SDO: Subspace Deconflicting Operator for Multi-Adapter Composition
Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends