Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Explorar
Noticias de IA
22116 elementos — filtrados, clasificados y sin duplicados
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Polic…
LSTM based IoT Device Identification
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Sc…
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications
CoVar: Confidence-Variance-Guided Pseudo-Label Selection for Semi-Supervised Learning
Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Schedu…
SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learni…
Mind the Perspective: Let's Reason Recursively for Theory of Mind
Lung-R1: A Knowledge Graph-Guided LLM for Pulmonary Diagnostic Reasoning
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Da…
HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distil…
Automated Mediator for Human Negotiation: Pre-Mediation via a Structured LLM Pipeline
Sonar-TS: Search-Then-Verify Natural Language Querying for Time Series Databases
From Explicit Elements to Implicit Intent: A Predefined Library for Auditable Behavioral …
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeli…
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
APPO: Agentic Procedural Policy Optimization
Non-frontal face recognition using GANs and memristor-based classifiers
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simul…
Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding
Runtime Enforcement of Hybrid System Properties
Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral G…
Tabular Foundation Models for Clinical Survival Analysis via Survival-Aware Adaptation
Multi-Rate Mixture of Experts for Accelerating Liquid Neural Network Training
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuou…
Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Lea…
DuoBench: A Reproducible Benchmark for Bimanual Manipulation in Simulation and the Real W…