← Todas las noticias

Latent Actions from Factorized Transition Effects under Agent Ambiguity

arXiv:2606.30544v2 Announce Type: replace Abstract: Latent Action Models (LAMs) learn action-like proxies from observation. However, in multi-object or distractor-rich scenes, observations contain not only agent motion but also distractors, camera dynamics, and background changes, making recovery of the underlying action intrinsically ambiguous without supervision. We argue that the appropriate unsupervised target is therefore not the true action itself, but a state-conditioned compositional summary of the transition effects present in the scene, enabling better action alignment and more effective utilization of action supervision. To this end, we propose a two-stage framework. We first pretrain Observed Transition Factorization (OTF) to discover reusable local transition primitives using a compositional codebook. We then aggregate these primitives into compact state-conditioned latent actions, instantiated as OTF-LAM-Pixel under the standard inverse-forward dynamics framework and OTF-LAM-Dino, a decoder-free variant operating in a frozen DINOv2 representation space. Experiments show that the learned transition primitives transfer across visual appearance and morphology, while the resulting latent actions exhibit substantially stronger action alignment than existing LAMs, make more effective use of downstream action supervision, and achieve competitive or superior downstream policy performance.
Leer el original en arXiv cs.AI →