Tuesday, September 01, 2026

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

This could be an interesting new paper by Li Fei Fei and her team!

From the abstract:
"Robot actions are inherently embodiment-specific and only weakly aligned with image-space visual changes, limiting their effectiveness as conditioning signals for robot world models.
In contrast, visual tracks provide an embodiment-agnostic representation of how task-relevant points move through a scene, offering dense image-space guidance for accurate and spatially precise future video prediction.
Building on this observation, we propose TrAct, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction.
TrAct consists of three components:
a Vision-Language-Action-and-Track model (VLAT) that jointly predicts candidate actions and corresponding visual tracks from the current observation and language instruction;
a track-conditioned world model (TWM) that predicts future visual outcomes conditioned on the proposed tracks; and 
a vision-language reward model (VLAC) that scores the predicted outcomes.
At inference time, VLAT generates candidate action-track pairs, TWM rolls out their visual consequences, and VLAC selects the track whose predicted outcome best satisfies the instruction; the action paired with the selected track is then executed by the robot.
Experiments on the proposed LIBERO-INTEGRAL benchmark and real-world Franka manipulation show that TrAct improves success rates from 27% to 55% in simulation and from 49% to 76% on real-world tasks compared with the strong VLA baseline π0.5.
Furthermore, TWM consistently improves video prediction quality over the action-conditioned world model (AWM).
These results demonstrate that visual tracks provide an effective shared interface between robot control and visual prediction, enabling more accurate world modeling and stronger robot generalization."

[2608.24101] TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks (preprint, open access)






No comments: