Door opening
Contact-rich manipulation that couples a stable handle grasp, timely rotation and a sustained push.

TECHNICAL REPORT · 2026
Inferring Action from Future Latent State for Robotic Manipulation
Fenghao Lei · Zhixiong Huang · Long Yang · Jiabao Chen · Jie Cheng · Peilin Huang · Cong Fang · Han Fu · Zhuo Li · Xiaoxue Ren
Dense video prediction spends capacity reconstructing how change looks. DELE-w0.5 jointly learns action and an action-relevant future latent, focusing its representation on the physical outcome.
Future prediction exists only during training. At inference, language and current vision map directly to action.

Language and current vision condition a shared model. During training, it jointly denoises action and future visual latents; masked attention ensures action never reads future targets.
At inference, future-state tokens are removed. The deployed policy maps language and vision directly to action.


Asymmetric attention lets future supervision shape the representation without entering the action path.
Pick-and-place is a primitive, not a destination. We evaluate complete household interactions—opening, retrieval, handoff and tool use—because a robot should be judged by the outcome it delivers, not an isolated motion it performs.
Every demonstration is shown uncut at 1× speed.
Contact-rich manipulation that couples a stable handle grasp, timely rotation and a sustained push.
A multi-stage sequence spanning refrigerator interaction, object extraction, bimanual handoff and closure.
A complete microwave interaction combining articulated opening, bimanual retrieval, door closure and terminal-state completion.
Tool use and repeated bimanual coordination across scoop pickup, ice acquisition, pouring and scene restoration.
Embodied intelligence has not reached its GPT-3 moment. Yet under interventions absent from training, DELE-w0.5 begins to show something beyond imitation: it holds the objective and composes a new response.
When the microwave door is closed mid-task, the model reopens it while retaining the package—or uses the package itself to recover access. Early evidence, not a claim. Emergence matters when it changes what a robot can do.
After the door is closed mid-task, the model retains the package and composes a new action to reopen it—an interaction absent from demonstrations.
When access narrows, the model uses the package already in hand to push the door open instead of returning to the demonstrated handle routine.
Under the same fine-tuning data and evaluation protocol, DELE-w0.5 reaches farther through every task and converts intermediate progress into complete execution more consistently than the compared policies.


DELE-w0.5 jointly learns future visual dynamics and action generation, building a shared representation of how actions reshape the world. At inference, it acts directly—without predicting future observations.