DELE-w0.5 controlling a dual-arm robot at a microwave
RESEARCH

TECHNICAL REPORT · 2026

DELE-w0.5

Inferring Action from Future Latent State for Robotic Manipulation

Fenghao Lei · Zhixiong Huang · Long Yang · Jiabao Chen · Jie Cheng · Peilin Huang · Cong Fang · Han Fu · Zhuo Li · Xiaoxue Ren

Learn the consequence of action

Dense video prediction spends capacity reconstructing how change looks. DELE-w0.5 jointly learns action and an action-relevant future latent, focusing its representation on the physical outcome.

Future prediction exists only during training. At inference, language and current vision map directly to action.

Paper Figure 1 comparing VLA, video-predictive approaches and DELE-w0.5

Joint learning for direct action

Language and current vision condition a shared model. During training, it jointly denoises action and future visual latents; masked attention ensures action never reads future targets.

At inference, future-state tokens are removed. The deployed policy maps language and vision directly to action.

Paper Figure 2 showing the DELE-w0.5 condition and noise streams
Paper Figure 3 showing training and inference attention masks

Asymmetric attention lets future supervision shape the representation without entering the action path.

Evaluation

Pick-and-place is a primitive, not a destination. We evaluate complete household interactions—opening, retrieval, handoff and tool use—because a robot should be judged by the outcome it delivers, not an isolated motion it performs.

Every demonstration is shown uncut at 1× speed.

Door opening

Contact-rich manipulation that couples a stable handle grasp, timely rotation and a sustained push.

Pepsi retrieval

A multi-stage sequence spanning refrigerator interaction, object extraction, bimanual handoff and closure.

Popcorn retrieval

A complete microwave interaction combining articulated opening, bimanual retrieval, door closure and terminal-state completion.

Adding ice

Tool use and repeated bimanual coordination across scoop pickup, ice acquisition, pouring and scene restoration.

Early signs of emergence

Embodied intelligence has not reached its GPT-3 moment. Yet under interventions absent from training, DELE-w0.5 begins to show something beyond imitation: it holds the objective and composes a new response.

When the microwave door is closed mid-task, the model reopens it while retaining the package—or uses the package itself to recover access. Early evidence, not a claim. Emergence matters when it changes what a robot can do.

Reopening while retaining the package

After the door is closed mid-task, the model retains the package and composes a new action to reopen it—an interaction absent from demonstrations.

Using the package to recover access

When access narrows, the model uses the package already in hand to push the door open instead of returning to the demonstrated handle routine.

Reliability across complete tasks

Under the same fine-tuning data and evaluation protocol, DELE-w0.5 reaches farther through every task and converts intermediate progress into complete execution more consistently than the compared policies.

62.5%
overall full-task success
81.3%
macro ordered-stage progress
87.5 ms
median core-model inference
Paper Figure 5 and Table 2 showing real-robot performance and efficiency
Paper Figure 6 showing stage-reach probability across four long-horizon tasks

A world-grounded foundation model for robot intelligence

DELE-w0.5 jointly learns future visual dynamics and action generation, building a shared representation of how actions reshape the world. At inference, it acts directly—without predicting future observations.

Read the paper