Preprint · 2026

Flex-π: A Multi-Stream World-Action Model with Compute Flexibility

One checkpoint. Any observed inputs, any generated futures, and a speed–accuracy operating point you pick at deployment.

Ge Yan*1, Jinghao Liu*1, Yuzhi Fan*1, Lei Cai1, Minwen Liao1, Jesse Zhang†1, Dieter Fox†1,2

1University of Washington 2Allen Institute for AI

*Equal contribution.   Equal advising.

Problem: World-action models (WAMs) predict the future to act better, but they mainly predict RGB latents trained for pixel reconstruction — with no explicit signal for the 3D geometry or object semantics that manipulation needs.

Method: We introduce Flex-π, a 6B-parameter WAM that jointly denoises RGB, 3D geometry, and object-centric DINO semantics with actions in a shared latent space. Per-stream dropout with cross-modality forcing yields a single checkpoint that runs on any subset of streams — from fast action-only inference to full joint generation — with no new sensors or visual priors.

Result: A demonstration-efficient policy that beats the strongest baselines by up to 2–7× on dexterous, precise, real-world bimanual tasks in and out of distribution — while running faster than π0.5.

Flex-π in action

Deployed on a bimanual YAM workcell across contact-rich, sub-millimeter and long-horizon tasks, including gripper self-repair.

Autonomous, 1.7x speed

Task completion on the real robot, averaged over the tasks in each panel. The line is the best baseline; the dark green is Flex-π's margin over it.

View as table
Task Flex-π (full joint) Flex-π (action-only) ManiFlow π0.5 Fast-WAM
Seen tasks 5 tasks, in distribution
Put Plate on the Rack95.084.275.872.512.5
Sort Utensils75.070.055.045.05.0
Kitchen Organization98.896.293.873.877.5
Self-Repair Gripper76.066.933.326.2
Soft-Bag Zipping70.064.931.942.8
Average83.076.458.052.131.7
Unseen conditions 3 tasks
Put Plate on the Rack95.085.055.072.533.8
Sort Utensils70.070.032.540.00.0
Soft-Bag Zipping63.357.56.917.2
Average76.170.831.543.216.9
Half the data 1 task
Put Plate on the Rack95.080.060.042.525.0

Task completion (%). Fast-WAM was not run on Self-Repair Gripper or Soft-Bag Zipping, so its averages cover 3 of 5 seen tasks and 2 of 3 unseen conditions.

Click on any video below to expand it.

Put Plate on the Rack
Sort Utensils
Kitchen Organization
Self-Repair Gripper
Soft-Bag Zipping

Three views of one moment

Given 1 RGB image, Flex-π encodes into latent visual streams for appearance, semantics, and geometry, and can generate all three futures alongside the action. What goes in and what comes out are runtime arguments rather than architecture, so all 56 combinations are deployable from the same weights — try them below. How that is trained is in How it works.

RGB · appearanceThe RGB image, encoded by a frozen Wan-2.2 VAE.
DINO · object semanticsObject-centric semantics, from a frozen DINOv3 encoder.
Pointmap · 3D geometry3D geometry, encoded by the same frozen Wan-2.2 VAE as the RGB.
Multi-stream world-action model
observed inputs generated futures init. Wan-2.2 Multi-Stream MoT 5B trunk RGB Wan VAE DINO DINOv3 Pointmap same VAE RGB future DINO future 3D future action expert · 1B action chunk
Inference latency
60ms
Faster than every baseline we compare against.
Task completion
76.4%
+18.4 points over the strongest baseline.
jump to
observed min
generated mout a
Action-only fast path

No future visual stream is read, so none is computed. This is the cheapest point on the frontier and recovers VLA-level latency.

configuration 49 of 56
Part I

What Flex-π does

An eight-stage self-repair, five contact-rich tasks scored against three baselines, and three conditions it was never fine-tuned on.

Flex-π leads on every task

One row per task, against whichever baseline was strongest on it. The action-only inference mode of Flex-π beats every baseline; joint future generation mode adds additional gains to every task.

Where that lands on the real robot
Averaged over the five tasks above

Task completion against inference latency. The shaded region is everything both slower and less accurate than the action-only path; hover a point for its numbers.

01Precision and long horizon

The hardest task in the suite: an eight-stage repair whose tightest stage is a sub-millimeter insertion, made with an electric screwdriver held between two moving arms.

Self-repair gripper

The robot repairs its own gripper over eight stages that must be completed in order, split between the two arms, with the tightest insertion near the end.

Autonomous · 2× speed

Driving the screw — the tightest stage

A 4 mm bit must enter a 4.5 mm socket: ±0.25 mm of lateral clearance. 802 demonstrations (11.8 h) + 660 ManiFlow-derived DAGGER corrections

5%
ManiFlow
55%
Flex-π success rate

Where π0.5 and ManiFlow lose the run

Two insertions and the screwing — the three stages they rarely get past. Each is shown twice: the baseline attempting the stage on the left, Flex-π completing it on the right. Stage 7 is the tightest of the eight, a 4 mm bit into a 4.5 mm socket.

STAGE 3Insert gripper — 19 mm part into a 20 mm holder, ±0.5 mm
Autonomous · 2×
π0.5 misses the holder
Autonomous · 1×
Flex-π (joint) seats the gripper
STAGE 5Insert screw — 4.5 mm M5 screw into an 8 mm hole, ±1.75 mm
Autonomous · 2×
ManiFlow cannot align the shank
Autonomous · 1×
Flex-π (joint) starts the screw
STAGE 7Screw in — 4 mm bit into a 4.5 mm socket, ±0.25 mm, the tightest stage
Autonomous · 2×
ManiFlow never seats the bit
Autonomous · 1×
Flex-π (joint) drives the screw home

Recovery behavior

What the policy does after a failed attempt, on the two stages with the least clearance to spare.

Retrying the gripper insertion
Retry
The first approach misses the holder. The arm pulls back, lines the part up again and seats it on the second try.
Retrying the screwing
Retry
The bit glances off the socket instead of dropping in. The arm lifts the driver clear, re-centers it over the screw head and comes back down.
Self-repair gripper
20 rollouts per method, identical training data

Eight stages in order, so we score both partial progress and the full sequence. Flex-π (full joint) finishes all eight in 11 of 20 rollouts; the best baseline manages it once.


02Dexterity on deformable objects

Zipping a soft bag: the fabric shifts under every grasp, so the zip never stays in one place. Five bags the policy was trained on, and one it was never shown.

Rollout
Autonomous · 2× speed

Soft-Bag Zipping

Unzip the bag, hold the opening wide, drop a pen in, then re-grasp the zip and close it — six scored steps on a deformable object, each one reachable only if the last succeeded.

42.8%
π0.5
70%
Flex-π Task Completion
Unzip the bag
Autonomous · 2× speed
π0.5 fails to unzip the bag
Autonomous · 2× speed
Flex-π (joint) unzips the bag
Zip the bag closed
Autonomous · 2× speed
ManiFlow fails to zip the bag closed
Autonomous · 2× speed
Flex-π (joint) zips the bag closed
Soft-Bag Zipping
20 rollouts per method per condition

On the unseen bag π0.5 falls from 42.8% to 17.2% and ManiFlow from 31.9% to 6.9%. Flex-π (full joint) goes from 70.0% to 63.3%.


03Robustness in cluttered scenes

Deployment scenes are rarely clean. We re-run the suite with novel distractor objects filling the workspace, and with objects swapped for types the policy has never handled. Seen is the setting it was fine-tuned on; Unseen is one it never saw.

Evaluation setting
Autonomous · 2.5× speed

Bimanual Put Plate on the Rack

A plate and rack of a type never seen in fine-tuning, with novel distractor objects filling the workspace. 300 teleoperated demonstrations

72.5%
π0.5
95%
Flex-π joint
Robustness in cluttered scenes
Number above each pair = the drop

Flex-π (full joint) holds 95.0% on Put Plate and 70.0% on Sort Utensils under the unseen conditions, 2.5 and 5.0 points below its in-distribution scores. ManiFlow has 3D inputs and drops 32.5 and 22.5 points.

Training on 50% of the data

Put Plate on the Rack

Faded bars are the full training set, solid bars half of it. Flex-π with half the data, action only, performs similarly to π0.5 with all the data. With full joint prediction, Flex-π half data outperforms every baseline with full data.

Part II

How it works

The shared latent space behind all of it, and the two masks that let one set of weights cover every input/output combination.

One latent space, one backbone

Appearance, geometry and object semantics normally mean three separate pipelines. Flex-π carries all three through a single shared backbone.

Three streams, one shared space

RGB and 3D pointmaps go through the same frozen video-generation VAE. DINOv3 semantics come from a separate frozen encoder and are projected into the same backbone by a linear adapter. Proprioception and the language instruction condition every stream. The diagram at the top of the page runs that path live, with the same two masks.

Stream V
RGB — appearance

Latent tokens that carry priors from internet-scale video pre-training.

Enc(ot) · frozen Wan-2.2 VAE

Stream D
DINO — object semantics

Object-level semantic tokens that ground appearance and geometry.

DINO(ot) · frozen DINOv3

Stream P
Pointmap — 3D geometry

A 3D pointmap pushed through the same frozen VAE, so no geometry-specific encoder is needed.

Enc(pt) · shared VAE weights

Why latents?

Flex-π uses a joint flow-matching prediction objective, but predicts future observations as latents from the pre-trained encoders rather than as pixels, all three streams in one space. Each encoder's priors carry over, the joint representation is stronger, and inference is faster. Actions are generated jointly with those latents under shared self-attention, so the policy reads the predicted future without decoding it.

RGB 4× speed
GT pointmap VAE reconstruction
Figure 3A video VAE trained only on RGB already encodes 3D. Drag the divider: left of it is the ground-truth pointmap, right of it the frozen Wan-2.2 VAE’s reconstruction of it — sweep across a moving episode and almost nothing changes (PSNR 31.1 dB, MSE 3.1×10−3 in the normalised space the VAE sees; 4.9 cm z-RMSE in metres).

Cross-modality forcing

Flex-π trains to predict future observations for each modality, even when one stream is missing. This alone raises RoboTwin success by 47% relative.

Robustness to a missing sensor is not the only reason we train this way. Requiring each modality to be predictable from the others keeps the shared backbone from splitting into three weakly-coupled channels, and pushes it toward a representation in which appearance, geometry and semantics are mutually predictive.


The futures behind the actions

Every action chunk is generated jointly with a latent future in each stream. Control never decodes those latents; decoding them afterwards shows what the policy predicted while it was acting.

Real rollout
Predicted future
Self-repairThe gripper-repair sequence.
Real rollout
Predicted future
Put the pen in the bagThe unzip, insert and close sequence.
Real rollout
Predicted future
Put the plate on the rackA flat plate seated in the rack without tipping.
Real rollout
Predicted future
Sort the utensilsSingulating thin objects from clutter.
Real rollout
Predicted future
Organize the kitchen rackSeating a bowl and then a plate.

The decoded future appears under each rollout, in whichever stream you pick.

Part III

How it compares to other VLAs and WAMs in simulation

RoboTwin and LIBERO benchmarks — and what the design costs.

Simulation benchmarks

Two suites that measure how well a policy fits its training tasks. Flex-π is fine-tuned separately for each, and within each one the same checkpoint runs either mode.

Success rate on RoboTwin and LIBERO
Flex-π is one flexible checkpoint run action-only or jointly · Flex-π* is fine-tuned for a single fixed mode

Flex-π leads both categories on RoboTwin, at 94.6% whether or not it generates futures. On LIBERO everything but π0 lands within two and a half points, so it reads as saturation rather than a ranking; the fixed-mode Flex-π* reaches 99.2%, level with the best published result there. Dashes are numbers nobody has published, not zeros — Flex-π* is reported on LIBERO only.


Conclusion and limitations

Flex-π embeds RGB, 3D and DINO semantics into a shared latent space, yielding a single checkpoint that supports any combination of input and output streams at inference. It matches or beats the strongest VLA in fast action-only mode and the strongest WAM at full joint generation across RoboTwin, LIBERO and a real bimanual YAM robot. Compute flexibility at inference comes from multi-stream latent dropout.

Still data-hungryFlex-π gets more out of each demonstration than the baselines we compare against, but the absolute number it still needs is large.
Joint generation trades latency for accuracyFull joint generation costs about 3× the latency of the action-only path (193 vs. 60 ms on an RTX 5090). The action-only path is faster than every baseline we compare against, but the two operating points cannot be had at once.

BibTeX

@article{yan2026flexpi,
  title   = {Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility},
  author  = {Yan, Ge and Liu, Jinghao and Fan, Yuzhi and Cai, Lei and
             Liao, Minwen and Zhang, Jesse and Fox, Dieter},
  journal = {arXiv preprint arXiv:2608.10860},
  year    = {2026}
}