ArXivIQ

ArXivIQ

World Models

Flex-π: A Multi-Stream World-Action Model with Compute Flexibility

Aug 15, 2026
∙ Paid

Authors: Ge Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox
Affiliations: University of Washington, Allen Institute for AI
Paper: https://arxiv.org/abs/2608.10860
Project: https://flex-pi.github.io/
Code: N/A
Model: N/A

TL;DR

WHAT was done? The authors introduce FLEX-π, a 6B-parameter world-action model (WAM) that jointly predicts continuous robot actions alongside future RGB appearances, 3D pointmaps, and DINOv3 [review] semantic representations. Crucially, 3D pointmaps derived from Depth Anything 3 are embedded directly into the latent space of a frozen RGB video variational autoencoder (Wan-2.2) without any pointmap-specific pre-training. Using visual stream dropout with cross-modality forcing within a Mixture-of-Transformers backbone, a single checkpoint supports flexible deployment across the inference spectrum, executing fast action-only inference (∼60 ms) or full joint generation (∼193 ms).

WHY it matters? Standard generalist robot policies face a trade-off between the fast reactive execution of Vision-Language-Action models (VLAs) and the rich internal predictive regularizations of RGB-only World-Action Models (WAMs). FLEX-π resolves this dichotomy by showing that physical grounding in 3D geometry and semantics can be injected into generative policies with zero extra sensor hardware, zero specialized 3D architectural pre-training, and zero mandatory inference slowdown. The model achieves a 2–6× success rate improvement over state-of-the-art baselines on sub-millimeter and deformable bimanual manipulation tasks.

Executive summary: Robotic manipulation requires understanding spatial depth, object semantics, and future visual consequences. Current AI models for robots typically predict only future camera pixels, which capture colors and textures but struggle with exact 3D geometry and object relationships. FLEX-π solves this by predicting 3D shapes and high-level object features simultaneously with robot actions. Remarkably, it requires no expensive 3D sensors or complex custom architectures: it extracts geometry and semantics directly from standard camera images using off-the-shelf foundation models and a standard video encoder. Furthermore, at runtime, the robot can dynamically choose between generating full mental simulations for maximum accuracy or generating actions only for high-speed, real-time control, outperforming existing systems on dexterous real-world tasks.

Details

The Appearance Bias in Generative Visuomotor Policies

Generalist robot policies have recently bifurcated into two main paradigms: Vision-Language-Action models (RT-2, π0​, π0.5​), which cast control as direct conditional imitation, and World-Action Models (UWM, DreamZero, Fast-WAM, LingBot-VA), which jointly optimize for action emission and future observation prediction. While WAMs inherit expressive spatiotemporal priors from internet-scale video generation models, their supervisory signal is almost exclusively confined to RGB latents optimized for pixel-level reconstruction. In contact-rich, high-precision manipulation, pixel appearance is a poor surrogate for explicit 3D Euclidean geometry and object-centric semantics.

Traditional attempts to incorporate 3D spatial priors—such as voxel grids in Perceiver-Actor or explicit point-cloud encoders in 3D Diffusion Policy and ManiFlow—introduce substantial practical barriers. They mandate active physical depth sensors at deployment, require specialized pre-training pipelines on geometric data, and incur fixed computational overheads that degrade inference control frequencies. FLEX-π removes these constraints by framing multimodal world-action modeling around a foundational insight: frozen video-generation autoencoders can serve directly as continuous field encoders for 3D pointmaps, allowing multi-stream geometric and semantic prediction without specialized sensor dependencies or dedicated 3D backbones.

User's avatar

Continue reading this post for free, courtesy of Grigory Sapunov.

Or purchase a paid subscription.
© 2026 Grigory Sapunov · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture