ArXivIQ

ArXivIQ

Video Generation Models are General-Purpose Vision Learners

Aug 04, 2026
∙ Paid

Authors: Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir and Cristian Sminchisescu
Paper: https://arxiv.org/abs/2607.09024
Project page: https://genception.github.io/
Code: N/A
Model: N/A

TL;DR

WHAT was done? The authors introduce GenCeption, a unified framework that repurposes a pre-trained text-to-video generative diffusion backbone into a single-step, feed-forward generalist vision model. By mapping heterogeneous dense tasks (such as depth, normal estimation, and camera pose) into a unified 3-channel RGB representation space and utilizing learnable queries for sparse tasks (like 3D keypoints), GenCeption performs multi-task video perception steered entirely by text instructions.

WHY it matters? This work demonstrates that large-scale video generative pre-training acts as a powerful “world model” that implicitly internalizes 4D spatiotemporal dynamics, geometry, and physical laws. Rather than relying on slow, iterative denoising, GenCeption extracts these rich representations in a single feed-forward pass, matching or exceeding task-specific state-of-the-art models with up to 500× less training data and showing exceptional zero-shot generalization to out-of-distribution categories.

Details

The Balkanization of Computer Vision and the Quest for a Unified Objective

Computer vision remains heavily balkanized, characterized by task-specific architectures and custom heads engineered for isolated objectives, such as Sapiens for human pose estimation or Depth Anything 3 for geometry. This stands in stark contrast to natural language processing, which collapsed task boundaries by adopting next-token prediction over unified architectures. To achieve a similar catalyst in vision, we need a pre-training objective that forces models to internalize spatiotemporal evolution, physical causality, and native language alignment. While previous visual representation learning paradigms like VideoMAE V2 or V-JEPA [see also] capture spatial correlations, they lack native language concept grounding and the capacity to model complex, continuous temporal physics. The authors address this fundamental bottleneck by demonstrating that large-scale text-to-video generation is the true visual analog to next-token prediction, serving as a superior foundation for general-purpose visual intelligence.

User's avatar

Continue reading this post for free, courtesy of Grigory Sapunov.

Or purchase a paid subscription.
© 2026 Grigory Sapunov · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture