AnalysisAI ModelsSeptember 30, 2026

PixelUMM unifies image and video understanding and generation without encoders

Read original source →arxiv.org

PixelUMM drops both VAE and ViT, feeding 16x16 image patches and 4-frame video tubelets into a single linear layer over a Qwen3-8B Mixture-of-Transformers backbone. It jointly handles autoregressive text prediction and pixel-space flow matching, with performance comparable to open-source baselines.

2 sources

More stories today

Open the live feed