LaunchAI ModelsJuly 25, 2026

Microsoft releases Mage-VL, a codec-native streaming multimodal model

The 4B model's visual encoder is trained entirely from scratch, using a codec-native, proactive-streaming design for image and video understanding. It targets the "Moravec's paradox" of VLMs — strong at offline reasoning, slow at streaming perception. Weights are available on Hugging Face.

2 sources

Microsoft by email

Get an email when Microsoft ships something

More stories today

Open the live feed