LaunchAI ModelsJuly 25, 2026
Microsoft releases Mage-VL, a codec-native streaming multimodal model

The 4B model's visual encoder is trained entirely from scratch, using a codec-native, proactive-streaming design for image and video understanding. It targets the "Moravec's paradox" of VLMs — strong at offline reasoning, slow at streaming perception. Weights are available on Hugging Face.
2 sources
Microsoft by email
Get an email when Microsoft ships something
More stories today
- Zvi Mowshowitz's AI update examines open-weight letter and alignment debate
- ML engineering interview guide covers GenAI, multimodal, agentic AI
- Chinese Hacker Commands DeepSeek via Telegram to Launch Autonomous Attacks
- JetBrains open-sources KotlinLLM, an LLM-powered IntelliJ Kotlin plugin
- AI use complicating relationships, dividing friends and families