User adds basic vision to DeepSeek V4 Flash with 40M connector

A developer froze DeepSeek V4 Flash and a 417M-parameter MoonViT encoder, training a 40.1M-parameter connector on 100K image-text examples to give the text-only MoE basic vision without retraining the LLM.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Enterprise AI agents limited by messy documents
- Seinfeld AI video shows George in GTA 6 using Minimax H3
- Claude Code adds unrequested corrections to spec
- Ethan Mollick: AI impact research must address older-model limits
- Hobbyist trains 1.2B game music generator on single H100