Training a 40M connector gives DeepSeek V4 Flash basic vision

A user froze DeepSeek V4 Flash and a 417M-parameter MoonViT encoder, then trained a 40.1M-parameter connector on 100,000 image-text examples. The experiment shows a text-only MoE can gain basic vision without retraining the language model itself.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs