40M connector gives text-only DeepSeek V4 Flash basic vision

Experiment adds basic vision to DeepSeek V4 Flash by training a 40.1M-parameter connector on 100K image-text pairs. Both the 417M MoonViT encoder and the MoE stayed frozen; only the connector was trained.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Li Auto expects second-generation Livis AI glasses in H2 2027
- Claude fable ultracode beats gpt luna max in Grok Imagine clone test
- Tool logs requests from coding agents like Claude Code and Codex
- Hierarchical AI agent framework enables multi-tool coordination
- Pipeline builds 3D dynamic scene graphs on the fly with SAM, BotSort, and VLMs