MiniMax H3 CLIP swap cuts VRAM from 15.7 GB to 4.5 GB

A user replaced MiniMax H3's Qwen3-VL-32B prompt encoder (truncated to 50 layers, 15.7 GB in NVFP4) with a Qwen3-VL-4B plus a learned linear projection, cutting VRAM to 4.5 GB. The swap outputs the same [seq, 5120] conditioning tensor as the original.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Demis Hassabis becomes Google DeepMind Chair, Alphabet Chief Scientist
- Student uses $3 chip to run Claude Code for automated betting
- Artist's AI-generated 'Found [You?]' footage project blends video and music
- Anthropic's Haiku 4.5 nears 12 months without an update
- China AI Chip Designer Moore Threads Plans Hong Kong Listing