AnalysisAI ModelsAugust 9, 2026

MiniMax H3 CLIP swap cuts VRAM from 15.7 GB to 4.5 GB

A user replaced MiniMax H3's Qwen3-VL-32B prompt encoder (truncated to 50 layers, 15.7 GB in NVFP4) with a Qwen3-VL-4B plus a learned linear projection, cutting VRAM to 4.5 GB. The swap outputs the same [seq, 5120] conditioning tensor as the original.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed