RTX 3090 MiniMax H3 speed test compares FP8 and INT8 ConvRot quantization

Community benchmark compares FP8 scaled vs INT8 ConvRot (W8A8) quantization speed for MiniMax H3. Tested on an RTX 3090 24GB with ComfyUI 0.30.0, CUDA 13.0, SageAttention, 17 Euler steps at 0.3 MP for 2-second clips.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Daydream launches AI companion virtual friend app
- EU AI Act prompts text watermarking; detection API to ship
- OpenAI launches ChatGPT and Codex desktop app for Linux
- Landscape map charts path of self-evolving AI agents
- Zitron: 70% of hyperscaler AI revenue flows from OpenAI, Anthropic