No headline generated
GGUF quantizations of the community-pruned Qwen3-VL-32B 'Heretic' (MiniMax H3) run from 6.7GB on local machines, using a text-encoder-pruned weight set. A separate NVFP4 variant is hosted on HuggingFace.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Gemma team to hold special event on August 20
- Situational Awareness invests $400M in chip startup Source Foundry
- Developer creates ChipTycoon to learn chip manufacturing via LLM simulation
- Reddit user creates AI-generated Big Bang Theory sitcom with ComfyUI
- Prompt injection is the most common way scammers attack people and agents