Hugging Face releases quantized MOSS-VL models for local use

FP8 and NF4 versions of MOSS-VL-Instruct and MOSS-VL-Realtime run locally in 24GB VRAM, covering image, video, and real-time streaming understanding. The technical report describes an open vision-language model family co-designed for real-time interaction.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills