Ling 3.0 Tiny MoE model released with 7.9B total, 1.3B active params

Ling 3.0 Tiny is a 7.9B-parameter MoE with 1.3B active per token, 256K context, and 32K max output. Community benchmarks report it rivaling Qwen3.5 9B reasoning at ~36 tokens/s on low-end PCs. llama.cpp support landed, and Vercel's AI Gateway offers it free until August 14.
How this story unfolded
13 days · 2 reports · 8 community posts · 10 of 11 shown
- Aug 6
- Aug 11
- Aug 12
- Aug 17
- Aug 18
- Aug 19
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs