inclusionAI releases Ling-3.0-flash, a 127.5B-parameter model
127.5B total params, 5.1B active, 512 experts with 8 active per token; MIT license with BF16 (~255GB) and official FP8 (~128GB) weights on Hugging Face. Reddit testers report ~80 tok/s decoding on a single DGX Spark.
How this story unfolded
3 weeks · 2 reports · 12 community posts · 14 of 15 shown
- Jul 23
- Jul 25
- Jul 27
- Jul 31
- Aug 4
- Aug 5
- Aug 8
- Aug 9
- Aug 11
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs