AMD releases Instella-MoE-16B-A3B, an open MoE LLM with 2.8B active parameters

AMD's fully open Mixture-of-Experts LLM has 16B total parameters with only 2.8B active per token. It was trained from scratch on Instinct MI300X and MI325X GPUs, and AMD is publishing weights from every training stage along with data mixtures.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Hailuo AI journey from 0 to scale highlighted by MiniMax
- Formula 1 adopts agentic AI on AWS to accelerate data operations
- H3 full precision weights showcased on Reddit
- OWASP AI Security Verification Standard offers a framework for secure AI apps
- Intology shows AI models training other models