AMD releases Instella-MoE-16B-A3B with 2.8B active parameters

The model has 16B total parameters but activates only 2.8B per token, trained from scratch on AMD Instinct MI300X and MI325X GPUs. AMD published weights from every training stage, along with data mixtures and training recipes, marking a fully open release.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Cogent AI releases VR-1 cyber reasoning model
- Orchestrator tool integrates 12 AI coding agents in Visual Studio Code
- Taste Skill rules file for AI coding agents crosses 70,000 GitHub stars
- Hugging Face Diffusers flaws allow arbitrary code execution
- 139 Agent Skills bring legal workflows to Claude, Codex and Gemini CLI