AMD releases Instella-MoE-16B-A3B, a fully open MoE LLM
Instella-MoE-16B-A3B has 16B total parameters but activates just 2.8B per token, trained from scratch on AMD Instinct MI300X and MI325X GPUs. AMD published weights from every training stage, plus data mixtures and training details, for full openness.
3 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Cloudflare announces Agents Week
- AI tool translates and dubs videos into 100+ languages with voice cloning
- MiniMax delays scheduled release
- AI engineering interview Q&A guide covers Gen AI and LLM engineer roles
- Richard Zemel discusses AI and the future of work in ICML keynote