WeChat AI Team Details WeLM Models Scaling to 617B Parameters

The models, WeLM-HD4-80B and WeLM-HD4-617B, use Hidden Decoding, which expands each token into multiple internal computation streams without growing the Transformer backbone. The 617B model activates 23B parameters (3B for the 80B) and beat autoregressive baselines across nine benchmarks, at 4.4x training cost (5.1x for 80B).
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills