HeyGen ports Avatar IV to Google Cloud TPUs

The 18B+ parameter video model hits 1.86x speedup for real-time streaming on an eight-chip Trillium (v6e) TPU host. Production code runs unmodified via torchax (PyTorch on JAX); optimizations target all-to-all collectives, partial sparse-attention blocks, and the softmax inner loop.
1 source
Google DeepMind by email
Get an email when Google DeepMind has news
No news that day, no email.
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills