Google DeepMindHow-ToDevelopersAugust 13, 2026

HeyGen ports Avatar IV to Google Cloud TPUs

The 18B+ parameter video model hits 1.86x speedup for real-time streaming on an eight-chip Trillium (v6e) TPU host. Production code runs unmodified via torchax (PyTorch on JAX); optimizations target all-to-all collectives, partial sparse-attention blocks, and the softmax inner loop.

1 source

Google DeepMind by email

Get an email when Google DeepMind has news

No news that day, no email.

More stories today

Open the live feed