NVIDIAAnalysisDevelopersSeptember 14, 2026

NVIDIA Transformer Engine speeds JAX MoE training 10.4x

NVIDIA reports DeepSeek-V3 training on GB200 rose from 103 to 1,068 TFLOPS/GPU using JAX with Transformer Engine kernel optimizations, a 10.4x gain. In the unoptimized baseline, inter-GPU communication consumed 84% of accumulated kernel time.

1 source

More stories today

Open the live feed