Avatar IV now runs on TPUs
HeyGen has ported its Avatar IV video generation model, which has more than 18 billion parameters, to Google Cloud's Trillium (v6e) TPUs. The port was done using torchax and XLA, running the model across an eight-chip mesh with FSDP and Ulysses sequence parallelism.
1.86x speedup for real-time streaming
The team achieved a 1.86x speedup for real-time streaming. To get there, they pipelined all-to-all collectives, aligned sparse attention block sizes to eliminate mask padding, and bypassed softmax serial dependencies using a precomputed Cauchy-Schwarz upper bound.
Quality gates
These custom Pallas kernel and compiler optimizations were deployed only after passing rigorous two-tier quality gates. The goal was to guarantee byte-identical or mathematically equivalent pixel outputs.



