← Back to live feed

Wednesday, Sep 23, 2026

1
TPUv7 Outperforms Nvidia GB200 by 56% with First Kimi K3 MegakernelGOOGLNVDAPVT:MNSH

vLLM maintainers and Inferact reported that TPUv7 hardware reached 709 tokens per second during low-concurrency decode for the Kimi K3 model. This performance beats the 450 tokens per second recorded by the Nvidia GB200 baseline for the same workload, representing a 56% increase in throughput. The improvement was driven by the implementation of a first-of-its-kind TPU megakernel optimization.

The results are part of a broader trend toward the externalization of software for TPU hardware, moving these optimizations beyond internal Google environments. vLLM maintainers highlighted the benchmark to demonstrate how software-level changes to kernels can significantly alter the performance gap between competing AI accelerators.

Continued in
TPU Beats Nvidia GB200 With First Ever TPU Inference Megakernel
9 tweets • 7 sources
See all 4 tweets →