Business only (hides war / politics / culture / sports)
Wednesday, Sep 23, 2026
1 TPUv7 Outperforms Nvidia GB200 by 56% with First Kimi K3 MegakernelGOOGLNVDAPVT:MNSH 🤖 AI Sep 23, 3:10 PM EDT 4/3
1
TPUv7 Outperforms Nvidia GB200 by 56% with First Kimi K3 MegakernelGOOGLNVDAPVT:MNSH
🤖 AI Sep 23, 3:10 PM EDT 4/3
vLLM maintainers and Inferact reported that TPUv7 hardware reached 709 tokens per second during low-concurrency decode for the Kimi K3 model. This performance beats the 450 tokens per second recorded by the Nvidia GB200 baseline for the same workload, representing a 56% increase in throughput. The improvement was driven by the implementation of a first-of-its-kind TPU megakernel optimization.
The results are part of a broader trend toward the externalization of software for TPU hardware, moving these optimizations beyond internal Google environments. vLLM maintainers highlighted the benchmark to demonstrate how software-level changes to kernels can significantly alter the performance gap between competing AI accelerators.