Google for Developers @developers.google.com · 22d

Scaling frontier Mixture-of-Experts (MoE) models takes more than trial-and-error tuning ⚙️

8 likes 1 replies

?

Replies

Google for Developers · 22d

Explore how Qwen 3.5-397B was optimized on Google Cloud TPUs (Ironwood v7x), delivering approximately 3.1× higher decode performance and 4.7× higher prefill performance through hybrid Attention DP + MoE Expert Parallelism, optimized routing metadata collectives, and JAX/Pallas kernel fusions.