Explore how Qwen 3.5-397B was optimized on Google Cloud TPUs (Ironwood v7x), delivering approximately 3.1× higher decode performance and 4.7× higher prefill performance through hybrid Attention DP + MoE Expert Parallelism, optimized routing metadata collectives, and JAX/Pallas kernel fusions.