MLCommons @mlcommons.org · May 14

MLPerf Training v6.0 has added GPT-OSS 20B. With 21B total parameters (but only 3.6B active per token), this new sparse MoE pretraining benchmark is designed specifically for accessibility—it can run on a single 8-GPU node. https://bit.ly/4noRr14

0 likes 0 replies

?