Last fall, I naively thought that I had a decent scaling recipe built on top of muP and a few other works on hyperparameter transfer! But when we scaled to larger (8B+) scales with compute support from the Google TPU Research Cloud, things quickly went wrong!
3 likes 1 replies
?