Will Held @williamheld.com · May 11

Before risking another run blowing up, I sanity checked that this new recipe seemed reasonably close to the empirical optimal across two hidden dimensions, four batch sizes, and three token counts. Complete(d)P is pretty impressive and generalized well to our setting!

2 likes 1 replies

?

Replies

Will Held · May 11

At this point, we pre-registered the new launch on Github and later on Twitter (sorry Bsky). Stressful to pre-register but important to avoid survivorship bias on scaling law extrapolation findings!! Thankfully, the new recipe scaled predictably for held-out PPL loss.