Will Held @williamheld.com · May 11

At this point, we pre-registered the new launch on Github and later on Twitter (sorry Bsky). Stressful to pre-register but important to avoid survivorship bias on scaling law extrapolation findings!! Thankfully, the new recipe scaled predictably for held-out PPL loss.

2 likes 1 replies

?

Replies

Will Held · May 11

Beyond pretraining loss, we find that downstream tasks also scale predictably if you do things carefully! The important trick is that you need to use soft metrics, e.g. NLL or BPB, to predict performance, following Rylan Schaeffer's "Emergence is a Mirage" work.