Will Held @williamheld.com · May 11

To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error. Getting there took some work 🧵

38 likes 2 replies

?

Replies

Will Held · May 11

Inspired by open scaling suites such as Pythia from @eleutherai.bsky.social, Delphi releases 3 things: - a recipe for choosing what to train at each compute budget - a suite of models trained from that recipe - a scaling law fit on smaller runs to predict larger ones

Sam Harsimony · May 11

Super interesting thank you! Would you mind summarizing any differences between your training recipe and Comple(d)P? Or is this fathfully applying Comple(d)P to large scale?