Will Held @williamheld.com · Jul 3

In our most similar setting to the original work (130M model), we don't see AdamC's benefits but - We use a smaller WD (0.01) identified from sweeps v.s. what is used in the paper (0.05). - We only train to Chnichilla optimal (2B tokens) whereas the original paper was at 200B.

0 likes 1 replies

?

Replies

Will Held · Jul 3

Have an optimizer you want to prove works better than AdamC/Muon/etc? Submit a speedrun to Marin! marin.readthedocs.io/en/latest/tu... For PRs with promising results, we're lucky to be able to help test at scale on compute generously provided by the TPU Research Cloud!