1 / 2
In our most similar setting to the original work (130M model), we don't see AdamC's benefits but - We use a smaller WD (0.01) identified from sweeps v.s. what is used in the paper (0.05). - We only train to Chnichilla optimal (2B tokens) whereas the original paper was at 200B.
0 likes 1 replies
?