Over the past week I trained a bunch of picoGPTs (~50M params) on single GPUs trying to get minimal loss with a fixed number of epochs. The biggest architectural impact I found was replacing some of the transformer Value tensors with Value tensors that read from the original input.
10 likes 2 replies
?