SFT - using a better pretrained model (Qwen 2.5) results in better scores for GSM8K and MATH. Kinda surprising since the final model is still Llama 3.1, wonder if they evaluated Qwen 2.5 on other datasets as well and if the performance was not good enough to be the base model?
0 likes 1 replies
?