Vivek Kalyan @vivekkalyan.com · Nov 24

DPO ablations - the number of unique prompts in the preference dataset matter. More unique prompts result in better downstream performance, while more generations for the same prompts does not. Using different prompts from SFT data also improve performance.

1 likes 1 replies

?

Replies

Vivek Kalyan · Nov 24

DPO ablations - on-policy data (generations from the base SFT model) improve model performance. Although, it's interesting to note that they don't seem to train the model for multiple rounds like Llama 3.1 (maybe cost concerns?)