DPO - preference data is made of a on-policy model pool (Tulu 3 7B/70B) and off-policy model pool (other models). GPT-4o is used to rate generations from 4 random models from the model pool from 1-5. The highest is taken as the chosen response and the rejected response is sampled from the rest.
1 likes 1 replies
?