Vivek Kalyan @vivekkalyan.com · Nov 24

DPO ablations - there is much focus on ensuring model follows instruction, mostly leveraging synthetic data methods to create {chosen, rejected} pairs. They also measure that you can take an existing dataset, regenerate it using their synthetic pipeline and it improves performance. Really cool.

1 likes 1 replies

?

Replies

Vivek Kalyan · Nov 24

DPO ablations - doing PPO is hard, primarily because of the reward model, which is hard to evaluate. They conclude that **for the budget** DPO works well enough for preference tuning.