We show that while Bon is almost optimal for win rate vs KL divergence, rewind-and-repeat exhibits a gap. We also show that Bon has different properties than win rate optimal, which we study further in a subsequent work on InfAlign: bsky.app/profile/abei...
0 likes 2 replies
?