Ahmad Beirami @abeirami.bsky.social · Feb 3

We show that while Bon is almost optimal for win rate vs KL divergence, rewind-and-repeat exhibits a gap. We also show that Bon has different properties than win rate optimal, which we study further in a subsequent work on InfAlign: bsky.app/profile/abei...

0 likes 2 replies

?

Replies

Deepak Ramachandran · Feb 3

This is such amazing work by @abeirami.bsky.social and collaborators. A deep investigation of a simple and practically important idea. Highly relevant to our own work and anywhere else RL is used for Gen AI.

Ahmad Beirami · Feb 3

See the paper: arxiv.org/abs/2401.01879 This is joint work with wonderful colleagues at Google: - Alekh Agarwal - @jonathanberant.bsky.social - Alex D'Amour - @jacobeisenstein.bsky.social - Chirag Nagpal - Ananda Theertha Suresh