Andrew White 🐦‍⬛ @andrew.diffuse.one · Jun 22

I have written up a 3.5k word/10 figure essay on how to write a reward function while avoiding reward hacking for chemistry. It covers all the ridiculous ways we had to avoid reward hacking for training ether0, our scientific reasoning model. diffuse.one/p/m1-000

23 likes 2 replies

?

Replies

Kjell Jorner · Jun 22

I love this! Is there some mapping of these to the functions in the GitHub repo? Would like to try out the subgraph matching asap.