Andrew White 🐦‍⬛ @andrew.diffuse.one · Dec 31

Aviary is a gymnasium of new scientific environments. Using behavior cloning, expert iteration, and consensus sampling we’ve trained Llamma-3.1 8B agents to very high accuracy on challenging multi-step tasks. And at low cost! www.futurehouse.org/research-ann...

1 likes 1 replies

?

Replies

Andrew White 🐦‍⬛ · Dec 31

A lot of effort in this work was framing the learning problem of agents. We settled on defining agents using stochastic compute graphs and splitting the environment and agent according to what we want to train. Here are some components of well-known agents as compute graphs: