Andrew White 🐦‍⬛ @andrew.diffuse.one · Feb 14

Models do pretty well on the tasks – with GPT-4o getting 72% and Llama-3.1 405B getting 68%. Some models, like Claude Sonnet, would do better but just couldn’t figure out NPT ensembles!

0 likes 1 replies

?

Replies

Andrew White 🐦‍⬛ · Feb 14

We found that custom built tools are better than just a python REPL, that llama-405B is a great open source model for this, and weaker models require carefully worded instructions. We checked for functionally getting a simulation set up and the correctness of the choices