Andrew White 🐦‍⬛ @andrew.diffuse.one · Feb 14

Work by Quintina Campbell, Sam Cox, Jorge Medina, Brittany Watterson. They wanted to see if agents could handle complex multi-step tasks like fetching a protein input structure, running the simulation and analysis. They built several tasks with varying complexity.

1 likes 1 replies

?

Replies

Andrew White 🐦‍⬛ · Feb 14

Models do pretty well on the tasks – with GPT-4o getting 72% and Llama-3.1 405B getting 68%. Some models, like Claude Sonnet, would do better but just couldn’t figure out NPT ensembles!