Nick Huntington-Klein @nickchk.com · Dec 17

Beyond that, the LLM results are very inconsistent. Which prompting methods generate the best results varies significantly across model and setting. And even minor things like "the order of the multiple choice options" greatly changes the results.

4 likes 1 replies

?

Replies

Nick Huntington-Klein · Dec 17

LLMs are consistently changing, and who knows what comes next, but for now this is not a task the mass-market LLMs are capable of. But if you think we've missed a trick or want to try this again when the next generation comes out, our code and data are very easy to update and are at osf.io/spzbu/