Beyond that, the LLM results are very inconsistent. Which prompting methods generate the best results varies significantly across model and setting. And even minor things like "the order of the multiple choice options" greatly changes the results.
4 likes 1 replies
?