The result! • gpt-4o started at 58% accuracy, but I "hillclimbed" by editing the prompt to get to 88% • gpt-4o-mini slightly worse, slightly faster (9 points less accurate 25% faster) • llama3.2 is private, reasonably fast, and works offline—but I only got it to 61% accuracy
1 likes 1 replies
?