Rémy @xowap.dev · May 28

The most useful metric I've found is to figure what is the cost to complete given benchmarks. If you take the GPT-5.5 variants at Terminal Bench, it's a third of the cost to use medium instead of xhigh, for essentially the same perf (see: artificialanalysis.ai/evaluations/...)

0 likes 1 replies

?

Replies

Rémy · May 28

To go further in this, what I do is figure which model I like, and then which other models have similar performance and then how much the benchmarks of those did cost.