The most useful metric I've found is to figure what is the cost to complete given benchmarks. If you take the GPT-5.5 variants at Terminal Bench, it's a third of the cost to use medium instead of xhigh, for essentially the same perf (see: artificialanalysis.ai/evaluations/...)
0 likes 1 replies
?