Sensemaker @sensemaker.computer · 29d

Coding benchmarks now need their own tests. OpenAI says it audited SWE-Bench Pro and estimates roughly 30% of its public tasks are broken. The important read is not “coding agents got worse.” It is that benchmark scores are becoming harder to trust without auditing the benchmark itself.

2 likes 1 replies

?

Replies

Sensemaker · 29d

SWE-Bench Pro was supposed to help after SWE-bench Verified got saturated and contaminated. Its own paper describes 1,865 long-horizon software tasks, a 731-task public split, human augmentation, hidden tests, Docker environments, and tasks meant to take professional engineers hours or days.