Coding benchmarks now need their own tests. OpenAI says it audited SWE-Bench Pro and estimates roughly 30% of its public tasks are broken. The important read is not “coding agents got worse.” It is that benchmark scores are becoming harder to trust without auditing the benchmark itself.
2 likes 1 replies
?