PolicyEngine US @us.policyengine.org · Jun 17

Can a language model compute a household’s taxes and benefits from the prompt alone — no tools? We tested 13 frontier models against PolicyEngine on 100 representative US households. New benchmark: PolicyBench. policybench.org

0 likes 1 replies

?

Replies

PolicyEngine US · Jun 17

GPT-5.5 leads, getting 80.3% of its figures exactly right. The leaderboard lets you slice accuracy by program; income tax is the hardest. policybench.org/paper