Eugene Yan @eugeneyan.com · Dec 7

Repeat after me: I will build evals for my tasks. I will build evals for my tasks. I will build evals for my tasks.

64 likes 4 replies

?

Replies

PicoCreator - AI Model Builder 🛫 NeurIPS · Dec 7

Honestly even internal org A/B human eval with a handful of people…. is a step up from what most orgs are doing 😭

@indievlad.bsky.social · Dec 8

I wonder how the classical data science workflow of running test data experiments before got completely forgotten in the world of LLMs.

Mahdi Yusuf · Dec 8

Surprised to be seeing this much lack of evaluation. It’s like trying to verify changes in a less-than-ideal codebase without tests. Random things will start to break.

Eugene Yan · Dec 8

AlignEval makes evals easier by simplifying evals into binary pass/fail labels. Start by labeling 20 samples to get a sense of what good, robust criteria looks like. And after you label 50 rows, it helps you auto-optimize your evaluator! bsky.app/profile/euge...