Ryan Steed @rbsteed.com · Feb 11

Automated benchmarks are not all you need, but they are popular tools in AI development. Hoping this doc is a foundation for future guidelines on field testing and other kinds of evals.

0 likes 1 replies

?

Replies

Ryan Steed · Feb 11

CAISI invites input on any aspect of this draft, including from orgs that conduct AI evals and from users of eval reports (for decision-making, procurement, integration, etc.) Public comment closes March 31 — details here: www.nist.gov/news-events/...