Evals
Evaluations for AI features
What it is
Evals are tests for AI features. You collect realistic inputs with known good outcomes, run the feature against them, and score the results automatically: with code, with another model acting as a judge, or with human review. They tell you whether a prompt change, a new model or a retrieval tweak made things better or worse before customers find out.
Enterprise buyers ask how you know the AI is accurate and how you catch regressions. A pass rate on a real set of tasks is a far better answer than a demo.
What it takes
About 4–8 engineer-weeks to build in-house, or 2–4 using Braintrust, LangSmith, Arize.
Answer these first
- Product: What does a good answer look like, written down precisely enough to score?
- Product: Which failures are unacceptable, and which are merely annoying?
- Engineering: Where do eval cases come from, and how do real production failures get added?
- Engineering: If a model grades the answers, how did we check the grader?
- Sales: Would we share eval results with buyers?
- Support: How do bad answers reported by customers end up in the eval set?
This page works best with JavaScript on. Every answer also has its own address, like /what/scim.