What is an eval, really?

If you come from traditional software testing, the word eval can be confusing. It sounds like a test — and it is — but it behaves differently enough that the old instincts mislead you. This is a short attempt to pin the term down.

A working definition

An eval is a repeatable way to measure whether an AI system’s output is good enough for a specific purpose. That’s it. The interesting part is hidden in two words: measure and good enough.

Why it isn’t just a test

Ordinary tests are deterministic: same input, same output, pass or fail. Model outputs aren’t. Run the same prompt twice and the wording changes. So an eval usually measures a property of the output rather than checking it against one golden string.

Aspect Unit test Eval
Output Deterministic Variable
Pass criteria Exact match Threshold or judgment
Grader Assertion Rule, model, or human
Result Pass / fail Score, distribution, or rate

A tiny example

A simple eval can still be plain code — a grader that checks a property instead of an exact value:

// "Good enough" here means: mentions the refund window and stays polite.
function gradeSupportReply(reply) {
  const mentionsWindow = /\b(30|thirty)[- ]day\b/i.test(reply);
  const isPolite = !/\b(no way|deal with it|not my problem)\b/i.test(reply);
  return { pass: mentionsWindow && isPolite, mentionsWindow, isPolite };
}

Run that over a hundred sample replies and the useful number isn’t any single pass/fail — it’s the rate: 92 out of 100 good enough, and here are the 8 that weren’t.

What’s next

Later posts will get into the harder parts: writing graders that don’t lie to you, using a model as a judge without fooling yourself, and deciding what “good enough” actually means before you start measuring.

← All posts