Most AI features ship without a way to tell whether a change helped. Building that first is the highest-return half day in the project.
Components
- Cases file. Inputs and expected outputs, stored as JSON in the repository.
- Runner. Executes the current system against every case and writes results to a file.
- Scorers. Deterministic checks first — schema validity, required fields, forbidden phrases — then a rubric score for the rest.
- Report. One line per case plus a summary, designed to be readable in a terminal.
Run it on every change
Prompt edits, model version bumps, retrieval tweaks. Recording scores over time turns "it feels better" into a trend line, and it catches regressions that a spot check misses.
Keep a holdout
Reserve part of the case set and never use it during development. Otherwise you tune against your own test set and the scores stop predicting real performance.
Start small
Twenty well-chosen cases beat two hundred scraped ones. Add a case every time production surprises you — that is how the set stays relevant.
Comments (0)
Log in to join the discussion
Log InNo comments yet