Evaluation is where most AI projects quietly fail. Without a fixed test set and an agreed pass criterion, every improvement claim is an opinion.
Building the reference set
Collect real inputs from production, not invented ones. Fifty to two hundred examples is enough to start if they are representative. Label the correct or acceptable output for each. Store it in version control so it can grow without losing history.
Layered scoring
- Deterministic checks. Does the output parse as valid JSON? Does it contain every required field? Does it stay within a length limit?
- Rubric scoring. A human rates accuracy, completeness and tone on a small integer scale.
- Comparative preference. Two outputs, which is better. Easier to judge consistently than absolute scores.
- Model-as-judge. Cheap and useful for regression tracking, but it must be validated against human labels before being trusted.
Watch for contamination
If the same examples were used to tune prompts, scores drift upward without real improvement. Hold back a portion of the set and never use it during development.
Comments (0)
Log in to join the discussion
Log InNo comments yet