Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Evaluating AI Output: Metrics That Mean Something

Evaluating AI Output: Metrics That Mean Something

Automated scores are useful for tracking regressions, not for deciding whether something is good. Human review on a fixed set remains the reference.

Evaluation is where most AI projects quietly fail. Without a fixed test set and an agreed pass criterion, every improvement claim is an opinion.

Building the reference set

Collect real inputs from production, not invented ones. Fifty to two hundred examples is enough to start if they are representative. Label the correct or acceptable output for each. Store it in version control so it can grow without losing history.

Layered scoring

  • Deterministic checks. Does the output parse as valid JSON? Does it contain every required field? Does it stay within a length limit?
  • Rubric scoring. A human rates accuracy, completeness and tone on a small integer scale.
  • Comparative preference. Two outputs, which is better. Easier to judge consistently than absolute scores.
  • Model-as-judge. Cheap and useful for regression tracking, but it must be validated against human labels before being trusted.

Watch for contamination

If the same examples were used to tune prompts, scores drift upward without real improvement. Hold back a portion of the set and never use it during development.

Comments (0)

Log in to join the discussion

Log In

No comments yet