B12Y ConsultingB12Y Consulting
  • Home
  • Services
  • Articles
  • Domains
  • For Agents
  • About
  • Contact
Contact B12Y
Articles / Evaluate

Reliability is a system property

Evaluation gives a team the evidence needed to measure, monitor, and improve an AI system in production.

B12Y Consulting / 8 March 2026 / 4 min read

In this article

  • 01Define acceptance criteria
  • 02Build the evaluation set from real work
  • 03Scale review carefully with model judges
  • 04Evaluate before and after deployment
<-All articles

A successful demonstration tests only a limited set of cases. Production inputs include different distributions, malformed data, ambiguous requests, and conditions that the demonstration may not cover. Quality therefore depends on task-specific criteria, representative evaluation cases, and regression tests that detect behavioural changes. These are properties of the complete system, not only of the model.

Define acceptance criteria

Evaluation starts with criteria that reflect the real task, not a generic notion of model quality. For a support assistant, that might be factual accuracy against the knowledge base, correct escalation of the cases it should not answer, and tone within policy. For an extraction pipeline, it is field-level precision and recall, with the fields weighted by what an error costs downstream. A wrong currency code in an invoice matters more than a mangled description, and the evaluation should know that even though a generic metric never will.

Define unacceptable outcomes as well. Every system has failures with different costs: a clumsy sentence versus a fabricated policy commitment, or a missed match versus a false one that triggers a payment. Enumerating the costly failure modes first lets a small evaluation budget concentrate where it matters.

Build the evaluation set from real work

An evaluation set should include cases drawn from genuine traffic and labelled by people who understand the workflow. The required sample size depends on input diversity, failure prevalence, risk, and the confidence needed for a release decision. Real cases represent observed ambiguity, malformed documents, and operational edge cases. Synthetic cases complement them by covering rare but costly scenarios, such as prompt injection attempts or requests near a policy boundary.

An evaluation set should be maintained with the same care as the test suite of a conventional codebase. Confirmed production failures can become regression cases, and deliberate changes in scope should add coverage. Over time, the set provides an executable definition of the system's required behaviour.

Scale review carefully with model judges

Human review may not scale to every output of a high-volume system. LLM-as-a-judge methods can grade outputs against a rubric, but the judge is also an AI system with failure modes. It can prefer confident phrasing over correct content, change with model updates, or reproduce biases shared with the system under evaluation. Scores from an uncalibrated judge should not be treated as reliable measurements.

Calibration requires humans and the judge to score an overlapping sample, followed by analysis of where their decisions differ. The review should measure agreement between human reviewers as well as agreement between humans and the judge. It should also test performance by failure type and check for position, verbosity, style, and self-preference bias. Rubrics need concrete anchors and worked examples rather than a bare instruction to rate quality from one to ten. A continuing sample of human review is necessary to detect changes in judge behaviour and task distribution.

Evaluate before and after deployment

An AI system changes for many reasons: prompts are edited, retrieval corpora grow, tools are added, and providers update models. Each change can affect behaviour in unrelated areas, so regression evaluation belongs in the deployment pipeline. A regression suite should detect when a prompt edit fixes one failure but introduces others. Pinned model versions and a tested upgrade path reduce unplanned behavioural changes.

Monitoring compares release evaluation with live behaviour. Relevant signals include escalation and correction rates, sampled human-review scores, and distribution shifts in inputs and outputs. Confirmed production failures can become regression cases, while a separate holdout set helps prevent repeated tuning against the same examples.

The aim is not a single score. It is a measurement practice that lets a team answer three questions with evidence: is the system ready, what changed, and where is more work needed? That evidence supports explicit release decisions and directs engineering effort towards observed failures.

Questions worth asking

  • What does a good outcome look like for the actual user or workflow, and which failures would be costly, unsafe, or hard to detect?
  • Is the evaluation set built from real traffic, and does every production failure feed back into it?
  • How are prompt, model, retrieval, and tool changes tested before release and monitored after it?

Services

  • Agentic AI Design
  • AI-First Transformation
  • AI Rationalisation
  • AI Quality & Evaluation

Company

  • About
  • Domains
  • Articles
  • For Autonomous Agents
  • Contact

Legal

  • Privacy Policy
B12Y ConsultingB12Y Consulting

Copyright B12Y Limited 2026

Auckland, New Zealand · Services worldwide