STACKOPTIMA Editorial · Published · Reviewed

Write a decision rule before collecting wins

A demonstration shows what a model can do under selected conditions. An evaluation asks how consistently it performs the work you intend to give it. The distinction is practical: an attractive answer is easy to celebrate, while an unrepresentative test set can send an expensive architecture in the wrong direction.

Before testing, write the conditions under which you would reject your preferred candidate. Examples include an unacceptable rate of incorrect amounts, unsupported factual claims, invalid structured outputs, or missed deadlines. The thresholds should come from the product's needs. They should not be invented after results arrive to preserve a favored model.

Build a small, representative test collection

Start with examples from the actual task, stripped of unnecessary identifying information. Include common cases, long inputs, ambiguous requests, difficult language, and cases where the answer is absent. A system that invents an answer when the evidence is missing may look helpful in a demo and create costly errors in production.

For an invoice workflow, separate ordinary invoices, credit notes, multiple currencies, scanned documents, and incomplete records. Report performance for each group. A single aggregate score can hide a serious failure in the group that carries the greatest financial risk. Weighting by expected traffic and weighting by potential harm answer different questions; show which you used.

Keep a development set for prompt improvements and a held-out set for the decision. If every difficult example is repeatedly used to tune the prompt, its eventual success says less about unfamiliar cases. Record how examples were selected so another person can identify gaps in coverage.

Match the judge to the question

Use deterministic checks for facts that can be checked deterministically: schema validity, exact identifiers, numeric tolerances, and required fields. Use a clearly written human rubric for judgments such as clarity, completeness, and appropriate escalation. Model-based judging can help scale review, but its decisions should be checked against human assessments.

Google Cloud's evaluation service documents approaches using evaluation datasets, metrics, and model-based assessment. That capability does not remove the need to choose meaningful criteria. A fluent but unsupported response should not earn a high overall grade because the evaluator was asked only whether the writing sounded professional.

Keep the rubric short enough that two reviewers can apply it consistently. When reviewers disagree, record the reason rather than averaging away the ambiguity. Sometimes disagreement reveals that the product team has not yet defined the desired behavior. That is a specification problem worth solving before selecting a model.

Control the comparison

Record the exact model identifier, endpoint, prompt version, tool configuration, output limits, and test date. Use the same task collection for each candidate. If a model requires a different prompt to work properly, document that adaptation and include the effort in your operational assessment.

Repeat enough cases to detect unstable behavior. Report the number of examples and the count of successes, rather than presenting a percentage without its denominator. For an original illustrative result, 19 successes out of 20 examples is 95%, but it is a much weaker evidence base than 950 out of 1,000. Neither automatically predicts a changing production workload.

Measure cost and completion time during the same evaluation. Include retries and tool calls in the task record. Otherwise the quality table and the cost table may describe different operating procedures, making the combined recommendation impossible to reproduce.

Turn failures into the next release gate

Inspect failures before celebrating the aggregate. Distinguish missing evidence, reasoning mistakes, retrieval failures, invalid formats, and infrastructure errors. Each class suggests a different remedy; replacing the model will not fix a document that never reached it.

Publish the decision with its limits: which task passed, under what conditions, and what remains untested. Keep a compact regression set and rerun it when prompts, models, retrieval, or tools change. The strongest evaluation is one that remains useful after the purchasing decision, when an attractive initial result has become a system people depend on.

Sources and further reading

Apply this to your workload →

All insights · Editorial policy and corrections