Service / Evaluation
Compare results,
not promises.
Model evaluation examines whether an AI system meets the requirements of a defined task. A fluent answer is not necessarily a correct answer, and a public benchmark does not establish suitability for your documents. Evaluate the complete system, including prompts, retrieval, validation and human review.

1. Write task-specific criteria.
Turn the proposed use into checkable requirements. For extraction, compare required fields against an accepted reference. For summarisation, check whether key facts are preserved and unsupported claims are added. For a knowledge assistant, examine both source selection and whether the answer is supported by those sources.
Separate dimensions that have different consequences. Instruction following, factual correctness, appropriate refusal and information disclosure should not disappear into one overall score. Define critical failures that prevent release even if routine outputs look acceptable. Agree the criteria with the process owner, not solely with the people implementing the model.
2. Build a representative evaluation set.
An evaluation set is a collection of inputs with expected outcomes or review instructions. Include ordinary work, difficult cases and cases where the correct response is to decline or ask a question. Preserve important variation in document layout, language, length and completeness. Use de-identified material where possible and respect permissions.
Keep development examples separate from release checks. Repeatedly adjusting a prompt to the same examples can produce a misleading impression of general performance. Record how examples were selected and what remains underrepresented. When staff disagree on the expected answer, resolve the requirement or retain the ambiguity explicitly rather than forcing a false reference.
3. Compare suitable deployment options.
Hosted model APIs, such as those described in OpenAI’s developer documentation and Anthropic’s documentation, expose models through managed interfaces. They move some infrastructure work to a supplier, while introducing contractual, retention and service-dependency questions. Compare the terms for the specific service rather than extrapolating from a consumer chat product.
Open-weight models provide access to model parameters, subject to their licences. They can support internally managed deployment, but require decisions about hardware, inference software, updates and security. “Open-weight” and “open-source” are not interchangeable descriptions. Check the actual licence and deployment obligations before adding an option to the comparison.
4. Combine automated checks with review.
Automated checks work well for exact matches, required fields and prohibited actions. They are less complete for nuanced summaries or answers with several valid formulations. Use a review guide that defines errors plainly and provides examples. Where practical, conceal the system’s identity from reviewers to reduce brand-related expectations.
A model can assist with judging outputs, but its judgement is another model output. Calibrate it against human decisions and examine disagreement. Do not treat model-based grading as an independent guarantee of correctness. Record the input, system configuration and result so failures can be investigated and comparisons can be repeated meaningfully.
5. Measure the complete operating path.
Latency is the elapsed time before a usable result is available. Include retrieval, validation and any approval step, not just the model response. Cost should include infrastructure, model usage, monitoring and staff correction work. A cheaper model may require more review; a larger model may add delay without improving the relevant task.
Document uncertainty in the estimates. Usage volume, input length and supplier terms can change the result. Consult suppliers’ published pricing and your own workload records when preparing a commercial comparison. Avoid turning a limited evaluation into a universal ranking: the conclusion applies to the task, examples and configurations examined.
6. Establish release and change checks.
Set release conditions before reviewing the final results. Record unresolved failures, manual safeguards and tasks excluded from use. Re-run relevant checks after changes to a prompt, model, retrieval index or tool integration. Maintain a route for users to report errors and a procedure for withdrawing a problematic configuration.
An evaluation engagement can include requirements, example selection, comparison procedures and a decision report. Agree whether the purpose is model selection, release readiness or diagnosing an existing system. These questions overlap, but they require different evidence. A useful report explains limitations and next actions as carefully as it explains the preferred option.