Start with the decision the evaluation must support

Name the proposed change before collecting a large dataset. Are you replacing a model, increasing tool access, expanding to a different user group or moving from drafted answers to automatic actions? Each change creates different evidence requirements. A general benchmark score cannot tell a business whether its own approval boundary should move.

For an illustrative knowledge assistant, the first question might be whether the pilot answers common policy questions from current documents and declines questions it cannot support. The evaluation must therefore include answerable questions, missing evidence, conflicting versions and requests outside the user’s access.

SWE-bench constructs software tasks from real repository issues and evaluates proposed changes in that context. The relevant point here is task design: a useful evaluation should resemble the work it claims to measure. A benchmark is still not evidence that an unrelated company workflow is ready to release.

Build the set with the people who know the work

Collect a modest starting set from approved real examples. Preserve enough context to explain the expected result, but remove unnecessary personal data. Ask the process owner and frontline users to identify costly mistakes and unusual but legitimate cases. Keep a separate set aside for checking changes so the team does not repeatedly optimise against the same examples.

For each case, record what makes an answer acceptable and what makes it fail. “Helpful” is too broad when two reviewers can reasonably interpret it differently. “Cites the current policy, includes its exception and does not disclose restricted material” gives reviewers something to inspect.

Weight the release discussion by consequence, not only by frequency. A rare unauthorised action should not disappear inside an otherwise strong average. Report the important categories separately, along with how many examples each contains.

Use several kinds of evidence

Deterministic checks are useful for structured fields, forbidden actions and required output shapes. Domain reviewers are useful for correctness that depends on business context. A model-based grader can help scale review, but it also needs calibration against human judgement and inspection of disagreements.

Anthropic’s evaluation guidance distinguishes the task, the trial, the grader, the transcript and the outcome. In particular, a successful-looking conversation is not enough evidence that an agent changed the world correctly. Inspect the final state of the system under test where the task calls for an action.

Run more than one trial for cases where output variability matters. Keep model, prompt, retrieval configuration and tool versions with the result. Otherwise a change in behaviour is difficult to reproduce, and a reviewer cannot tell which version the release readout describes.

RAGAS offers component-level evaluation for retrieval-augmented generation, including questions of retrieved context and answer support. Such measures can help locate a failure. They still need interpretation against the task, especially when a plausible answer is wrong in a way the business considers costly.

Turn results into a decision

A useful readout says which categories improved, which regressed and whether the important boundaries held. Include latency, operating cost and the amount of human review required. Attach representative failures so the product owner can understand what the aggregate numbers mean.

When a result fails, classify the cause before rewriting the prompt. The source may be missing, the retrieval may have selected the wrong passage, the tool contract may be ambiguous or the expected answer may itself be disputed. Each calls for a different correction.

After release, bring approved examples from real failures into the next evaluation. Keep the original held-out checks, maintain the policy behind expected answers and make the test set part of the product’s ongoing ownership.

EvidenceWhat it tells you
Task checksWhether required outputs and boundaries held
Outcome inspectionWhether the intended state actually changed
Human reviewWhether the result fits the business context
Repeated trialsHow stable the behaviour is across runs

Keep the disagreement in the readout

Two reviewers disagree about an answer. Do not immediately average their ratings. Ask whether the source is ambiguous, whether the rubric omitted a condition or whether one reviewer knows something the system cannot see. The disagreement may reveal a product requirement.

Record a short explanation beside the result and update the rule only after its owner agrees. Keep a version history for evaluation criteria, just as you do for the prompt.

Sources & further reading

  1. Anthropic — Demystifying evals for AI agents
  2. Jimenez et al. — SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
  3. Es et al. — RAGAS: Automated Evaluation of Retrieval Augmented Generation

Primary sources inform the technical background. The examples and proposed working methods are MGLO’s own; Business scenarios are identified in the text.