Build an agent test set

Collect ordinary cases, ambiguity, missing information, and failures with outcomes a reviewer can actually judge.

Build an agent test set editorial photograph

A test set is a collection of situations in which you know what acceptable behavior looks like. That sounds straightforward until the team starts writing the expected answers. One person wants a complete response, another wants a cautious escalation, and a third realizes that the underlying policy has never been settled.

Those disagreements are worth finding before launch. A useful test set gives them a place to be resolved and keeps the resolution attached to a concrete example. It also helps the team distinguish an implementation defect from an unclear requirement. The goal is a dependable review process, not a large folder of prompts that nobody can interpret.

Start with a small set that the team can review carefully. The examples concern an internal request-triage agent and are illustrative. They do not establish a benchmark, a required sample size, or a performance guarantee for any product.

Define the behavior you are testing

Start with the job described in the workflow’s handoff brief. An agent that prepares an internal summary should be tested on whether the summary preserves the relevant facts and routes uncertainty correctly. It should not receive credit for sending a message if sending is outside its authority. The evaluation needs to match the actual product boundary.

Write down the observable result. For a triage case, this might include the assigned category, the responsible queue, and whether a clarification is required. Keep style preferences separate from business-critical conditions. A slightly awkward sentence and an incorrect assignment may both deserve attention, but they should not be indistinguishable in the review record.

Anthropic’s evaluation guide for agents distinguishes the task, a particular trial, the grading process, and the resulting outcome. That vocabulary helps avoid a common mistake: treating one successful attempt as proof that the whole case is reliably handled. Preserve the case separately from each run against it.

Collect examples from the work

Ask an operator for recent ordinary requests and several that required extra thought. Remove information that is unnecessary for the test, and follow the organization’s rules for using real records. If you create synthetic examples, label them. Synthetic cases can be useful for isolating a condition, but they do not prove that your set reflects the incoming workload.

For each example, save a source note explaining why it was selected. Perhaps it represents the most common request type, a recurring ambiguity, or a past incident. That note helps later reviewers understand the case’s purpose. Without it, a difficult example may be deleted as unrealistic simply because the next maintainer has not seen the same situation.

Cover four different kinds of situation

Include ordinary success cases first. These establish that the workflow can handle the work it was designed for. An ordinary case should contain enough information to reach a clear result, with an expected outcome a reviewer can explain. Avoid making every example unusually complicated in an attempt to be thorough.

Add ambiguity cases in which more than one interpretation is plausible. A request might mention two projects or use an outdated department name. The expected behavior could be to ask a specific question rather than choose a convenient interpretation. Explain what makes the ambiguity material; otherwise reviewers may disagree about whether clarification was necessary.

Add missing-information cases. The agent may have a clear understanding of the request but lack a required identifier or attachment. Finally, include failures in the surrounding system, such as an unavailable lookup or a denied permission. These categories expose different behaviors and should remain visible when results are summarized.

Write an expectation that allows valid variation

A reference answer can be helpful, but exact wording is often the wrong standard. For a summary, list the facts that must be preserved, the claims that must not be introduced, and the action that must follow. Several different sentences may satisfy those conditions. The grader should focus on meaning where the task permits variation.

For structured outputs, exact values may matter. If the next system expects a specific queue identifier, a plausible synonym is still wrong. State that requirement plainly. A test set can combine semantic judgments with strict checks, provided the team knows which kind of judgment each field requires.

Build one case record completely

Consider a request asking for access to two systems while supplying an employee identifier for only one. The case record includes the original request, the permitted reference data, the workflow version, and the expected behavior. The agent should recognize both access requests and flag the missing information rather than treating the whole request as ready.

The review criteria might ask whether both systems were identified, whether the supplied identifier was associated correctly, and whether the missing detail was made visible to the responsible team. A separate prohibited-action check confirms that no access was granted. The reviewer records a reason for any failure and can mark the case unresolved if the policy itself is unclear.

This complete record is more useful than ten loosely described examples. It gives the team a model for the next case and makes the assumptions visible. Once the format works, collecting additional examples becomes easier because contributors know what information is needed.

Separate development cases from later checks

Builders naturally learn the examples they use while improving a workflow. Keep a separate group of cases for a later check so you can see how the change behaves beyond the familiar set. The appropriate split depends on the available work and the evaluation purpose. Do not present a small collection as statistically representative merely because it was divided into two folders.

When a new failure appears in real use, decide whether it should become a development case, a held-out check, or both through distinct examples. Record that decision. Repeatedly moving difficult cases out of the reported set makes the result look better while reducing its usefulness.

Review the reviewers

Have two people judge a small sample independently. Compare their reasons, not only their labels. If one sees an unsupported promise and the other sees an acceptable inference, the criteria need work. Resolve the policy question before using the disagreement as a model-quality score.

The NIST AI Risk Management Framework offers a broader reference for documented evaluation and management of AI risks. A local test set should remain explicit about its own limits. Passing these cases does not establish that every possible request is handled or that the system meets an external certification standard.

Report the gaps alongside the results

Show which cases ran, which failed, which were skipped, and which could not be judged. Keep failures grouped by the behavior that matters: wrong source, missing clarification, unauthorized action, or incorrect routing. A single average can hide a small number of serious defects inside a large number of easy successes.

End each review with an action. Fix a specific behavior, clarify a policy, collect a missing category of examples, or hold the release. A test set is useful when it changes a decision. Maintain it as the workflow evolves, and keep enough history that the next person can tell whether an apparent improvement came from better behavior or simply from changing what was tested.

Make Agentso.com yours.

A distinctive .com for your next agent venture.

Inquire about Agentso.com