Choose your first agent workflow

Compare three candidates using repetition, reversibility, and the effort required to review the result.

Choose your first agent workflow editorial photograph

The first workflow is often chosen because someone has a convincing demonstration. A model reads a message, writes a plausible reply, and makes a complicated process look almost finished. The missing work becomes visible later: checking the account, resolving an exception, asking permission, or explaining why the reply should not be sent.

A better starting decision comes from the actual queue. Choose a recurring task that a team understands well enough to describe, review, and stop. The aim of the first pilot is to learn whether a bounded piece of work can be handled usefully. You need a candidate that exposes the important questions while keeping the consequences manageable.

The worksheet below is a discussion aid; its scores do not predict performance. Use it to structure a conversation with the people responsible for the work. The supplier example is illustrative.

Name the unit of work

Write one sentence that begins with a real trigger and ends with an observable result. For example: when a supplier asks for an order update, prepare an internal draft using the current order record. That sentence identifies an input, a source, and an output. It also leaves sending the answer outside the initial scope.

Avoid starting with a department-wide ambition such as improving customer service. That could include dozens of tasks with different risks and owners. Break the ambition into work someone can recognize from yesterday’s queue. If the team cannot agree where a task starts and stops, spend more time mapping it before discussing the technology.

Anthropic’s guidance on effective agents recommends starting with simple approaches and distinguishes fixed workflows from agents that direct their own actions. That is a useful reminder during selection. A good candidate may ultimately need a rule, a form, or a single drafting step rather than a broadly autonomous system.

Bring three candidates to the table

Choose three tasks from the same team’s work so the comparison uses a reasonably shared context. One might be preparing supplier replies, another classifying incoming requests, and a third changing order details. Ask the operator to provide a recent example of each, including one that was awkward. The manager’s description alone may miss the workarounds that keep the process moving.

For each candidate, record frequency, sources, reviewer, completion condition, and likely consequence of a mistake. Use rough observations if precise data is unavailable, but label them as estimates. The point is to reveal which unknowns matter. A detailed spreadsheet built from guesses can make uncertainty harder to see.

Score repetition and input stability

Give a candidate a higher repetition score when the same recognizable pattern occurs regularly. This does not mean every input must be identical. It means the team can explain the common shape and identify important exceptions. A task that happens once a year may still be valuable, but it offers fewer opportunities to learn from a short pilot.

Then consider input stability separately. Does the relevant information arrive in a known location? Can the team identify the current version of the record? A frequent task with unreliable source material may be a poor first choice. Write down the missing input rather than assuming the agent will somehow find it.

A simple three-level scale is enough: low, medium, or high, each with a sentence explaining the judgment. Those sentences are more useful than the total. They let another person challenge a score and show the team what would have to change for a candidate to become more suitable.

Check reversibility before convenience

Ask what happens if the output is wrong. An internal draft can be discarded. An external commitment may create a problem even if a later message corrects it. Changing a system record can affect downstream work before anyone notices. These differences should influence the boundaries of the first experiment.

Reversibility is not the same as low importance. A draft used by a busy reviewer can still cause harm if it contains a convincing error. The pilot needs a review process that actually works under normal conditions. Name the person who will review outputs and confirm that they have time, context, and authority to reject them.

The NIST AI Risk Management Framework provides a broader reference for considering AI risk in context. For this worksheet, translate that into a concrete discussion of consequences and responsibility. Do not turn a low score on a homemade worksheet into a claim that a system is safe.

Measure reviewability as its own property

A task is easier to review when a person can compare the output against accessible evidence. Preparing a summary from a short source document may be reviewable if the original remains visible. Producing a recommendation from many changing records may require much more investigation. The output’s length tells you little about the effort needed to check it.

Ask the prospective reviewer to review a sample now. Observe where they look, which facts they verify, and what they cannot determine. If they must repeat the entire task to judge the result, the pilot may still be useful for consistency or training, but the expected time saving needs to be reconsidered.

Work through the comparison

In the supplier example, preparing a draft reply might score well on repetition and reversibility, with moderate review effort. Classifying requests might be easier to check but less valuable if the current sorting process is already quick. Changing order details might offer a larger apparent benefit while carrying more serious consequences and requiring permissions the pilot does not have.

A sensible first choice could be draft preparation for a narrow set of order-status questions. Exclude disputed invoices and changes to delivery commitments. Keep the original message and order record beside the draft. The team can learn whether the preparation helps before deciding whether any later action should become automatic.

This is a worked decision, not a recommendation that every team choose supplier messages. Another team may have clean internal documents and a difficult external inbox. The same worksheet could point toward preparing internal summaries instead. Preserve the reasons for the choice so the pilot can be evaluated against them.

Write the pilot boundary on one page

The boundary should name the included request types, permitted source records, reviewer, and stopping conditions. State what happens when required information is missing. Include a manual fallback and the person who can pause the experiment. A pilot that cannot be stopped cleanly is harder to learn from because every defect becomes an operational emergency.

Choose a small set of representative cases before the first run. Include ordinary work and several known exceptions. Keep some cases aside for a later check so that improvements are not judged only on examples the builder has already seen. The companion guide to building an agent test set explains how to organize those examples.

Decide what evidence will change your mind

Before running the pilot, write down why the candidate looks promising and what result would undermine that belief. Perhaps review takes longer than manual preparation. Perhaps missing inputs occur much more often than expected. Perhaps employees value a clear status record more than the draft itself. Those are useful findings, even if they point away from the original idea.

End the selection meeting with one owner and one next action: collect the examples, confirm the source access, or observe the current task. Do not treat selecting a candidate as approval to connect every system it touches. A good first workflow gives the team a focused way to learn, with enough structure to recognize when the proposed approach is not helping.

Make Agentso.com yours.

A distinctive .com for your next agent venture.

Inquire about Agentso.com