An automation pilot can produce a convincing result and still add work to the team. Someone has to check the output, correct an exception, maintain the source connection, and explain the new process to colleagues. If those activities disappear from the measurement, the apparent saving can be much larger than the benefit people actually experience.
Measure the complete unit of work before deciding to expand. That includes the manual baseline and the work that remains after automation. It also includes quality: a faster process is not necessarily a better one if the next person has to repair its output. The purpose of the measurement is to make a decision about this workflow, under the conditions of this pilot.
The calculations below use invented numbers to show the method. They are not performance claims for Agentso.com or any tool. Replace them with observations from your own process and keep the assumptions visible beside the result.
Establish a baseline people recognize
Choose a clearly bounded task, such as preparing an internal response to a routine order-status request. Observe several examples before introducing the pilot. Record active handling time, waiting time, and the reason for any exception separately. Waiting for a supplier to respond is different from a coordinator spending ten minutes finding a record.
Ask the operator what else happens around the task. They may batch similar requests, use a saved template, or check several systems while handling another job. A stopwatch can miss that context. Write a short description of the manual method so that the later comparison is between recognizable processes rather than two unexplained numbers.
Use a representative mix where possible. If the baseline contains difficult cases and the pilot contains only easy ones, the difference says little about the proposed change. If you cannot match the mix, report the limitation and compare the groups separately. A small honest observation is more useful than a precise-looking claim built on different workloads.
Count the work the pilot creates
Break the assisted process into preparation, review, correction, and follow-up. Include time spent opening the result, locating its evidence, and deciding whether it can be used. If a person must repeat the original task to check the answer, record that rather than counting the output as complete when it appears on screen.
Track exceptions separately. Some will be ordinary work that the pilot was designed to escalate. Others will be defects introduced by the new process. Both consume time, but they imply different improvements. A deliberate escalation may be working exactly as intended; a lost attachment may need an implementation fix.
Anthropic’s guidance on effective agents notes that more complex agentic approaches can trade cost and latency for improved task performance. For a business pilot, that makes it sensible to compare the whole workflow rather than assuming that more autonomous behavior automatically creates more value.
Use a transparent time calculation
Suppose the observed manual process takes twelve minutes of active work per ordinary case. In the pilot, preparing the input takes two minutes, reviewing the draft takes four, and routine corrections take another two. The apparent saving is four minutes per ordinary case before accounting for exceptions and maintenance. Keep each component visible so the team can question it.
Now suppose the team handles fifty comparable cases during the observation period. Four minutes each would imply two hundred minutes of gross time saved. If handling pilot-specific exceptions takes sixty minutes and maintaining the workflow takes another ninety, the remaining observed saving is fifty minutes. That is a very different story from multiplying the manual time by every automated output.
These numbers are deliberately simple. A real comparison may need separate groups for request types, operators, or days. Do not hide variation by reporting only an average. A workflow that helps routine cases but creates long delays for a small group may need a narrower scope before it should expand.
Keep elapsed time distinct from effort
A draft may appear almost immediately while the request still waits hours for approval. That can be useful if it makes the review easier, but it does not automatically improve the customer’s experience. Measure the time from request arrival to the agreed completion point separately from the team’s active effort.
Look for a shifted bottleneck. The preparation queue may shrink while the review queue grows. In that situation, adding more automated preparation could increase the backlog. The next useful change might be a clearer review screen, a revised approval policy, or a smaller intake scope. The measurement should help reveal that choice.
Include quality in the decision
Define a few outcomes that matter to the receiving person. For a response-preparation workflow, those could include correct source selection, no unsupported commitments, and enough context for a reviewer to act. Use the same criteria for the manual and assisted processes where the comparison is meaningful.
Keep serious errors visible on their own. A single unauthorized action should not disappear inside a high percentage of acceptable drafts. The NIST AI Risk Management Framework is a broader reference for considering benefits and risks in context. Your local measurement should explain what was checked and what remains unknown rather than claiming universal safety from a successful pilot.
Account for maintenance and adoption
Someone will update instructions, review new failure cases, renew access, and answer questions from colleagues. Record that work during the pilot. Some effort is temporary setup; some will recur. Separate the two so the team can judge whether a larger rollout would spread a fixed cost or multiply an ongoing burden.
Adoption also has a cost. A new interface can interrupt a familiar routine even if its individual steps are faster. Ask operators where they hesitate or leave the tool to finish elsewhere. Their observations can explain why a technically working pilot is rarely used. Do not treat non-use as a training problem until you understand the friction.
Decide what scaling would change
Expansion can alter the workload. Another department may use different source records, receive more ambiguous requests, or have less time for review. List those changes before applying the pilot’s result to a larger population. The measured benefit belongs to the conditions you observed, not automatically to every similar-looking process.
A sensible next step may be another bounded pilot with one changed condition. For example, keep the request type fixed while adding a second team. That makes the new evidence easier to interpret. Changing the model, the interface, the users, and the task at once makes it difficult to understand why the result improved or worsened.
Write a decision note, not a victory slide
Close the pilot with the baseline, the assisted process, the observed quality, the remaining work, and the limitations. State whether the evidence supports expansion, a narrower revision, or stopping. Include the person responsible for the next decision and the conditions that would trigger a review.
The companion guide to choosing a first workflow can help if the pilot points toward a better candidate. A useful result may be learning that a simpler process change solves the problem. The value of the exercise is a better operational decision, even when that decision is to keep the work manual for now.

