Is Your AI System Still Right Six Months Later? Nobody Knows Unless Someone Measured
An automated workflow goes live. For the first month, everyone watches it. By the third, the people who used to do the work have moved on to other things, the vendor has moved on to other clients, and the system is producing outputs that nobody checks because it has not obviously broken. By the sixth month, the customer price list has changed, a supplier has started sending a new invoice format, and the model that was accurate on the cases it was built against is quietly wrong on a fifth of the cases it now sees. Nothing has failed. Nothing has alerted. The system is still running, and it is no longer right.
This article explains, for a COO or an operations director rather than an engineer, what an evaluation suite is, why it is the only honest answer to "is it still right," and what to ask for so that the answer exists six months from now. It is written from the position that most of a workflow should not be AI at all, and that the part which is needs a number attached to it, forever.
Key Takeaways
- An evaluation suite is a golden dataset plus a monthly report. Cases your own team handled correctly, scored step by step, re-run against the live system every month.
- Trust is the wrong instrument. The people building these systems distrust them more than they trust them: 46% of developers actively distrust AI tool accuracy against 33% who trust it (Stack Overflow, 2025). A COO asked to trust one in production is entitled to a number.
- Score the steps that need a model, not the whole system. Most steps are lookups and rules. In one European manufacturer's order handling, three of eleven steps needed a model (SUPALABS engagement data, 2024–2026). Those three get the golden dataset.
- Buy it as a line item. A bundled evaluation is the first thing cut when the build runs late. Quoted separately, it survives.
What "Still Right" Means, in Operational Terms
A workflow is right when it produces the output the person who used to do the work would have produced, on the cases that actually arrive. Both halves matter. "The output the person would have produced" means the standard is your own team's judgement, not a benchmark the vendor chose. "On the cases that actually arrive" means the standard has to include the exceptions, the customer who sends orders as a photo, the invoice that references a purchase order that does not exist, because those are where a system drifts first.
Drift has three ordinary causes, none of them dramatic. The world changes: a price list, a supplier's document format, a regulation, the mix of customers. The system changes: a model version is updated by its provider, a connector is patched, a rule is edited to fix one case and silently breaks another. And the people change: the operator who used to catch the wrong outputs on sight has been promoted, and the new one does not know what wrong looks like. A system exposed to all three for six months without measurement is not "probably fine." It is unmeasured.
What an Evaluation Suite Is
An evaluation suite has two parts, and both are ordinary enough that a COO can specify them without an engineer.
The first part is a golden dataset: a set of real cases from your own history, orders, invoices, claims, drafts, each paired with the output your team produced and signed off. Not synthetic examples. Not the vendor's demo data. Your cases, including the awkward ones, chosen so that the set covers the exceptions in the Exception Ledger and not just the happy path. A few hundred cases is usually enough for one workflow; the number matters less than whether the exceptions are represented in proportion to how often they actually occur.
The second part is a per-step pass rate. The workflow is scored not as a whole but step by step, and only on the steps that involve a judgement. A lookup does not need scoring; it is either implemented correctly or it is not. A step where a model reads a supplier's PDF and decides which catalogue code it refers to does. In one European manufacturer's order-handling workflow that SUPALABS mapped, three of eleven steps genuinely needed a model, and the other eight were parsing, lookups, validation and routing (SUPALABS engagement data, 2024–2026). The golden dataset scores those three. The other eight are tested once, like any software, and left alone.
Put the two together and run them every month against the live system, and you have a monthly accuracy report: for each judgement step, the share of golden cases the system got right this month, compared with last month and with the threshold you set. That report is the answer to "is it still right." It is one page. It is written by the system about itself, which makes it the only testimonial worth having.
What a Monthly Accuracy Report Contains
| Line | What it tells a COO |
| Pass rate per judgement step, this month vs last | Whether any step is drifting, and which |
| Pass rate vs the threshold you set | Whether the step keeps its current authority tier or drops back |
| Golden cases added this month | Whether the dataset still reflects the cases that actually arrive |
| Regression alerts raised, and what changed | Which of the three drift causes was responsible |
| Cost per run, this month vs last | Whether a cheaper model could do a step at the same pass rate |
Why Trust Is the Wrong Instrument
The alternative to measurement is trust, and trust in AI output is going in the wrong direction as usage goes up. The 2025 Stack Overflow Developer Survey found 84% of developers using or planning to use AI tools, and in the same survey 46% actively distrusting the accuracy of those tools against 33% trusting it. Those are the people who build these systems for a living, with the most exposure to how they fail. If they will not run on trust, an operations director responsible for the output of a supplier-invoice queue should not be asked to either.
This is also why a vendor's assurance at launch is not evidence. A launch accuracy figure is a measurement taken on the day the world, the system and the people were all as the vendor left them. The monthly report is the same measurement taken after all three have moved. The first tells you the system worked. The second tells you it works.
How the Report Governs Authority
The evaluation suite is not only a monitor. It is the mechanism by which an automated task earns the right to act. Every task starts at a Drafted tier: the system prepares the work, a person finishes it, and the golden dataset scores the drafts. When the pass rate you defined as the threshold is met on real cases, the task moves to Approved: the system prepares the complete action and a named person authorises it. Only for reversible steps, and only after the pass rate holds, does a task move to Autonomous, where it acts and the team audits afterwards. A regression in the monthly report sends the task back down a tier until the pass rate recovers. Irreversible steps, anything that files with an authority, ships goods or moves money, stay at Approved regardless of accuracy, because a reversible mistake is a cost and an irreversible one is a liability.
The thresholds are yours. Nothing is promoted on the vendor's say-so. That is the difference between "human in the loop" as a slogan and as a number, and it is set out in full on the method page.
What to Ask For, Before the Build
Four things, all of them checkable on a quote or in a scope document.
- The evaluation suite as its own line item. Not bundled into the build. A bundled evaluation is the first thing cut when the build runs late, and the build always runs late somewhere. Quoted separately, it survives the schedule.
- A golden dataset built from your cases, covering the exceptions. Ask how the cases will be chosen and whether the Exception Ledger drives the selection. A dataset of clean cases will report a high pass rate on the day it stops being true.
- Per-step scoring on the judgement steps only. Ask which steps will be scored and why. The answer should be a short list, and it should match the AI Boundary Map. A partner who proposes to score "the system" has not classified the steps.
- The monthly report, with a named recipient, that continues after the partner leaves. Ask who receives it, what happens when a line crosses the threshold, and whether the golden dataset stays with you if the service stops. The right answer to the last one is yes: the system runs in your accounts, the dataset is yours, and the only thing that stops is the report.
The Six-Month Test
Here is the whole article as one question to put to any vendor, or to your own team about a system already live. Six months from now, on a Tuesday, the supplier-invoice queue produces a wrong match that nobody notices, because it looks like all the others. What, in the current setup, would have told you that morning? If the answer is a person who happens to be paying attention, you are running on trust. If the answer is a line on a monthly report that crossed a threshold you set, you are running on measurement. Only one of those survives the person being promoted. The evaluation suite is not the expensive part of an AI workflow. It is the part that makes the rest of it worth having, and we cover what else a partner should hand over in the five documents to ask for.
A Number, Not an Assurance
Every SUPALABS build ships with an evaluation suite quoted as its own line item, a golden dataset built from your cases, and a monthly accuracy report that continues after we leave.
See how an engagement runs →Sources & References
- Stack Overflow, "2025 Developer Survey: AI", source of the 84% adoption figure and the 46% distrust versus 33% trust figures on AI output accuracy.
- SUPALABS engagement data, 2024–2026: three of eleven steps needing a model in a European manufacturer's order handling. Anonymised by engagement; no client is named. Published with sources at /en/work/.
إحصائيات رئيسية (2025)
قراءة إضافية
الأسئلة الشائعة
Innovation9 min2026-09-09

