Three of eleven steps needed a model.
The method exists to find out which three.
Most of a workflow should not be AI. In one European manufacturer’s order handling, eight of the eleven steps were parsing, lookups, validation and routing that ordinary software does more cheaply, faster, and with an answer you can reproduce (SUPALABS engagement data, 2024–2026). Knowing which steps are which, before anything is built, is the whole difference between a system a team can run and a pilot that quietly dies. This page describes how we find out: five documents, produced in a fixed order, three levels of authority a task has to earn, and a four-rung engagement you can stop after any rung.
- Five named artifacts, produced in the same order on every engagement
- Three authority tiers, set per task type, promoted only by evidence against your thresholds
- Built on top of what you already run. Stop the service and keep the system.
Five artifacts · Three tiers · Four rungs · Europe
The same four workstreams, whoever the client is
Assess maturity, prioritise use cases by ROI, deploy on a governed platform, upskill the workforce. The large consulting firms sell some version of this to every client, because consulting economics reward a methodology that can be resold. A framework that is identical across every buyer cannot produce advantage for any of them: if the company down the road bought the same four workstreams from a different logo, you have both paid for parity.
Source: The State of AI, “Accenture and Deloitte Are Selling the Same Brain to Every Company in Your Category” (21 August 2026), which calls the convergence “a structural outcome rather than a hypocrisy”.
Five documents about your process, and nobody else’s
The Exception Ledger for your order handling is useless to your competitor. So is your Boundary Map. They are produced by sitting beside the people who run the process, and they are the scope the build is priced against. The output is not reusable, which is exactly why it is worth something.
Produced in this order, on every engagement, whatever the sector.
You cannot inspect a client we will not name. You can inspect a document. Each artifact answers the same three questions, because that is how a buyer compares this to a framework: what is it, who in your organisation reads it, and what can they decide once they have it.
| Artifact | What it is | Who reads it | What it lets them decide |
|---|---|---|---|
| 1. Exception Ledger | Every real deviation from the documented process: how often it happens, what triggers it, and who absorbs it today. | The COO or operations director. | Whether the process is worth automating at all, and which of its steps the automation has to survive. |
| 2. AI Boundary Map | Each step classified as deterministic code, model judgement, or human approval. The headline number is the determinism ratio. | The CIO and whoever owns the budget. | How much of the system is actually AI, and therefore what it will cost to run, how fast it will be, and what can be shown to an auditor as a rule rather than argued for as a sample. |
| 3. Evaluation suite | A golden dataset built from your own cases, per-step pass rates, and a monthly accuracy report. Quoted as its own line item, never bundled. | The head of transformation, and the board pack. | Whether the system is still right, six months after the vendor left, and whether a task has earned more authority. |
| 4. Decision log | Every automated action, its inputs, its confidence, and who approved it, in a surface that can be opened without asking an engineer. | Compliance, finance, and an acquirer’s diligence team. | Whether any single decision the system took can be explained and defended. |
| 5. ERP-additive covenant | A written commitment, made at the start, that we build on top of the systems you already run and never propose replacing them. | The CIO, and whoever would carry a migration. | That this engagement is not a replatforming in disguise. |
The order is not decorative. The Boundary Map cannot be drawn until the exceptions are written down, because an exception nobody recorded becomes a step the model is quietly expected to handle. The evaluation suite cannot be built until the Boundary Map says which steps have a judgement to evaluate. The decision log only means something once the suite defines what a correct decision looks like. Skip a step and the next one is a guess.
Shadow mode: measured beside the humans before it is trusted.
Between the Build and the first live task, the system runs alongside the people who do the work without acting. It produces what it would have produced, the team keeps doing the job as before, and the evaluation suite scores the difference. Shadow mode ends when the pass rate you set is met on real cases, not on the demo. Until then, nothing it produces reaches a customer or a system of record.
Three tiers, set per task type, and you set the thresholds.
Every automated task runs at one of three levels of authority. The tier is set per task type rather than per system, so the same workflow can draft one step, wait for approval on another, and act alone on a third. A task moves up only when the evaluation suite shows the pass rate you defined as the threshold, and it moves back down the moment the monthly report shows a regression.
Nothing is promoted on our say-so. The thresholds are yours, the report that tests them is a named line item, and the decision log shows every action that was taken at every tier.
Four rungs. You can stop after any of them.
The Mapping Sprint is how the first two artifacts get made, and its findings are the scope the Build is priced against. Discovery and build are one motion, run by the same people, which is why the quote is a finding rather than an assumption.
Stop the service. Keep the system.
The system runs in your accounts, on top of the software you already own, and it is documented for your team from the first week of the Build, because handover is a deliverable rather than a favour at the end. If you stop the Run stage, the workflow keeps running, the golden dataset stays with you, and the only thing that stops is the monthly report. An engagement is finished when your team runs the system without us, and that is written into the scope, not hoped for.
Where the data lives and who can reach it →The same method, three buyers.
When you should not buy this.
If one of these is your situation, we will say so on the call rather than after the invoice.
Frequently asked questions
Thirty minutes to find out whether there is anything worth mapping.
The call is free, and we will tell you if the answer is no. Bring one workflow that annoys the people who run it.