The method

Three of eleven steps needed a model.
The method exists to find out which three.

Most of a workflow should not be AI. In one European manufacturer’s order handling, eight of the eleven steps were parsing, lookups, validation and routing that ordinary software does more cheaply, faster, and with an answer you can reproduce (SUPALABS engagement data, 2024–2026). Knowing which steps are which, before anything is built, is the whole difference between a system a team can run and a pilot that quietly dies. This page describes how we find out: five documents, produced in a fixed order, three levels of authority a task has to earn, and a four-rung engagement you can stop after any rung.

  • Five named artifacts, produced in the same order on every engagement
  • Three authority tiers, set per task type, promoted only by evidence against your thresholds
  • Built on top of what you already run. Stop the service and keep the system.

Five artifacts · Three tiers · Four rungs · Europe

What a framework produces

The same four workstreams, whoever the client is

Assess maturity, prioritise use cases by ROI, deploy on a governed platform, upskill the workforce. The large consulting firms sell some version of this to every client, because consulting economics reward a methodology that can be resold. A framework that is identical across every buyer cannot produce advantage for any of them: if the company down the road bought the same four workstreams from a different logo, you have both paid for parity.

Source: The State of AI, “Accenture and Deloitte Are Selling the Same Brain to Every Company in Your Category” (21 August 2026), which calls the convergence “a structural outcome rather than a hypocrisy”.

What the method produces, you are here

Five documents about your process, and nobody else’s

The Exception Ledger for your order handling is useless to your competitor. So is your Boundary Map. They are produced by sitting beside the people who run the process, and they are the scope the build is priced against. The output is not reusable, which is exactly why it is worth something.

The five artifacts

Produced in this order, on every engagement, whatever the sector.

You cannot inspect a client we will not name. You can inspect a document. Each artifact answers the same three questions, because that is how a buyer compares this to a framework: what is it, who in your organisation reads it, and what can they decide once they have it.

ArtifactWhat it isWho reads itWhat it lets them decide
1. Exception LedgerEvery real deviation from the documented process: how often it happens, what triggers it, and who absorbs it today.The COO or operations director.Whether the process is worth automating at all, and which of its steps the automation has to survive.
2. AI Boundary MapEach step classified as deterministic code, model judgement, or human approval. The headline number is the determinism ratio.The CIO and whoever owns the budget.How much of the system is actually AI, and therefore what it will cost to run, how fast it will be, and what can be shown to an auditor as a rule rather than argued for as a sample.
3. Evaluation suiteA golden dataset built from your own cases, per-step pass rates, and a monthly accuracy report. Quoted as its own line item, never bundled.The head of transformation, and the board pack.Whether the system is still right, six months after the vendor left, and whether a task has earned more authority.
4. Decision logEvery automated action, its inputs, its confidence, and who approved it, in a surface that can be opened without asking an engineer.Compliance, finance, and an acquirer’s diligence team.Whether any single decision the system took can be explained and defended.
5. ERP-additive covenantA written commitment, made at the start, that we build on top of the systems you already run and never propose replacing them.The CIO, and whoever would carry a migration.That this engagement is not a replatforming in disguise.

The order is not decorative. The Boundary Map cannot be drawn until the exceptions are written down, because an exception nobody recorded becomes a step the model is quietly expected to handle. The evaluation suite cannot be built until the Boundary Map says which steps have a judgement to evaluate. The decision log only means something once the suite defines what a correct decision looks like. Skip a step and the next one is a guess.

Before it acts

Shadow mode: measured beside the humans before it is trusted.

Between the Build and the first live task, the system runs alongside the people who do the work without acting. It produces what it would have produced, the team keeps doing the job as before, and the evaluation suite scores the difference. Shadow mode ends when the pass rate you set is met on real cases, not on the demo. Until then, nothing it produces reaches a customer or a system of record.

How authority is earned

Three tiers, set per task type, and you set the thresholds.

Every automated task runs at one of three levels of authority. The tier is set per task type rather than per system, so the same workflow can draft one step, wait for approval on another, and act alone on a third. A task moves up only when the evaluation suite shows the pass rate you defined as the threshold, and it moves back down the moment the monthly report shows a regression.

1
DraftedThe system prepares the work and a person finishes it. Every draft carries its evidence: the source documents, the rule or model that produced it, and its confidence. This is where every task starts, and where shadow mode runs before anything is shown to the team at all.Gated by the Exception Ledger — the task is not automated at all until its real deviations are written down.
2
ApprovedThe system prepares the complete action and a named person authorises it before anything consequential moves. The approval, the approver and the inputs go into the decision log.Gated by the evaluation suite — promotion from Drafted needs the per-step pass rate you set, measured against the golden dataset.
3
AutonomousThe system acts and the team audits afterwards, on a sample or on the exceptions. Any irreversible step stays at Approved regardless of accuracy, because a reversible mistake is a cost and an irreversible one is a liability.Gated by the monthly accuracy report — a regression alert sends the task back down a tier until the pass rate recovers.

Nothing is promoted on our say-so. The thresholds are yours, the report that tests them is a named line item, and the decision log shows every action that was taken at every tier.

How you get it

Four rungs. You can stop after any of them.

The Mapping Sprint is how the first two artifacts get made, and its findings are the scope the Build is priced against. Discovery and build are one motion, run by the same people, which is why the quote is a finding rather than an assumption.

0
Qualification call — free, 30 minutesA straight answer on whether there is anything here worth mapping. No deck, no proposal, no follow-up sequence.
1
Mapping Sprint — five days, with the people who do the workPaid discovery. Produces the Exception Ledger, the AI Boundary Map, the evaluation plan and either a fixed-price build quote or a written no. If it shows you nothing you did not already know, you do not pay.
2
Build — six weeks to productionScoped from what the sprint found. Majority deterministic software on top of your existing systems, every task starting at Drafted, shadow mode before anything acts, the decision log live from day one.
3
Run — ongoing, and optionalMonthly accuracy report against the golden dataset, regression alerts, cost-per-run tracking, and the tier promotions your thresholds allow. Stop it whenever you like and keep the system.
When it ends

Stop the service. Keep the system.

The system runs in your accounts, on top of the software you already own, and it is documented for your team from the first week of the Build, because handover is a deliverable rather than a favour at the end. If you stop the Run stage, the workflow keeps running, the golden dataset stays with you, and the only thing that stops is the monthly report. An engagement is finished when your team runs the system without us, and that is written into the scope, not hoped for.

Where the data lives and who can reach it →
Who buys it

The same method, three buyers.

Enterprises and corporate groupsOne operational workflow that hurts, existing systems you intend to keep, and no internal AI delivery team yet. Embedded Operators →
Private equity sponsors and operating partnersOne Mapping Sprint per portfolio company, a written yes or no, and an accuracy report that goes in the board pack. For private equity →
Italian mid-market and SMEsLo stesso metodo, in italiano, costruito sopra il gestionale che usi già. Sprint di Mappatura Operativa →
Read this before you book

When you should not buy this.

You can staff a standing practiceIf you are large enough to run your own internal delivery team, hire it. Embedded delivery is how you learn what to hire for, not a permanent substitute for it.
The problem is arithmetic, not judgementIf the rules are knowable and stable, a rules engine or a well-built spreadsheet beats anything we would put in front of you, and costs less to maintain. We will tell you on the call.
No executive owns the outcomeWithout a named owner who can clear an approval boundary, the work stalls at the first one. That is true regardless of who builds it.

If one of these is your situation, we will say so on the call rather than after the invoice.

FAQ

Frequently asked questions

Is this a framework?It is a method, and the difference matters. A framework is the same document for every client, which is why it can be sold at scale and why it cannot produce advantage for any one buyer. The five artifacts here are produced from one company’s actual process, by sitting beside the people who run it. The Exception Ledger for your order handling is useless to anyone else, and that is the point.
Why are the artifacts the proof, rather than client names?An embedded operator sits inside a client’s operation, and confidentiality is a condition of that access, so no client is ever named. What you can inspect instead is what every engagement produces: the same five documents, in the same order, regardless of sector. On a call we can walk through anonymised examples of each.
What are the authority tiers, and who sets them?Every automated task runs at one of three levels. Drafted means the system prepares the work and a person finishes it. Approved means the system prepares the complete action and a named person authorises it before anything moves. Autonomous means the system acts and the team audits afterwards. The tier is set per task type, not per system, and it only moves up when the evaluation suite shows the pass rate that you defined as the threshold. You set the thresholds. Nothing is promoted on our say-so.
What happens if we stop working with you?You keep the system. It runs in your accounts, on top of the systems you already own, and it is documented for your team from the start, because handover is a deliverable of the Build, not a favour at the end. The monthly accuracy report is the only thing that stops when the Run stage stops, and the golden dataset it runs against stays with you.
How much of the system will actually be AI?Usually a minority of it. In one European manufacturer’s order-handling workflow that SUPALABS mapped step by step, three of eleven steps genuinely needed a model; the other eight were parsing, lookups, validation and routing that ordinary software does more cheaply, faster, and with an answer you can reproduce. That is one engagement rather than a statistic, but the shape repeats, and the AI Boundary Map is how we show you yours before anything is built.
How do I get the five artifacts?Through the four-rung ladder, and you can stop after any rung. The thirty-minute qualification call is free and tells you whether there is anything worth mapping. The five-day Mapping Sprint is paid discovery and produces the Exception Ledger, the AI Boundary Map, the evaluation plan and a fixed-price build quote, or a written no. The Build puts one workflow into production in six weeks with the decision log and the evaluation suite in place. The Run stage keeps the monthly accuracy report going. Everything after the call is quoted per engagement, because the scope is a finding rather than an assumption.

Thirty minutes to find out whether there is anything worth mapping.

The call is free, and we will tell you if the answer is no. Bring one workflow that annoys the people who run it.