The method, in plain terms
Everything we use on an engagement comes from research published since 2023, never from 2010s consulting frameworks. Here are the seven tools, with their sources, so you can check them.
The L0–L5 maturity ladder
Six levels describing what a company knows about its own work. L0, everything lives in people's heads. L1, a few people experiment on their own. L2, the work is written in a form a machine can follow. L3, the data is under control. L4 and L5 describe systems that adapt and learn on their own.
These are dependencies, not a schedule: you do not skip a level. The step from L1 to L2 is where most projects live or die, and it is where we spend the most time.
We work from L0 to L3. L4 and L5 are not sold to an SMB. On engagements. We place the company in the first week of the diagnostic, department by department. [1]
Fixed flow or reasoning agent
Two questions are enough to classify a workflow. Same input, same path every time? Fewer than twenty or so branches in the diagram? If so, it is fixed-flow automation: two to six weeks, a predictable cost, and behaviour you can test end to end.
If the path changes case by case and something has to be decided in an unpredictable context, it is a reasoning agent: two to four months, with an uncertainty no quote removes. Six or seven ideas out of ten fall in the first bucket. That is where the return comes fastest, and where we start.
On engagements. Every automation kept in a diagnostic carries its category in writing. [2]
The seven moves
What an agent does at each step comes down to seven moves: wait for a trigger, check, classify, enrich, produce, execute, ask. Decomposing a workflow means writing down which ones chain together and where a person decides.
The point is practical. Once the workflow is written as moves, you can see which steps are risky, which verify themselves automatically, and where to put the human approval. The diagram below shows a supplier invoice, from the file arriving to the accounting entry.
The green step is where a person decides. On engagements. It is the output format of the analysis phase: one diagram per workflow kept. [3]
The autonomy ladder
Five levels answering one question: who decides. Operator, the person does the work. Collaborator, the agent produces pieces. Consultant, the agent proposes and the person approves every time. Approver, the agent acts on low-risk cases. Observer, the agent acts alone and the person audits afterwards.
Every deployment starts as a consultant. Moving to approver is conditional: thirty days above 95% accuracy, and only on low-risk actions. Irreversible actions never become observer, and the list is written into the contract.
On engagements. The planned level is set per decision point, before anything is built. [4]
One agent, unless verification is automatic
We build one agent per workflow. On tasks nothing verifies automatically, every added agent is one more place where an error slips in silently, then propagates to the next agents, which treat it as valid input.
An analysis published in May 2026 measured a 17x error amplification between multi-agent and single-agent systems on the same task. Multi-agent architectures still make sense when every result verifies itself without a human, like a mathematical proof or a passing test. Your invoices and client emails do not verify like that.
On engagements. It is why we turn down projects that start with a team of agents. [5]
The five production conditions
An agent that works in a demo and an agent still working six months later are two different objects. The second has recovery points, memory separated from working context, a written list of what it may do, explicit failure handling, and an isolated environment with limited access to your tools.
These five are not options. A deployment missing one will hold for the length of the demonstration, then drift without anyone noticing.
On engagements. It is the checklist for the production sprint. We do not ship below five out of five. [6]
Where the saved time goes
An automation produces nothing by itself. It frees hours, and those hours have to go somewhere: task, time saved, usable capacity, reassignment, result. If the chain breaks at reassignment, the hours dissolve into email and the return stays at zero.
Second question, more awkward: who still does this work the old way, in parallel? The shadow process is the most common cause of a measured gain of zero.
Then the J-curve. For two or three months, a deployment that works and one that fails look the same from the outside: extra costs, no visible effect. So we set leading indicators before spending, and stop at week eight if they have not moved.
On engagements. We write the reassignment of hours with you, in the diagnostic deliverable. [7]
Sources
- [1]Turing Post, six-level maturity ladder, April 2026.
- [2]Lenny's Newsletter, April 2026; Anthropic, "Building effective agents", 2024.
- [3]Turing Post, decomposition into moves, May 2026.
- [4]Feng, McDonald, Zhang, Knight First Amendment Institute, 2025.
- [5]AlphaSignal, error amplification in multi-agent systems, May 2026.
- [6]Turing Post, production conditions, May 2026.
- [7]Turing Post, July 2026; Exponential View, July 2026; McKinsey, State of AI 2025 (21% of companies redesigned their workflows).
This is the method we apply on engagements. To place your own company, the three diagnostic questions take two minutes.