The review desk: why checking costs what it costs

The diagnosis, stated fairly The argument runs like this.

The same work block appears alone on one desk and inside a supporting verification frame on another.

The diagnosis, stated fairly

The argument runs like this. AI use is scaling faster than the capacity of the people supervising it. Sustained cognitive offloading degrades the independent judgement it is supposed to free up. Attention fragments. The organisation ends up with more output and less capacity to tell whether the output is right.

From that McKinsey derives five principles: calibrate cognitive load across a role rather than stacking only high-intensity work; protect recovery as a performance requirement rather than a wellness perk; enable focus by batching interruptions and giving triage to one person instead of everyone; build the skills that do not offload well, such as critical thinking and learning agility; and shape the culture and the framing so people can say what they actually think about the tools. Their summary of the intent is that the human stays at the steering wheel with AI in the loop, not the reverse.

Worth being precise about what that piece is. It carries no survey of its own. It is an editorial frame assembled over other people’s research, which is a legitimate thing to write and a reason to cite it for its argument rather than for its evidence.

The argument holds. Five sensible principles, and an organisation that adopted all five would be better run than most.

What the five principles do and do not reach

Two of them are genuinely work design rather than wellbeing. Load architecture changes what a role is made of. Concentrating triage in one person instead of having everyone check everything changes the total review burden of the organisation, and does so directly. This is not a list that only tells people to rest.

What the five work on is the distribution and the conditions of review: who does it, how often they are interrupted, how recovered they are, how skilled, how safe they feel saying what they found. Those are real levers and they move real numbers.

The variable they do not work through is what sits inside one check. Give the same reviewer the same case in two different packages, and the work of accepting or rejecting it is not the same work. That difference is a design choice, it is separable from how rested the reviewer is, and it does not appear in the five.

So the question we would add is narrower than a disagreement: what does one quality-adjusted check contain, and what does that do to its unit cost?

Self-reported outcomes have opposite signs: 14% more mental effort under supervision and 15% less burnout under substitution.
01 / Self-reported outcomes have opposite signs: 14% more mental effort under supervision and 15% less burnout under substitution.

The evidence that the mode matters more than the tool

There is a survey that makes this concrete, and its most useful finding is one almost nobody quotes.

BCG’s Henderson Institute surveyed 1,488 US employees in January 2026, published in Harvard Business Review as “When Using AI Leads to Brain Fry”. The headline is cognitive fatigue: supervising AI output is associated with 14% more mental effort, 12% more fatigue and 19% more information overload, and 14% of AI users report the acute version.

Now the part that matters for design. In the same study, AI that replaces routine work is associated with 15% lower burnout.

Same technology. Opposite sign. What differs is the mode of engagement: substituting for toil versus requiring intensive oversight. That is not a fact about the resilience of the people involved. It is a fact about how the work was arranged around them, which is the argument of this piece delivered by somebody else’s data.

One more figure, because it changes the urgency. Under this kind of load, the study reports minor errors up 11% and major errors up 39%. Review does not degrade smoothly. It holds on small things and gives way on large ones, which is precisely the wrong order, and it happens before anyone reports feeling unable to cope.

These are self-reported measures from a single survey at one point in time, in the US, and correlations rather than demonstrated causes. Read them as a strong signal about direction rather than as coefficients.

Four hypotheses about what makes a check expensive

These are hypotheses drawn from our own work and from the vocabulary this series uses, not a measured decomposition. We put them forward because they are cheap to test on a real queue, and because each names something a team can change.

They apply to reviewing a decision before it takes effect. Confirming afterwards that an authorised action actually landed is a separate stage with its own cost, and the two get conflated more often than they should.

No reasoning is attached, so it gets reconstructed. The output arrives as a conclusion. Accepting or rejecting it may mean rebuilding the path to it, which is a version of the original work done backwards.

Nothing disagrees. If the system presents its answer without the answers it discarded, there is no contrast to think against, and the reviewer either supplies the counter-position from memory or stops supplying it.

It is not stated whether this case is ordinary. Without a declared class, every item can be the exceptional one, so attention gets spread evenly over cases that do not need it.

Uncertainty is not exposed as a signal. Not how fluent the output reads, but whether the system reports a usable measure of how confident it should be. Without one, nothing can be triaged, and everything is opened to the same depth.

Then, separately, the execution stage: nothing comes back from the system that was supposed to change, so confirming that an approved action took effect becomes a second manual job with its own minutes attached.

None of these is a human failing. Each is a design decision, usually taken implicitly, in the layer around the model.

Small navy blocks remain intact while larger brass blocks split under review load.
02 / Small navy blocks remain intact while larger brass blocks split under review load.

Which makes cost per check a design variable worth testing

Here is what follows, stated as carefully as we can.

If cost per check is effectively fixed, supervisory capacity is mostly a property of people, and the right responses are the ones on that list: spread the load, protect recovery, concentrate triage, train the skill.

If cost per check is partly set by what the system packages into each item, then it is a design variable, and it is one almost nobody is currently measuring. That is a testable proposition rather than a result, and it can cut against us: attaching reasoning and rejected options adds material to read, and could raise time per item rather than lower it. Faster triage on a confidence signal could raise throughput and lower quality at the same time. Neither outcome is obvious in advance, which is the argument for measuring rather than for assuming either way.

So the measurement has to carry more than speed. A week of tally marks on a real queue is enough to start, and five columns cover it: minutes per item; how many systems the reviewer had to open; how often they rebuilt reasoning that existed upstream and was discarded before it reached them; how much of what they approved came back as rework; and how many errors got through and how large they were.

The last column is the one that decides. A design change that cuts minutes and raises escapes has made things worse, and given the 11% against 39% split above, the damage shows up in the expensive tail rather than in the average.

We have not run that comparison at scale and are not claiming a number for it. What we are claiming is that the variable is real, that it is separable from how rested anyone is, and that it is currently being decided by default in most deployments.

Review errors rise by 11% for minor errors and 39% for major errors rather than failing smoothly.
03 / Review errors rise by 11% for minor errors and 39% for major errors rather than failing smoothly.

The objection worth answering

The sharpest response to all of this is not that we are wrong about cost. It is: why check everything at all?

Which is the right question, and it has a real answer. Not every decision earns the same review. A markdown inside policy depth, reversible in a day and cheap to get wrong, does not need what a supplier exit needs. The cost of a check should track the cost of being wrong about that class of case.

That is why the class matters more than any single design fix on the list above. Without a declared class, the only available policies are check everything at the same depth, which is where the fatigue numbers come from, or check by feel, which is where the escapes come from. With one, most items can be sampled and the expensive attention goes where consequence is.

So the sequence is: declare the classes, set a review regime per class, and only then argue about what each individual check should contain. Doing the third without the first two is optimising a queue you have not sorted.

Where recovery still belongs

None of this argues against the McKinsey list, and it would be a poor reading to take it that way.

Recovery, load variation and skill-building matter for a reason this piece does not address: the capacity to judge is maintained by exercising it, and a reviewer who has spent a year approving rather than deciding is a different reviewer. That is a real problem with a real remedy, and the remedy is close to what they propose.

The distinction worth holding is that the two solve different things. Their principles protect the reviewer’s capacity. What a system packages into each item determines how much of that capacity the item consumes. Doing only the first distributes a shortage more evenly. Doing only the second engineers a system nobody is fit to run. Both are needed, and the cheap move is to find out which one your numbers point at before committing to either.

For our part, this is the layer SHEPORD is built for, and the honest way to put it is that it is designed to make these variables visible rather than to promise a number against them. A decision class stated explicitly, so ordinary cases can be marked ordinary. The evidence behind a decision assembled with the options that were rejected. Disagreement shown rather than smoothed away. Operstead, our execution harness, is designed to return a receipt from the system whose state was supposed to change, which separates confirming the decision from confirming the execution instead of leaving both to a person.

Whether that lowers the cost of a check in your operation is a question your own measurement answers, not us.

The bottleneck may well be human. What the human is holding was designed by somebody.