Approving what you cannot restate
What the experiment actually did In October 2025, Ivan Aslanov, Patricio Felmer and Ernesto Guerra ran a study with 102 participants, split into three groups.
What the experiment actually did
In October 2025, Ivan Aslanov, Patricio Felmer and Ernesto Guerra ran a study with 102 participants, split into three groups. Everyone was asked how well they could explain a set of questions. One group then asked ChatGPT for an explanation. A second group received the same explanatory text directly, without the AI framing. A third got nothing. Everyone then wrote out their own explanation and rated themselves again.
The group that went through ChatGPT showed the largest drop between what they predicted at the start and what they conceded after having to explain. Their written explanations were also less accurate than those of the group that read the same material without the AI.
The finding is not that the AI made people worse at the subject. It is that the AI made people worse at knowing how well they understood the subject. Those are different failures, and the second one is the dangerous one in an organisation.
It is a preprint, and it is students answering four questions in a lab, not executives working a real decision. Everything that follows is my extension of that mechanism into an organisational setting, not a result the study reports. I think it is worth extending, because of what confidence is used for at work.
Confidence is not a feeling. It is a routing rule
In a company, confidence does not stay inside a person’s head. It routes work.
It decides when analysis stops and a decision starts. It decides whether something goes up a level or gets settled here. It decides whether to ask a second person, commission the deeper review, or sign. Nobody writes “I felt sufficiently confident” in the minutes, but that is the trigger being pulled every time.
So a distortion in confidence is not a private matter. It is a change to the stopping rule of the whole organisation, and it moves in one direction: earlier. Work stops sooner, escalates less, gets a second opinion less often. Everything downstream still looks orderly, because the artefacts are all there.
That is the part worth sitting with. Nothing in the paperwork tells you this has happened. A decision made with well-calibrated confidence and a decision made with inflated confidence produce identical documents.
Agreeing with the output is not a check
Most review processes run on agreement. Someone produces analysis, someone senior reads it, and if it seems sound it is approved. That has always been a weak test, and a fluent explanation weakens it further, because agreement is exactly the thing fluency is good at producing.
The experiment suggests a cheaper test than any of that, though the suggestion is mine and not a finding. The gap between claimed and actual understanding only became visible once people had to produce the explanation themselves. Everyone in the study did that, so the design does not prove that explaining is what corrects the overconfidence. What it does show is that the gap was invisible until then.
So the useful question in a review is not “do you agree with this”. It is “can you state the reasoning without the document open”. If the answer is a summary of the conclusion rather than the reasoning that produces it, you have found the gap, and you found it before the decision rather than after.
This is also a fairer test than it sounds. It does not ask anyone to out-argue the machine or to be an expert on everything. It asks whether the reasoning has been taken in.
The cost, stated honestly
This is slower. That has to be said plainly, because speed is the whole reason the tool is there in the first place, and a control that quietly eats the benefit will be abandoned within a quarter, usually without anyone announcing it.
So the answer is not to apply it everywhere. Most decisions do not deserve it. The ones that do share two properties: consequences that land on someone else, and a cost of reversal that is high. Pricing a promotion for one store is reversible next week. Changing a replenishment parameter across a region, committing to a supplier, or standing behind a number that goes to a board is not.
There is a second cost, less obvious. The test is uncomfortable in a way that lands unevenly. Asking a senior person to explain their reasoning out loud reads as a challenge to their standing unless the practice is universal and routine. If it applies to everyone including the person who introduced it, it becomes a professional norm. If it applies selectively, it becomes a weapon and it dies quickly.
The hole in the governance chart
It would be too neat to say nobody controls for this. Serious organisations already have instruments that get close. Four-eyes review, credit and investment committees, independent challenge, model validation, pre-mortems, decision journals and post-decision review all exist precisely because judgement needs checking, and several of them examine the basis of a decision and not only its output.
The narrower claim is the one I would defend. Most ordinary approval flows, the ones that carry the daily volume rather than the exceptional cases, do not explicitly test whether the approver can reproduce the reasoning. They test whether the approver agrees, and they infer the rest.
That inference used to be reasonable. Producing a structured, confident-sounding case took real work, so the ability to present one was decent evidence of having been through it. What has changed is the cost of producing the case, and not the cost of understanding it. I cannot prove that this is what has shifted approvals, and I would not claim the instruments above are useless. The point is narrower: the cheap inference at the centre of routine approval is worth re-examining, because the thing it was standing on has moved.
Where we are building
SHEPORD is designed as a decision operating system for retail, and this is one of the seams it is designed around. The intent is that a decision carries not only what was decided and by whom, but what it rested on and what was checked before it was authorized, so that authority attaches to a stated basis and not to a signature alone. The point of that design is to let a later review ask what was understood at the time, and not only what was approved.
We’re in design-partner conversations with teams working on exactly this.
Worth looking at in your own approval process: find a recent significant sign-off and ask the approver to state the reasoning without opening the document. Whatever comes back is the actual control you have.