Human-in-the-loop is a perishable control

For COO / risk leaders — your strongest AI control may be decaying without any signal reaching you.

A human control fades as the underlying decision practice recedes.

The checkbox every board accepts

When an AI roadmap reaches the board, one phrase does most of the reassuring: human-in-the-loop. Autonomy will expand, the slide says, but a person approves the consequential calls: the markdown beyond policy depth, the replenishment override. It is a reasonable design. Authority stays with someone who can be held to account.

The reassurance works because everyone pictures the gate the way it looked on day one: an experienced merchant who knows the category, reads the recommendation sceptically, and rejects it when it smells wrong. Nobody pictures the same gate two years and ten thousand approvals later. And almost nobody retests it.

The assumption nobody retests

A human gate is only as good as the judgment of the person standing at it. That judgment is not a fixture; it is a skill. The most direct evidence we have on what is happening to it comes from the people running these organisations.

In June 2026, BCG published a study of 70 C-suite and senior executives on what it calls distributed de-skilling: the collective erosion of thinking skills as generative AI substitutes for them. This is perception data, and worth reading as exactly that. It reports what leaders observe in their own organisations, not a longitudinal measurement. What they observe is uncomfortable enough. Half say de-skilling is already visible, and most expect it to become a material threat within three to five years. Asked which skills erode fastest, they put judgment and decision making at the top of the risk ranking, with problem framing and causal reasoning close behind. The same skills they rate most critical to performance. The same skills the person in your loop is there to supply.

Deloitte’s 2026 research on AI adaptation reaches the same spot from the other side. It separates “adapters” from mere “adopters” by three behaviours, and the first is judgment: knowing when to trust, challenge, or reject an AI output. It also flags the condition that starves that behaviour. Most employees going through significant AI change report rising workload, and lack of time is the top barrier they name. It is not hard to guess where that pressure points a reviewer: toward approving, the one act that takes no time.

A human approval becomes a repeated rubber stamp rather than an independent decision.
01 / A human approval becomes a repeated rubber stamp rather than an independent decision.

Why you would not know

In the leaders’ telling, the way a review gate hollows out is mundane, which is why it is easy to miss. Roughly 90% in BCG’s study cite over-reliance on AI outputs without stress-testing or challenge: the recommendation looks good enough, so it goes through. Many describe a quieter shift alongside it, ownership dissolving until “the AI suggested it” becomes a shield.

None of this shows up on a dashboard. The review metric still reads 100%. Every decision had its human, every box was ticked. What the metric cannot see is whether the human did the work of judging or performed the gesture of approving. A reviewer who has not rejected a recommendation in a quarter looks, on paper, identical to one who catches every error. And barely anyone is looking: in the same study, only one company in ten reports an organisation-wide strategy for de-skilling.

One number in BCG’s material points the other way. In its AI at Work survey of 11,749 people, 52% said that AI had increased the time they spend reviewing and correcting output. Read that precisely: it is the share of respondents reporting more review time, not a measure of how much more, and it is not limited to companies that lead on AI. It is also not proof that their review works. What it does show is that serious human oversight consumes real capacity, which means it cannot be bolted on as a free checkbox once the agents multiply.

A control-decay diagram pairs weakening review practice with rotation, sampling, and independent reconstruction.
02 / A control-decay diagram pairs weakening review practice with rotation, sampling, and independent reconstruction.

Review is work. Design it like work, then retest it.

BCG’s own conclusion about at-risk skills is that they survive only through active use. If that is right, the fix for a hollowing gate is not another approval step. More boxes produce more ticking. The fix is to design the review so it exercises judgment, and to check periodically that it still does. Four practices follow. Their effect on judgment has not been measured, by us or by anyone, and that is precisely the point: they make the control testable instead of assumed.

Together these turn reviewer calibration into something you can put on the same page as model performance, measured and discussed. A control you do not exercise and do not retest is a control you assume.

Where SHEPORD sits in this frame

SHEPORD, Valnative’s retail decision operating system, is designed with this failure mode in view. Each recurring decision type (a markdown, a replenishment override) carries a short, explicit definition: what evidence the reviewer must be shown, including the options the system considered and rejected, and whose name the approval requires. Every approval or refusal is captured with its reasoning in the decision’s own record. The intent is that the quality of review becomes inspectable after the fact rather than presumed. You can see whether your gate judges or rubber-stamps.

One boundary worth naming: SHEPORD’s job is the decision and its authorization. Carrying an authorized decision into other systems is a separate concern, and in the Valnative portfolio a separate product, Operstead, acts only on what SHEPORD has authorized.

What the design gives you is narrow and concrete: the gate is inspectable. Whether a particular review practice keeps judgment sharp is exactly the kind of thing an organisation should measure for itself rather than take from a vendor.

A designed review boundary keeps human judgment active at selected decision points.
03 / A designed review boundary keeps human judgment active at selected decision points.

Three questions before you grant more autonomy

  1. Which of your human gates have a named owner, someone who can be wrong in a way that matters to them?
  2. When did a reviewer last reject an AI recommendation and turn out to be right? Would your systems show you?
  3. What share of review time does your autonomy budget actually pay for — and is it growing or shrinking as the agents multiply?

If the answers are “unclear”, “no idea”, and “shrinking”, the honest conclusion is not that your control has failed — it is that you cannot currently tell. For a control that underwrites expanding autonomy, not being able to tell is the finding.

If that is the gap you recognise, the useful next step is a conversation, not a pilot pitch. We are looking for design-partner discussions on the human side of governed decision loops, and we are glad to walk through the bounded runtime evidence behind what is claimed here.