Copilot Is Not a Stage of Maturity. It's a Wiring Problem.

For COO / CDO leaders in retail and CPG: (1) the copilot plateau BCG describes (67% still approving every AI step, only 9% in autopilot, from a 39-executive survey) is not a sign of caution paying off; it is a sign autonomy was assigned to the wrong noun.

Different decision classes receive different levels of autonomy rather than one universal mode.

The plateau everyone recognizes

Walk into most retail or CPG operations running AI at any scale and you will find the same picture. Demand forecasts, replenishment suggestions, markdown recommendations: the models are everywhere, and a person signs off on every one of them. BCG and the Consumer Goods Forum put a number on this in their June 2026 board brief, drawn from a survey of 39 retail and CPG executives conducted that April. Sixty-seven percent of companies operate this way, AI proposing and a human approving each step. Only nine percent have moved decisions into autopilot, where the system acts inside defined guardrails and a person intervenes on exceptions rather than on every case.

It is tempting to read this as a maturity curve, with most companies still climbing and a handful further along. That framing is comfortable because it implies the fix is patience: keep deploying, keep building trust, and autopilot arrives on schedule. The data does not actually support that story. A company can broaden its AI footprint for years, adding models to forecasting, pricing, and replenishment, while every single decision still waits on a human click. Breadth of deployment and depth of autonomy are different axes, and the plateau sits on the second one.

Autonomy assigned to the wrong noun

Our reading of why the plateau holds is simpler than a courage problem or a model-quality problem, and less comfortable. In the patterns we see, autonomy gets granted or withheld by naming an agent or by setting a company-wide rule, rather than by classifying the decision in front of it. Neither the BCG survey nor Gartner’s note establishes that mechanism; it is our thesis about what the numbers describe. “This bot handles replenishment, so it’s cleared.” Or, in the more common and more defensive form, “our policy is human-in-the-loop,” applied uniformly across every agent regardless of what it touches.

Both units of assignment are wrong for the same reason. A single agent, say the one recommending replenishment orders, touches decisions with wildly different consequences: a routine reorder of a fast-moving staple carries almost no downside if it is slightly off, while a first order from a new supplier or a reorder that breaches a contractual minimum carries real financial exposure. Clearing the agent clears all of it at once, or none of it. A blanket organisational policy has the identical flaw at a larger scale, treating a low-stakes reorder and a high-stakes liquidation decision as the same kind of thing because they happen to run through similar software. Gartner’s 26 May 2026 finding names this directly: applying uniform governance across AI agents leads to enterprise agent failure. The unit of governance has to be the decision, not the agent and not the organisation.

A borrowed approval stamp becomes the bottleneck in an otherwise faster system.
01 / A borrowed approval stamp becomes the bottleneck in an otherwise faster system.

One replenishment order, three modes

The clearest way to see this is to watch a single decision type move through three different classes, because the type stays constant while the risk does not. Take a replenishment order for a single retailer.

Three modes, one decision type, and the difference between them is not which bot is running or which company-wide policy is in force. It is which class the specific order falls into, decided by exposure, novelty, and track record, not by label.

Three contracts define the decision, the allowed autonomy, and the operating evidence.
02 / Three contracts define the decision, the allowed autonomy, and the operating evidence.

What a contract of autonomy actually contains

Turning that logic into something an organisation can operate, rather than argue about case by case, means writing it down as a contract, one per decision class. A usable contract names five things. The decision class itself, defined precisely enough that a person or a system can tell which class a given case belongs to. The autonomy mode assigned to that class: autopilot in guardrails, mandatory sign-off, or human-only. The threshold that defines the class’s boundary, expressed as a number wherever possible, an order value, a supplier tenure, a deviation from forecast. The owner, a named role accountable for the class and for any exception it produces. And the revocation right: who can pull a class back to a stricter mode, and under what condition, held in the same document rather than scattered across separate escalation policies.

Writing these five fields down for every recurring decision type is not a formality. It is the difference between an organisation that can point to why a given order ran unattended and one that can only say the bot handles it.

What moves a decision between modes

A contract that never changes is not a contract, it is a rule pretending to be one. The fields above only earn their keep if something concrete triggers a reassessment. A breach of the class’s own threshold, an order that exceeds the capped value or a forecast deviation past its band, should trigger review of that instance and, if repeated, of the class boundary itself. A new supplier or a SKU without enough history to have earned a track record starts in the stricter mode by default, moving up only once the record accumulates. An audit finding or an exception rate that climbs past what the owner expects should trigger a downgrade, tightening the mode rather than waiting for an incident. And the owner named in the contract can invoke the revocation right directly, moving a class back to mandatory sign-off or human-only without needing to argue the point through a company-wide policy change.

None of these triggers require guessing at intent. They are observable events: a number crossed, a history absent, a rate climbed, an owner’s call. That is what makes the contract enforceable rather than aspirational.

A class-specific revocation boundary narrows autonomy when operating conditions change.
03 / A class-specific revocation boundary narrows autonomy when operating conditions change.

Where SHEPORD sits, and the question worth asking next

SHEPORD, Valnative’s retail decision operating system, is designed to carry exactly this structure: a contract per decision class, with mode, threshold, owner, and revocation right captured alongside the evidence pack and decision trace for whichever mode a class is assigned. The intent is that a class can sit in autopilot inside its guardrails, and its owner can still see every action logged and can still pull it back the moment a trigger fires, without renegotiating a company-wide policy to do so.

Before the next internal debate about whether to grant an agent more autonomy, the more useful question is narrower: which of your recurring decision types actually has a written contract naming its mode, its threshold, its owner, and who holds the right to pull it back? If the honest answer is that the policy lives in a slide about human-in-the-loop and not in a document per decision class, that is the plateau, named.

If that gap is one you recognise, the useful next step is a conversation about how a decision-class contract would be structured for your highest-volume decisions, not a pilot pitch. We are open to design-partner discussions on autonomy design in retail and CPG decision operations, and glad to walk through how the contract structure works.