Groupthink used to need a group

For COO and decision owners — AI does not remove group bias from decisions, it reproduces the worst dynamic of a meeting more cheaply and less visibly.

A fractured navy ring loses one brass shard, breaking apparent agreement.

The signatures transfer, and one of them transfers wrong

Hold the four signatures against a working session between an executive and an assistant and three of them arrive intact. A request carries a premise, and the model tends to develop that premise rather than contest it, which produces unanimity without anyone agreeing to anything. Justification is instant, so a weak idea acquires a persuasive shape before a doubt has finished forming. And the tone of an answer bears little relation to the uncertainty behind it, so a shaky inference and a solid one arrive in the same register.

The fourth signature transfers differently, and the difference matters. Janis described pressure on dissenters inside a group. In practice the pressure has not disappeared, it has changed target. Nobody is arguing with a machine. They are arguing with the colleague who brought the machine’s analysis into the room, and that colleague now holds a stronger position than they did an hour earlier, because their case arrives with apparent independent support attached. Dissent has not become socially awkward. It has become more expensive.

BCG’s June 2026 study of seventy senior executives found roughly nine in ten pointing to over-reliance on AI outputs that nobody stress-tests, and more than half describing a softening of ownership, where “the AI suggested it” becomes a shield. That is perception data rather than a measured effect, and it is the people closest to the problem describing what they see. The causes transfer as cleanly: familiarity with a tool does the work cohesion used to do, even an intact, well-assembled context window works as an echo chamber, because it contains what you put into it and little else, and time pressure needs no translation, with Deloitte’s 2026 survey finding sixty-nine percent of respondents reporting heavier workloads under AI-driven change, and forty-three percent naming lack of time as their main barrier to adapting.

The part that actually costs money

None of the above is the real problem. The real problem is a specific illusion the pair creates and a meeting never could.

When a group agrees with you, you know it shares your incentives, your information and possibly your blind spots, and you discount accordingly. When a system agrees, it feels like an independent check performed by something without ego or politics, using more data than you have. That feeling is the error. You supplied a framing and received it back, elaborated and confirmed.

The consequence is behavioural and expensive. An executive who knows the room is deferential asks a second person. An executive who believes an outside opinion has already been obtained asks no one.

Here is what it looks like in an operation. A category team suspects demand for a slow line is fading, and asks the system to size the exposure and recommend an action on that basis. The analysis comes back thorough: sell-through curves, ageing stock, a recommended reduction in the next order. What it does not contain is the fact that would have changed the answer, because nobody asked for it and the framing did not invite it. A competitor two kilometres away has been running a deep promotion on the same category for three weeks, which is why the line looked soft, and that promotion ends on Sunday. Cut the order now and you are short precisely when demand returns. The recommendation is not wrong given its premise. The premise was never tested, and the confident output made testing it feel unnecessary.

There is a moment worth watching for in yourself and in your teams, and it deserves a name: the point at which someone stops forming their own position before consulting the system. Their capacity to judge is intact; they simply stop exercising it first. Not the point where they use the system, and not the point where they trust it, but the point where they no longer arrive with a view of their own to be tested. Call it cognitive surrender. It looks like efficiency, it feels like fluency, and it removes the only thing the system could usefully have disagreed with.

Four distinct review signatures separate independent checking from repeated agreement.
01 / Four distinct review signatures separate independent checking from repeated agreement.

Why “think critically” is not the answer

The reflex is to prescribe vigilance: train people to challenge outputs, remind them that AI errs. It tends to fade at operating scale, for the same reason that telling a deferential room to speak up rarely holds. The behaviour is produced by the structure, and exhortation leaves the structure intact. Under deadline, in the twentieth decision of the day, vigilance competes badly.

The alternative is to change what a recommendation is allowed to look like. If the system must show disagreement, the person does not have to manufacture it.

The practical form is one page attached to a recommendation, with five fields, each carrying an owner and the date its data was current.

The originating hypothesis, written down as the person actually framed it, because everything downstream inherits it and it is normally invisible by the time a decision is reviewed. The strongest counter-argument, stated at its strongest rather than as a straw man. The evidence that pointed the other way, since almost every real decision has some and its absence from a recommendation is itself informative. The reversibility of the action, because a cheap mistake and an expensive one deserve different scrutiny and rarely get it. And the boundary of confidence, expressed so that a well-supported inference is distinguishable from a plausible guess.

A sixth line is worth adding for the discipline it imposes: what would change this answer. It converts a conclusion into something falsifiable rather than something to accept or reject wholesale.

An answer containing none of these does not show that it was checked. It shows that it was echoed.

The industry is already building the antidote

Something worth noticing: the people building serious agent systems have been converging on the same answer, and they arrived at it as an engineering requirement rather than as a nod to open-mindedness.

Sakana’s Fugu line runs what amounts to a conductor pattern. Rather than sending a hard problem to one model and accepting what comes back, it fans the problem out to several specialist models and then synthesises across their answers. The architectural bet is explicit: the value is in the spread between the responses, and a single response has no spread to inspect. OpenRouter’s fusion approach makes a similar wager from another angle, taking answers produced independently and merging or ranking them into one result, so that a conclusion has to survive contact with alternatives before it is served.

Both designs quietly concede the point of this article. If agreement from one capable model were sufficient, neither pattern would be worth its latency or its cost. They exist because a single articulate answer is not evidence, and because the useful signal appears only when several attempts, made from different starting points, are laid side by side and disagree in specific places.

The word for what they are buying is diversity, and it helps to strip the term of its social connotations here. In this context it is a purely mechanical property: independent sources fail differently. Two models with the same training lineage, given the same framing, tend to be wrong in the same direction, which is why asking one system to check another’s work inside the same context adds confidence without adding information. What produces a genuinely new signal is a different starting point, whether that comes from a different model family, a different prompt framing, a different data window, or a person who has not read the first analysis.

That is the same principle as the opponent rule below, arriving from the technical side. Fan out, then synthesise. Disagreement is not friction in these architectures. It is the product.

A second opinion echoes the first path instead of creating independent evidence.
02 / A second opinion echoes the first path instead of creating independent evidence.

The opponent, named

The second move is organisational and older than any of this technology. On decisions where being wrong is expensive, someone should hold the explicit job of arguing the other side, and it should be a role rather than a temperament, assigned and rotated, so that dissent is not left to whoever happens to be brave that week.

The reason to formalise it now is that the pair has removed the accidental friction that used to serve this purpose. A meeting produced at least the possibility of a raised eyebrow. A session with a system produces none, so the friction has to be designed back in.

One rule makes the difference between a real opponent and a ritual one. Above a defined cost of being wrong, the challenge must start from a different set of premises, not continue the same conversation. Asking the same system to critique its own output, inside the same context, is not a second opinion. It is the same meeting with a different chair. In practice that means a separate brief, written by someone who has not seen the first analysis, a recorded premise delta showing which hypothesis, data window or success criterion was changed, and a route by which the finding can actually stop the decision rather than be noted.

A review record requires rejected options, conflicting evidence, a confidence boundary, and what would change the answer.
03 / A review record requires rejected options, conflicting evidence, a confidence boundary, and what would change the answer.

Where SHEPORD sits in this

SHEPORD, Valnative’s retail decision operating system, is designed so that a decision arrives as a case rather than a conclusion. The evidence pack that accompanies a recommendation is designed to carry the options that were considered and rejected, the evidence that pointed elsewhere, and the constraints that bounded the choice, with the approval and its reasoning recorded alongside. The intent is that the disagreement is in the artefact, where a reviewer cannot miss it, rather than in the reviewer’s own supply of scepticism at the end of a long day.

One question to run this week

Take a recommendation you accepted in the past week, ideally one you felt good about. Then ask what it disagreed with you about.

If nothing comes to mind, there are two possibilities. Either your framing was correct in every particular, or you have been having a very agreeable meeting with yourself.