The order was right. The shelf was empty.
A counting result that should change what you fund In a European grocery chain, researchers ran a field quasi-experiment.
A counting result that should change what you fund
In a European grocery chain, researchers ran a field quasi-experiment. One store received a full stocktake, a matched store did not, and both were tracked for two months either side. Across 5,657 directly comparable SKUs, the audit was associated with an estimated 11% store-wide sales lift.
The size of the lift is not the interesting part. Where it landed is. The gain concentrated on items where the system believed it held more than the store actually had.
It is worth being exact about what the intervention was, because it was not only counting. Before the count, a preparation team tidied the back room and the shop floor, and easy-to-spot misplaced items were brought back to their correct locations. Two things therefore happened together: goods that were in the store but unavailable became available, and records that overstated reality were corrected.
That bundle is the finding, not a flaw in it. The authors describe the audit as acting “not simply as a counting device, but as a visibility-restoring intervention that makes overstated records observable and therefore correctable”.
And it lands directly on the measurement problem. Because the intervention bundled both, the experiment cannot attribute the lift between them. Neither can a weekly aggregate, unless the two pathways were instrumented separately before the fact. That is the whole argument for doing it beforehand: after the fact there is nothing to separate with.
The order was locally right, and the system was still wrong
There is a comfortable version of this story, and it is worth refusing. The comfortable version says the replenishment decision was correct on the inputs it had, so the fault lies elsewhere. Given a record of 100 units, ordering nothing is the right arithmetic.
But a system that treats a stock record as ground truth is not neutral. It has made a silent assumption, and the assumption is wrong in a way that is knowable in advance. Records drift, and they drift more in some places than others. A replenishment engine that consumes a quantity without any notion of how reliable that quantity is has not been let down by its inputs. It was built to ignore a property of its inputs that matters.
So the defect is not the arithmetic. It is that the number arrives without a confidence attached to it, and nothing downstream is allowed to ask.
That reframing matters because it moves the work. If the order was simply misinformed, you go and clean the data forever, which is a project with no end. If the decision is missing a reliability signal, you can change the decision: treat a quantity you have reason to doubt differently from one you do not, and put counting effort where overstatement is most likely.
The trap: two tidy numbers, equally false
The obvious response is to split the availability number in two. Nothing in the store is one failure. In the store but not on the shelf is another. Different processes, different owners, different fixes.
The split is right, and it fails in a specific way that is worth naming, because it is the most common way this work dies.
If you classify the two using the same stock record you already distrust, you have not built two numbers. You have built one untrustworthy number twice. Every case where the record overstates reality gets filed under “in the store, not on the shelf”, because the record says the goods are there. The store gets a shelf-execution problem it does not have. Replenishment gets a clean scorecard it has not earned. Both numbers look sharper than the one they replaced, and both are wrong in the same direction.
This is why the third row of the table below matters more than the first two.
| What actually happened | Who can act | What the record says |
|---|---|---|
| Nothing in the store | replenishment, logistics, supplier | zero, correctly |
| In the store, not sellable or not on the shelf | store execution: receiving, filling, quality, planogram | a quantity, correctly |
| The record says more is here than there actually is | on paper inventory control, store operations, loss prevention; in practice nobody until someone counts | a quantity, incorrectly |
The third row usually has owners on paper. Inventory control, store operations and loss prevention all have a claim on it. What it lacks is a trigger: nothing in the daily flow raises it, so ownership only activates after somebody has already found it by other means. That is worse than having no owner at all, because the org chart looks complete. It is invisible to the ordering system and to the store scorecard, and it surfaces mainly through an activity that costs money and produces no output of its own. It is also the class where the experiment’s gain concentrated.
There is a market reading alongside this. IHL Group surveyed executives at more than 200 of the largest US retailers in November 2025 and reported that fewer than one in four reach 80% accuracy on on-shelf availability, planogram compliance and promotional execution. Three separate disciplines with three separate owners, failing together. My reading, not theirs: they fail together because they are managed as one line on one dashboard.
And even the first two rows are cleaner on paper than in a store. A single empty facing can pass through a short delivery, a receiving error, a back room, a planogram change and a filling routine on the way to being empty. What you are separating is not really departments. It is decisions: what was ordered, what was accepted, what was believed, what was placed. Departments come attached to those decisions, but the decision is the thing with an owner.
Reason codes become a political system
The standard advice at this point is to attach a cause to every stock correction, on the argument that a corrected quantity fixes today and a cause changes tomorrow. That is true and it is not sufficient, because a cause is not an observation. It is a statement made by somebody with an interest in the answer.
Watch what happens when the same case is touched by several hands. The store selects “delivery”. The distribution centre selects “receiving”. An overnight correction rewrites the quantity and the sequence of events is gone. Nothing in that flow is dishonest, and the output is still unusable, because three parties recorded three versions of one event and the system kept the last one.
The design questions that decide whether any of this works are unglamorous and they are the whole game.
Can the record hold two conflicting versions of the same case without collapsing them into one. Who has the authority to close a case, and is that person different from the people who filed the causes. Does a cause survive the next overnight batch. And when a cause turns out to be wrong three weeks later, does anything reopen.
A cause that cannot be contradicted is not evidence. It is a vote.
Two numbers worth asking for
If you want a short test of whether this loop exists in your organisation and not only on a slide, there are two numbers worth asking for.
The first is the time from the first signal that an item is unavailable to the moment it is back on the shelf, at the median and at the ninetieth percentile. The median tells you about the good days. The ninetieth tells you what the loop does when it is under load, which is when it matters.
The second is the share of cases that cannot be assigned an owner within twenty-four hours. The reaction to that request often tells you as much as the answer would.
Neither of these requires a new system to start measuring. Both of them are uncomfortable, which is the point.
Where the intervention makes things worse
Counting more is not a general answer, and the study points at exactly the goods where the temptation is strongest and the risk is highest: record inaccuracy was higher on perishables, on items restocked frequently, and where inventory levels were higher.
The next step is mine, not theirs. Those same goods, particularly the slow-moving and high-shrink end of them, are where an availability programme most often turns into a waste programme. If the response to a gap is to order more, you convert an availability number into a write-off, and the store learns quickly that the programme costs it money.
So the distinction has to be held firmly: correcting a record downward is not the same instruction as ordering more. The first improves what the system knows. The second changes what the system commits to. The audit study measured the first. It did not measure a policy of ordering more, and neither should anyone quoting it.
There is a companion finding in the same work that cuts against intuition. Inaccuracy was lower on items running promotions. My reading rather than their conclusion: attention is part of the mechanism, and the goods that people are actively watching drift less. Which suggests the targeting question is not only about product characteristics but about where nobody is looking.
The part that is not a data problem
It would be convenient to conclude that retailers have not seen the evidence. That is not what is happening.
The cost of counting sits in store labour. The benefit shows up in category sales. Different budget, different owner, different reporting line, and the person who pays is not the person who gains. That is why a good result does not spread on its own merits, and why another dashboard does not move it either.
The honest version of the recommendation is not “measure availability better”. Decide who owns each failure by the decision that produced it. Express the number in lost selling time rather than counts. Keep causes that can be contradicted, and name who may close a case. And put the cost of counting in the same conversation as the sales it protects, because if those sit in two conversations the mechanism regenerates the loss every day, whatever the forecast does.
Where we are building
SHEPORD is designed as a decision operating system for retail. The seam above is what it is built around. It is designed so that a decision carries its inputs, the authority under which it was made and its trace, so that an outcome can be resolved back to the decision that produced it and to the picture of the world it was made on. In this problem that means a stock correction is designed to behave as a decision with an owner and a stated cause rather than as a data-entry event, conflicting versions of a case are designed to survive rather than be overwritten, and the cause is designed to reach the replenishment decision instead of resting in a comment field. Operstead is the layer designed to execute what has been authorized and return a receipt for it.
We are in design-partner conversations with retailers working on exactly this seam.
The question worth taking into the next availability review is not which store is failing. It is which decision produced this empty shelf, what it believed at the time, and who was allowed to tell it otherwise.