The remainder is harder: what automation leaves behind, and what moving people costs
What actually happened at one company Ingka Group, which operates most IKEA stores, has a customer service bot named after the Billy bookcase.
What actually happened at one company
Ingka Group, which operates most IKEA stores, has a customer service bot named after the Billy bookcase. Over two years the share of customers it could assist went from 47% to 74%.
Hold that verb. Fortune’s word is assist, not resolve, and the denominator is customers rather than contacts, conversations or intents. One customer may contact you five times; an assisted conversation may end in a call back. So the figure does not translate into a share of the queue removed, and anyone modelling from it, including us, has to say which quantity they think they are using.
The people in that queue were not made redundant. Around 8,500 call centre staff were retrained into remote sales roles, including kitchen design. Reported by Fortune at the end of July 2026, the full reskilling took two years.
Some caution about what that does and does not show. Remote sales grew 15 to 20% year on year and revenue rose from 1.08 to 1.25 billion euros, but digital sales went from negligible before the pandemic to 30% of the total, so the channel was moving underneath all of this and the article establishes no causal link between the bot and the revenue. An internal customer satisfaction score of 89%, against 60% before the bot, is the company’s own measure with an undisclosed method. Both companies in the group did lay off hundreds of corporate staff this spring, and the position that none of it was attributable to AI is theirs rather than an audited finding. One company in retail is not a market.
One more thing the reporting does not settle. Retrained is not the same as retained on the same terms. We do not know how many of the 8,500 completed the programme, how many landed permanent roles, what happened to pay and shift patterns, or what attrition looked like across two years. The case shows the scale and the intent of the programme. It does not show the outcome for the people in it.
What survives all of that is still unusual: a concrete, public transition horizon. Two years, roughly 8,500 people. Whatever else the case does or does not show, somebody stated the duration out loud, and that is rarer than it should be.
Why there has to be a transition at all
The tempting reading of a 74% assist rate is that a quarter of the work remains and therefore roughly a quarter of the people. It is worth being precise about why that reading is risky, because the precision is where the argument earns its keep.
Automation does not take a random sample of the work. It takes what it can take, and what it can take correlates with how standardised a case is. If standardised also means easier for a person, then the share automated falls as difficulty rises, and the work left behind is harder on average than the work that went in. Not as a law of nature. As a consequence of those two conditions.
They are conditions worth testing rather than assuming, because each can fail. A system that takes a flat share across every difficulty band leaves the mix unchanged. Some hard cases are highly standardised, and those go early. Demand mix can shift underneath the whole exercise. Better tooling can lower the effective difficulty of what people retain. One condition also fails in the other direction and makes things worse: easy cases that the machine mishandles come back as escalations, arriving harder than they started.
Our hypothesis is that the conditions hold often enough to plan around, and the reason is mechanical rather than empirical: a system is deployed against the cases somebody could specify, and specifiability and difficulty tend to move together. That is an argument, not a measurement, and your queue may not behave that way.
Which is worth checking rather than assuming, though not with a single number. Handling time alone is a poor proxy for difficulty: it also absorbs staff experience, slow internal systems, waiting on another team and documentation rules. A long case is not always a hard one. Look instead at several signals together on the cases automation actually took: variability of handling time, escalation rate, how many systems or teams a resolution touched, what authority level it needed, and the repeat-contact rate afterwards.
Where the conditions hold, remaining volume falls while average difficulty rises. Those move in opposite directions, and typically only one of them appears in the business case.
What the hard cases are actually made of
Two remarks from IKEA staff in that piece describe the residue better than any framework.
One worker in Sheffield noted that the AI is good on the formal terms and conditions but does not know what the customer has already been offered. That is not a knowledge problem that a better model solves. It is missing context about this specific situation, held by a person or scattered across systems nobody connected.
A kitchen specialist in Helsingborg said customers want to know what a work surface feels like under the fingertips. That one is not a context problem at all. It is a boundary on what can be transmitted through the channel the bot occupies.
Between them they mark out the residue fairly precisely: cases where the missing piece is situational rather than general, and cases where the thing being asked for is not really information. Neither category gets smaller as the model improves.
Worth being clear about what those remarks are. They are two employees describing their work to a journalist, not a measured change in case difficulty. Nobody has published a before-and-after distribution of handling times at IKEA, and the harder-remainder mechanism is our hypothesis about why transitions take as long as they do, not something that case demonstrates.
There is one broader signal pointing the same way. PwC’s 2026 Global AI Jobs Barometer, which analysed more than a billion job advertisements across 27 countries, reports that entry-level roles most exposed to AI now ask for traditionally senior, human-intensive skills far more often than entry-level roles with low exposure: advanced skills make up 52% of new requirements in the first group against 7% in the second. Note the comparison, because the popular version of this figure garbles it. It is not junior roles against senior roles. It is entry-level against entry-level, split by whether AI arrived. PwC also finds the two tracks diverging in volume, with openings for these seniorised entry roles up 35% since 2019 while other entry roles fell 10%.
Treat that as correlation, because that is what it is. Advertisements are not hires: employers copy boilerplate, ask for more than the job needs, and often settle for less. The exposure taxonomy is PwC’s own. And 2019 is a baseline containing a pandemic. Several explanations fit the same data: AI arrives earlier in sectors whose requirements were already higher, some genuinely entry-level roles disappeared so the label now sits on different jobs, or advanced economies both adopt faster and demand more. Establishing that AI made the work harder would need the same occupation tracked before and after against a control group, which this is not.
What can be said is narrower and still worth knowing: across 27 countries, employers now describe AI-exposed entry roles as needing more advanced skills than comparable roles where AI has not arrived.
The price nobody puts in the model
If the remaining work is harder, the people doing it need to be good at a different thing, and getting them there takes time. At IKEA the programme ran two years. That is the number worth carrying away: not because it transfers to your situation, but because it is a real reported horizon for moving a workforce of that size into a changed role, and most models carry no horizon at all.
That is the missing line. Not severance, which business cases do include, but the cost of carrying people across to the shape of work that automation leaves behind.
The alternative branch is not free either, though its cost is harder to pin down. An organisation that automates the entry-level work without designing a transition still needs people who can handle the residue, and it has removed one of the routes by which people used to acquire that ability. Whether that becomes a staffing problem depends on things you can actually estimate: how much of the required judgement was built on the automated work rather than elsewhere, what your attrition looks like among the people who already have it, and whether you can hire it in at a price you would accept. Those are scenarios to model, not a fate. The reason to model them is that if the risk does land, the input is time, and time is the one input you cannot buy back at the point you discover you need it.
Does any of this mean keeping the headcount
No, and it is worth saying plainly, because the argument is easy to over-read.
Average difficulty rising does not by itself determine how many people you need. Staffing follows arrival volume, handling time, service levels, what tooling does to productivity, and the economics of redeploying against hiring in. A much smaller flow of harder cases can need fewer people, the same number, or more, and the difficulty mix alone cannot settle which.
What the mix does determine is the composition. The remaining work asks for judgement the previous job may never have required, so a headcount answer on its own is incomplete rather than wrong. You still need to model how many. You additionally need to model who, and at what authority level, and the second calculation does not fall out of the first. A proportional cut applied to a changed job keeps the wrong proportion of the wrong skills, and it does so with a spreadsheet that looks entirely sound.
What is actually being decided
Which reframes the decision on the table.
Handing a class of work to a machine looks like a cost decision, gets modelled as a cost decision, and is signed off as a cost decision. It is really a decision about what the people who did that work become. Both branches cost money. One is visible immediately and lands in the training budget. The other is invisible for a couple of years and lands somewhere nobody is looking.
Choosing the first branch on purpose is a defensible position. Choosing it by default because the model only had one line in it is not a choice at all.
What to put in the business case
Three lines, none of them exotic.
State which class of work is being handed over, precisely enough that somebody can disagree with you. “Customer queries” is not a class. “Order status and delivery rescheduling within policy” is.
Name the residue and the skill it requires, before the transition rather than after. If the answer is that the remaining cases need judgement the current team has not had to exercise, that is the finding, and it has a duration attached.
Then put a duration in the model, with a range and the reasoning behind it. One published case ran to two years for a workforce of several thousand. Yours will differ, and a business case carrying no transition horizon at all is not conservative. It is incomplete.
One thing we cannot settle from the outside, and would like to hear from people who have run this. Where you have automated a queue, did the residual work actually get harder, or did better tooling absorb the difference?
For our part, we build SHEPORD around the first of those three. It is designed to make a decision class an explicit object: what falls inside it, what evidence supports a decision in it, who may authorise, and what conditions push a case back to a person. It is designed so that a class can be handed over deliberately and so that cases falling outside it surface for review rather than passing quietly.
The queue shrank. The work that was left got harder. Those are the same event, and only one of them was in the model.