Your agent didn't get dumber. Its window did.
For COO/CIO — when an agent "gets worse", the cheap and usually correct first move is window inspection, not model replacement: capability gains do not buy consistency (Princeton 2026), window discipline is a property of the harness you run and fully under your control; require that the team can rep
“The agent got worse this week.” That sentence starts more model migrations than any benchmark. Before it starts yours, one check costs an hour: what exactly was the model looking at when it failed?
A practitioner deep dive from The Carbon Layer (June 2026) walks one of the author’s own sessions. An agent sharp for thirty turns starts reopening files it already edited and proposing fixes it had ruled out. The model had not changed. The window had: filled with logs, half-read files and a stale design document, until the thing that mattered was buried.
The talk’s inventory of what you find the first time you actually inspect a window is the persuasive part. Half the window is a log nobody remembers loading. Three copies of the same notes. A tool result from twenty turns ago with no business still being there. And the instruction file everyone assumed was steering the agent, never in the payload at all.
There is a vocabulary for this rot, credited to Drew Breunig. Poisoning: a wrong fact gets in and keeps being cited. Distraction: the window grows so large the model leans on it over what it knows. Confusion: irrelevant material degrades the answer. Clash: two sources disagree. In that session the design document said the fix was deployed, production disagreed, and the agent believed the document.
Now the budget angle. A model migration is a real project with real cost, and Princeton’s 2026 agent-reliability study found 24 months of capability gains produced only small reliability improvement across major providers. Our read: a smarter model reads the same rotten window. The expensive version of that mistake, entirely hypothetical but easy to picture, is a months-long model migration aimed at a problem that lives in a log file. Window discipline belongs to the harness you run, and it is far more directly controllable by your own team.
So the gate we would put in front of any “agent got worse” escalation: reproduce the window first. If nobody can show you what the model saw, the team is debugging blind and the model is an easy scapegoat.
We build Operstead, our execution harness for agent work, on this premise: what reaches the model is an engineered, inspectable input, not an accident of accumulation.
As the talk puts it, context management is invisible when it works and looks like stupidity when it fails. An hour with the window is cheaper than a migration.