Field notes / Valnative

Decisions worth
operating.

Evidence-bounded essays on decision intelligence, durable execution, and the operating systems around applied AI.

RSS feed

20 long reads · 4 notes

Operstead

The score belongs to the pairing, not the model

The two numbers that sit side by side The cleanest result comes from Claw-SWE-Bench, published in June: 350 real GitHub issue-resolution tasks across eight programming languages and 43 repositories, with nine models swept against five variants of the same agent framework.

Read article

SHEPORD

The retail AI prize is on the shelf, not at the storefront.

For retail chain and FMCG supplier executives: (1) retail's AI divide is not a technology-access gap, most of the same tools are available to slow and fast movers alike; it is an investment-priority gap between where attention goes (customer-facing layers still 'too early to tell' on value) and wher

Read article

SHEPORD + Operstead

Field notes

  1. Operstead 2 min read

    Your organisation has always run on pass^k

    For COO/CIO — stop evaluating agents on demos (pass@1) and price supervision like you already do for people: score the chain not the step, put checks where error cost × observed inconsistency is highest, relax supervision as pass^k accumulates, keep blind spot-checks after autonomy.

  2. Operstead 2 min read

    Your agent didn't get dumber. Its window did.

    For COO/CIO — when an agent "gets worse", the cheap and usually correct first move is window inspection, not model replacement: capability gains do not buy consistency (Princeton 2026), window discipline is a property of the harness you run and fully under your control; require that the team can rep

  3. Operstead 3 min read

    An agent's context is assembled, not accumulated

    For COO/CIO — ask whether your agent stack has a POLICY for three flows (what enters the window, what leaves it, what gets remembered and in which kind of memory) or whether all three happen by accident; the four memory kinds need different handling, and dumping everything into chat history is the d

  4. SHEPORD 3 min read

    The thirty seconds before the prompt

    I had a very good week recently.