Back to Insights

Nobody Took the Before Picture

Your team says the AI got worse this month. Without a dated record of your own work, that argument has no floor and no end.

8 min readBy The Bushido Collective
AI AdoptionMeasurementSmall BusinessTechnology LeadershipAdvisory
Share:LinkedInX
A developer posting as ninjahawk has been running livenerf daily since Claude Opus 5.5 shipped, and when it reached the front page of Hacker News on September 29, 2026, the README opened on a problem that has almost nothing to do with AI research.

“For months there have been reports that Anthropic ‘nerfs’ models some days or weeks after release,” it reads. “It could also mean nothing happened and people are pattern-matching on noise. Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes.”

Read that last sentence again with your own company as the subject.

The intake summaries feel thinner. Ask when it changed and you get a shrug, because nobody wrote down what good looked like in July.

Here’s what it takes to answer that question properly. livenerf screened 2,336 questions at four samples each to find the 78 that Claude Opus 5.5 only sometimes gets right, froze the prompts, pinned the CLI at 2.1.280, locked the graders, keeps the raw logs forever, and runs the whole panel once a day for thirty days. It’s built on Inspect, the open-source eval framework from the UK AI Security Institute and Meridian Labs, and its statistics follow Adding Error Bars to Evals, Evan Miller’s paper written at Anthropic, so the numbers are borrowed rather than homebrew. Day 0 is 2026-09-22, the day Opus 5.5 shipped, and the pre-registered protocol won’t let a result be called until the full series has run.

That’s the rigorous version: one question, one model, one panel, one person, and no answer yet.

The instrument is weaker than you would hope

The part of that README worth an owner’s attention is the section labeled “The limit.” As a validation test, livenerf swapped the older Opus 5 in for Opus 5.5 to see whether the benchmark would notice a different model behind the same interface. It did not. The accuracy difference came back at 3.8 points down, plus or minus 6.3, a range that straddles zero, and the README says so plainly: “This instrument can’t detect a same-family model swap of that size in a validation’s worth of samples.”

That limit is narrower than it sounds: a same-family swap, of that size, in a validation’s worth of samples. That’s a narrow finding rather than a verdict on benchmarks, and what it lands on is the metric almost anyone would reach for first. On accuracy, a different model behind the same name did not register.

Something did register. Output tokens dropped 23 percent. The README gives its token deltas as bare percentages, without the intervals it puts on accuracy, so hold the figure loosely; what stands is the gap between a 23 percent move and an accuracy difference that straddles zero. The same gap runs through the effort tests: “Lower effort shows up much more clearly in tokens than in accuracy.” Dropping effort to low cut output tokens 62 percent against an accuracy move of 8.3 points, plus or minus 4.5; medium effort cut tokens 26 percent against 4.2 points, plus or minus 3.9.

Which makes us read the hunch in your shop differently, and this is inference rather than evidence: “the summaries feel thinner” is a report about length, and length is where livenerf’s change showed up. Whether your team’s thinness is the same thing as a token count is exactly what a dated record would tell you and a hunch never will. What the hunch certainly cannot settle is whether the answers got worse, because that is where even a pre-registered benchmark came back ambiguous.

These are 78 hard academic questions, though, not your quoting workflow. livenerf’s author does claim an ordering, writing that if a model quietly starts thinking less, “this is where it shows up first, often before accuracy moves at all,” but that is their reason for tracking tokens, not a published result. What we infer is narrower: in that comparison, cost was the easier of the two to see. Watch only whether the output still looks acceptable and you are watching the harder channel.

The rate card is not your cost per unit of work

That same day, September 29, Artificial Analysis reported that GPT-6.1 Sol had replaced GPT-6 Sol after seven days. Nothing on the price list went up: per-token pricing held at $2 and $10 per million input and output tokens, and the cache read discount improved from 90 to 95 percent. In the same measurement, the new model used roughly 10 to 30 percent more output tokens than the one it replaced.

Those two pull against each other. More output tokens at an unchanged output price raises the bill; a better cache discount lowers the input side for anyone sending long, repeated prompts. Output-heavy work got more expensive, cache-heavy work with short answers may have gotten cheaper, and a rate card cannot tell you which one you are. In the Hacker News thread on the replacement, one commenter, Dinuda, described how that lands on a subscription: “After 5.5, I basically don’t notice a jump in model performance, other than my usage ending sooner.” An allowance running out sooner is what more output tokens per task would feel like on a flat plan, but a per-token rate card can’t confirm it, and one comment doesn’t establish it. The durable part is the arithmetic: the published price and your cost per finished quote are two different numbers.

The explanation we can’t rule out

The Hacker News thread on livenerf is worth reading for its skeptics. Some of its sharpest comments argue the complaints are psychological, and the best of those came from nsarrazin: “Perceived performance is actual performance over expectations and the latter just keeps increasing over time.”

Take that seriously, because it cuts against us as much as the vendors. If expectation drift is what is happening in your case, your team’s report that the tool got worse is weak evidence, and so is your own impression that it didn’t. Two things may have moved under you at once, the tool and the standard you judge it by, and you have a record of neither.

Record the small thing, not the benchmark

You’re not trying to prove a vendor did something. You’re trying to answer your own question the next time somebody says it got worse. That’s a far lower bar. It’s still real work, since somebody has to pull the cases and agree on what a good answer was. But livenerf spent four samples on each of 2,336 questions just to choose which 78 to keep, then runs that panel every day. Yours is twenty or thirty cases, judged once, by somebody who already knows what a good answer looks like.

In priority order, start with the one workflow where a customer or a dollar depends on the output, not the demo that impressed everybody. Freeze a panel of real past cases from it, twenty or thirty, each with the answer you would have accepted and the date you accepted it. Write that answer down while you still hold the old standard, because a judgment re-made from memory next year drifts in exactly the way the thread says your expectations do.

Then record the whole trace rather than the verdict: the input, the output, what it cost, and the version of every tool involved. On a flat subscription, cost means whatever units your provider actually shows you, against the allowance. livenerf pins a CLI number and a harness hash for exactly this reason, and it is the cheapest part of the whole exercise to get right.

Put cost per case on the same page as quality, since that was the easier channel to read on livenerf’s panel and the cheaper of the two to log. Then decide now what a bad reading means you will do, while nothing is on fire and nobody is defending a purchase.

Nobody can show you a finished public version of this, including us. livenerf’s own answer isn’t in yet.

Skip all of it if AI touches nothing a customer or a dollar depends on. You don’t have this problem yet, and building the panel would be the science project you were right to refuse.

If you have an engineering team with real evaluation discipline, this is theirs to own. Ask them for the frozen panel and the dated log, and if they produce both, you are finished and you don’t need outside help.

Public leaderboards and vendor dashboards are genuinely useful, and they are structurally incapable of the thing you need, because nobody outside your company holds a panel of your quotes. That is a boundary, not negligence.

We work this way for a selfish reason too. Every engagement carries a number before work starts and a monthly report in the client’s own numbers, revenue, margin, hours back. Without a dated before, we would end up arguing with a client exactly the way that thread argued with a vendor, and we would deserve to lose. If you want that discipline pointed at a workflow that is already carrying weight, that’s what we do.

Want this looked at in your business?

Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.

Get your rough map, free

Not ready to talk? Stay sharp anyway.

We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.

Keep reading

Share:LinkedInX