⏱️ Reading time: 11 min

Claude Opus 5.5 launched on September 22, 2026, and nine days later there’s already a repository publicly measuring AI model degradation day by day, instead of waiting for the community to sense it through scattered comments on social media.

📑 En este artículo
  1. TL;DR
  2. What is AI model degradation?
  3. Why it matters
  4. How a tracking benchmark works
  5. Practical examples and how to replicate the statistical comparison
  6. How to start running your own tracking benchmark
  7. Real-world use cases
  8. Common mistakes and best practices
  9. Comparison with alternatives
  10. Going deeper: the statistics behind the verdict
  11. Frequently Asked Questions
    1. What’s the difference between model drift and a model just having an off day?
    2. Does LiveNerf need access to the Anthropic API?
    3. How many days do you have to wait for a verdict with LiveNerf?
    4. Does AI model degradation affect open and closed models equally?
    5. Can I adapt LiveNerf’s method to another provider like OpenAI or Google?
    6. What happens if the question panel has errors in the correct answers?
  12. References

The project is called LiveNerf and tackles a concrete problem: when a model performs worse weeks after launch, all that usually exists are screenshots and complaints, never a baseline to actually compare against.

TL;DR

  • Detecting whether a model got worse requires comparing the same panel of questions against its launch baseline.
  • Lowering the reasoning effort cuts output tokens by up to 62% before the score drops.
  • A calibrated panel discards questions the model always gets right or always gets wrong, leaving only the uncertain ones.
  • With frozen prompts and a CLI pinned to one version, the only variable that changes is the model itself.
  • The method uses clustered standard errors, the same statistic Anthropic recommends for comparing evaluations.

What is AI model degradation?

AI model degradation is the gradual loss of quality a language model suffers after launch, without the provider announcing a version change. It can be caused by quantization, routing to a cheaper variant, or reduced reasoning effort, and it can only be confirmed by comparing against a fixed baseline.

The term became popular because of a specific accusation: that Anthropic “nerfs” its models days or weeks after releasing them. The accusation might be true, might be statistical noise, or might be a mix of both. The underlying problem isn’t whether the accusation is real, but that almost nobody measures it with a method that would hold up to serious scrutiny.

That’s where LiveNerf’s real contribution lies: LiveNerf doesn’t defend or accuse anyone, it runs the same panel every day since Opus 5.5 launched and lets the statistics speak first.

Why it matters

A team serving a model in production can’t rely solely on the feeling that “something changed.” If a provider lowers reasoning effort to save on compute, the final score may take a while to shift, but token consumption moves first. LiveNerf documents exactly that pattern: at low effort, output token consumption drops 62% while the score falls only 8.3 ± 4.5 points; at medium effort, consumption drops 26% and the score falls 4.2 ± 3.9 points, according to the validation published by the project.

That asymmetry explains why token count works as an early warning signal. Before accuracy shifts enough to be statistically significant, compute spend has already changed. For anyone paying per token, that data point only matters if it comes from a fixed, repeated panel, not an isolated anecdote on a forum.

💭 Key point: the project measures two signals, not one: accuracy and token count. When a model starts thinking less, it shows up in the tokens before it shows up in the score.
Daily tracking began on September 24, 2026. Foto de Enayet Raheem en Unsplash

How a tracking benchmark works

LiveNerf’s design solves an uncomfortable technical problem: a model with reasoning turned on can’t be made deterministic. There are no sampling parameters to fix, no way to switch off internal thinking. The solution isn’t to fight that, but to freeze everything else: the prompts, the exact version of the Claude Code CLI (pinned at 2.1.280), and the hash of the harness that runs the evaluation (461391b6fce64167 across the six runs logged as of September 29, 2026).

With those three elements fixed, any difference in the result has only one possible explanation: the model changed. The project runs on Inspect, the open evaluation framework from the UK’s AI Security Institute, and follows the statistical methodology Anthropic describes in its work on standard errors in evaluations: paired comparison per question, with clustered errors so each item’s difficulty cancels out instead of mixing in with the signal.

The question panel wasn’t chosen at random. LiveNerf started with 2,336 questions from GPQA Diamond, MMLU-Pro, competition math, and AIME 2025-26, filtered them with 4 samples per question, and found that Opus 5.5 gets about 93% right on the first try, and that 97% of the questions always come out correct or always incorrect. Only 78 questions landed in the middle zone, the ones it sometimes gets right and sometimes gets wrong, and those are what make up the final panel.

flowchart TD
A["2336 candidate questions"] --> B["4 samples per question"]
B --> C{"Result"}
C -->|"93% always correct"| D["Discarded: too easy"]
C -->|"97% always correct or always wrong"| E["Discarded: no variation"]
C -->|"sometimes correct"| F["Final panel: 78 questions"]

Practical examples and how to replicate the statistical comparison

The part a developer can copy without depending on LiveNerf is the paired comparison: for each question, store today’s result against the baseline result, and calculate the difference with a standard error clustered by question instead of treating each sample as independent. That way a hard question doesn’t distort the average any more than the others.

import numpy as np

def diferencia_pareada(baseline, corrida_actual):
    # baseline and corrida_actual: lists of 0/1 per question, same order
    diffs = np.array(corrida_actual) - np.array(baseline)
    delta = diffs.mean()
    se = diffs.std(ddof=1) / np.sqrt(len(diffs))
    return delta, se

baseline = [1, 1, 0, 1, 0, 1, 1, 0]
corrida_actual = [1, 0, 0, 1, 0, 1, 0, 0]
delta, se = diferencia_pareada(baseline, corrida_actual)
print(f"delta={delta:.3f} se={se:.3f}")

With that 8-question example panel, the output is delta=-0.250 se=0.164: a 25-percentage-point drop with a standard error of 16.4 points, meaning still within the noise for such a small sample. With LiveNerf’s real 78 questions and a window of ten daily runs, that same calculation has enough power to tell apart a real drop from a bad streak.

sequenceDiagram
    participant Cron as Daily cron
    participant CLI as Headless CLI
    participant Model as Model under test
    participant Grader as Automated grader
    Cron->>CLI: triggers the daily run
    CLI->>Model: sends the panel questions
    Model-->>CLI: returns answers and tokens used
    CLI->>Grader: passes answers for grading
    Grader-->>Cron: append-only log with the result
    Note over Cron,Grader: same prompt, same CLI, every day

How to start running your own tracking benchmark

LiveNerf runs on claude -p, the headless variant of Claude Code, and requires a Claude Max subscription instead of an API key. The repository uses uv as its dependency manager, with pyproject.toml and uv.lock already checked in.

macOS and Linux:

curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/ninjahawk/livenerf.git
cd livenerf
uv sync

Windows (PowerShell):

powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
git clone https://github.com/ninjahawk/livenerf.git
cd livenerf
uv sync

uv sync installs the dependencies pinned in uv.lock, so you run exactly the same environment the project used for its six logged runs. The exact commands to launch an evaluation (which script under scripts/ to invoke and with what flags) change between versions, so it’s best to follow the repository’s official README instead of guessing.

⚠️ Heads up: without an active Claude Max subscription there’s no way to run the evaluation as designed; the project explicitly doesn’t support running it against an API key.
Only 78 of 2,336 questions turned out uncertain enough for the panel. Foto de Trnava University en Unsplash

Real-world use cases

A team that pins a model version in production can run a small panel before each update, to avoid inheriting a silent regression. A company paying per token can watch output token counts as an early warning, before accuracy shifts enough to show up in business metrics.

An independent researcher can use the same method to separate a real complaint from statistical noise when the community says “the model got worse” on social media, without relying on vibes or scattered screenshots.

Common mistakes and best practices

Not freezing the exact prompt invalidates the comparison: a paraphrase, even if it says the same thing, changes the difficulty the model perceives. Confusing a change in model family with AI model drift within the same version is another common mistake: LiveNerf tested swapping Opus 5.5 for Opus 5 and couldn’t distinguish it at 99% confidence (−3.8 ± 6.3 points, −23% tokens), so this method has a clear sensitivity limit.

Picking “hard” questions without correcting for selection bias artificially inflates uncertainty; that’s why LiveNerf recalculated the accuracy rate on fresh samples before setting the statistical power. Ignoring the serving path also distorts results: sometimes the safety classifier responds with a different model or refuses to answer biology or math questions, and those samples need to be rejected and counted separately, not averaged in as if they were normal answers.

Finally, not version-controlling the evaluation harness itself is the most silent mistake of all: if the code that runs the benchmark changes without leaving a trace, there’s no way to know whether the difference comes from the model or the harness.

Comparison with alternatives

OptionWhen to use itAdvantageLimitation
Scattered reports on social mediaAs a first community warning signalAppear before any formal measurementNo baseline or statistical control
General uncalibrated public benchmark (MMLU, GPQA)To compare different models against each other onceAlready exists, no need to design itMost questions are always right or always wrong, noise masks small changes
Custom calibrated benchmark, LiveNerf-styleTo monitor a specific model over weeksDetects changes of a few points with statistical significanceRequires upfront design and months of daily runs before the first verdict

Going deeper: the statistics behind the verdict

LiveNerf’s main metric is the per-item paired difference against the baseline, with clustered standard errors, following the approach Anthropic describes in its research on error margins in evaluations. Clustering by question makes each item’s difficulty cancel out in the subtraction, instead of adding up as extra noise.

The project also documents its own data issues: an audit of the panel’s 78 questions (plus 2 excluded afterward) found 8 correct answers that look wrong and 30 ambiguous questions. Instead of simply discarding those questions, the design includes a pre-registered sensitivity analysis that recalculates the result without them, documented in its pre-registration file. That discipline, more than the final result, is what separates a serious measurement from an anecdote with a chart.

The timeline is deliberately long: the first ten days are baseline only, followed by two windows of ten days each, and the first possible comparison doesn’t land until around October 24, 2026, with the first row of results published after day 20.

Your next step: clone the LiveNerf repository, run uv sync, and check docs/DESIGN.md to replicate the same “sometimes correct” filter with your own question set.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What’s the difference between model drift and a model just having an off day?

An off day is single-sample noise; model drift is a difference that holds up against a fixed baseline, with enough runs for the standard error to back it up.

Does LiveNerf need access to the Anthropic API?

No. It runs on a Claude Max subscription using the headless Claude Code CLI (claude -p), without an API key.

How many days do you have to wait for a verdict with LiveNerf?

Ten days of baseline and then windows of ten days each; the first possible comparison lands around October 24, 2026.

Does AI model degradation affect open and closed models equally?

LiveNerf’s method doesn’t distinguish between the two by design: any model served through a fixed API or CLI can be audited the same way, as long as the provider doesn’t change the prompt or the harness.

Can I adapt LiveNerf’s method to another provider like OpenAI or Google?

Yes, the statistical part (paired comparison with clustered errors) is provider-independent; what changes is how the daily run gets triggered without using an API with variable sampling parameters.

What happens if the question panel has errors in the correct answers?

LiveNerf documented 8 possibly wrong answers in its panel and runs a sensitivity analysis that recalculates the result without those questions, instead of ignoring them.

References

📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Kashifa Sharif en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: Tech NewsTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.