AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The livenerf benchmark, which tests whether Anthropic’s Claude Opus 5.5 gets quietly worse after launch, is still collecting its baseline: 6 of 30 days done as of Sept 29, 2026. No verdict exists yet — the first possible call is around Oct 24. Validation showed the tool detects accuracy changes of ~7.5 points per 10-day window but could not distinguish a full swap to Opus 5.

The answer to the question “Has Opus 5.5 been nerfed?” is, as of September 29, 2026, that nobody can yet say — including the benchmark built specifically to answer it. The livenerf project, an open-source append-only monitor designed to detect whether Anthropic quietly degrades its frontier model after release, has completed only 6 of 30 planned days of data collection on Claude Opus 5.5, all within the baseline window. Its first statistical verdict is expected around October 24, 2026.

Claude Opus 5.5 launched on 2026-09-22. Livenerf began its daily measurement runs on 2026-09-24 at 22:10 UTC, roughly 2.5 days after launch. The protocol runs once a day for 30 days: days 1–10 form the baseline, followed by two 10-day comparison windows. Because the baseline is still being collected — 6 of 10 baseline days complete as of September 29, with none missed — there is currently no post-baseline data to compare against.

All six completed days ran the full 90 samples per day on the same harness hash (461391b6fce64167) and a pinned CLI version (2.1.280), according to the project’s progress log. One deviation is documented: Day 5 ran with the budget guard overridden once. The project notes that pinning the CLI is mandatory because “a changed harness looks exactly like a changed model.”

The benchmark panel consists of 78 questions selected from 2,336 screened GPQA Diamond, MMLU-Pro, competition-math and AIME 2025–26 items — specifically, questions the model answers inconsistently. The methodology follows Anthropic’s own statistical guidance (“Adding Error Bars to Evals”) and runs on Inspect, the UK AI Security Institute’s open-source evaluation framework. The panel and analysis were pre-registered before data collection.

At a glance
reportWhen: ongoing; 6 of 30 days collected as of 2…
The developmentSix days into its 30-day monitoring run, the livenerf project confirms it cannot yet say whether Claude Opus 5.5 has been degraded since its Sept 22, 2026 launch.

Why Quiet Degradation Testing Matters

Claims that Anthropic “nerfs” models after launch — through quantization, routing changes, or reduced effort — have circulated for months, but as livenerf’s documentation puts it, every argument has ended up as “vibes versus vibes” because nobody had a clean day-0 baseline. A pre-registered, statistically grounded instrument changes that: it can turn a recurring, unresolvable social-media dispute into a measurable question with a dated answer.

The stakes extend beyond one model. If the method works, it provides a reusable template for holding any AI lab accountable for post-launch changes, and its honest reporting of limits — including what it cannot detect — models how such tools should be built. Improvements are reported as prominently as regressions, which guards against the benchmark becoming a nerf-hunting machine.

Amazon

Top picks for "livenerf opus nerf"

As an affiliate, we earn on qualifying purchases.

How the Benchmark Was Calibrated

Livenerf’s design addresses a known measurement trap: questions selected for being “sometimes right” look harder than they are. On fresh samples, the panel’s pass rate rose from 54.7% to 62.0%, so the power calculations use the fresh rates. A report-only audit found 8 likely-wrong answer keys and 30 ambiguous questions among the 80 reviewed; none were dropped, but a pre-registered sensitivity analysis reruns results without them.

Validation runs quantified the instrument’s sensitivity. Deliberately lowered effort showed up in tokens before accuracy: the low setting cut output tokens by 62% while dropping accuracy by 8.3 ± 4.5 points; medium cut tokens 26% with a 4.2 ± 3.9 point accuracy drop. One full run per day can detect an accuracy change of roughly 7.5 points per 10-day window. Notably, output token count is treated as the leading indicator — “if a model quietly starts thinking less, this is where it shows up first, often before accuracy moves at all.”

A stated limitation: a validation-strength swap of Opus 5 for Opus 5.5 was not statistically distinguishable (−3.8 ± 6.3 points, −23% tokens). The project also rejects and excludes samples touched by a serving-path safety classifier that sometimes answers with Opus 5 or refuses certain questions.

“For months there have been reports that Anthropic ‘nerfs’ models some days or weeks after release… It could also mean nothing happened and people are pattern-matching on noise.”

— livenerf project documentation (GitHub)

What the Instrument Cannot See

Several things remain unresolved. First, no result exists yet — the baseline is incomplete, so any current claim that Opus 5.5 has or hasn’t been nerfed is unsupported by this benchmark. Second, the tool’s sensitivity has limits: a full model swap to Opus 5 was not detectable at validation sample sizes, so smaller changes — subtle quantization or routing tweaks — could plausibly pass unnoticed. Third, the benchmark runs through a Claude Max subscription rather than an API, and the serving path includes a safety classifier that alters some samples; the project rejects and excludes these, but this means the measured path may not represent all traffic. Finally, the pinning of the CLI to version 2.1.280 means results describe that specific harness configuration.

Timeline to the First Verdict

The baseline finishes around October 4, 2026 (day 10). The first results row appears after day 20 (approximately October 14), covering the first 10-day comparison window. The project’s primary decision — whether Opus 5.5’s scores drifted from its launch-week baseline — becomes possible around October 24, 2026, after both comparison windows close. The repository will maintain a running 10-day table reporting paired per-item score differences with clustered standard errors, plus median output token counts, against the locked baseline. The pre-registered sensitivity analysis excluding the flagged questions runs alongside the main result.

Key Questions

Has Claude Opus 5.5 been nerfed so far?

There is no evidence either way yet. Livenerf is still collecting its baseline (6 of 10 days done as of Sept 29), so no post-launch comparison has been made. The first possible call comes around October 24, 2026.

What exactly is livenerf measuring?

Daily runs of 78 calibrated questions through a pinned, frozen harness, tracking paired accuracy deltas against the launch-week baseline and output token counts, which often reveal reduced effort before accuracy drops.

Could the benchmark miss a real nerf?

Yes. The project itself reports it could not distinguish a full swap from Opus 5.5 to Opus 5 at validation sample sizes. Subtle same-family degradations may be below its detection threshold of roughly 7.5 accuracy points per 10-day window.

Why does it run through a Claude Max subscription instead of the API?

Version 0 was built around headless Claude Code (claude -p) with no API key, reflecting how many users actually access the model. Samples touched by a serving-path safety classifier are rejected and excluded.

Is the methodology trustworthy?

The panel, analysis, and validation criteria were pre-registered, the statistics follow Anthropic’s published eval guidance, and the framework is the UK AI Security Institute’s Inspect. All raw logs are retained, and known question flaws are handled via a pre-registered sensitivity analysis rather than silent removal.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Multi-Vector (Late Interaction) Embedding Models With Sentence Transformers

Sentence Transformers v6.0 adds MultiVectorEncoder, enabling ColBERT-style late-interaction retrieval for text and visual documents, with larger indexes and complex scoring.

12 Cutting-Edge AI Devices Transforming Home Automation In 2026

Explore 12 cutting-edge AI-powered home automation devices shaping smart homes in 2026, from local control hubs to DIY solutions and AI voice processors.

How to Choose AI Automation Software

Learn how to set up and use AI automation software to streamline your workflows efficiently and accurately.

Transform Your Home With These 11 AI-Powered Devices In 2026

Discover the top 11 AI-driven home automation devices for 2026, enhancing comfort, security, and energy efficiency with smart technology.