Livenerf begins 30-day Claude Opus 5.5 benchmark
The GitHub project has completed six daily runs against a launch-week baseline, but has yet to publish comparative results.
A GitHub project called livenerf has begun tracking Claude Opus 5.5’s performance after release, using the model through Claude Code on a subscription. The 30-day benchmark compares results with a launch-week baseline and is designed to detect changes in either direction.
As of 29 September, six daily runs were complete, with none missed. The project says all six used the same benchmark harness and pinned Claude Code version; one run included a budget-guard override. It has not published a results table.
Livenerf’s pre-registered rule requires a change to clear a 99% interval threshold in two consecutive 10-day windows, reach an effect of at least three points and not show the same movement in the control arm. The project plans its first results table after day 20.
The benchmark uses frozen prompts, pinned tooling and fixed graders, with raw logs made public, according to the project. It also tracks output token counts as a secondary signal.
The project cautions that it measures Claude Opus 5.5 as served through Claude Code on a subscription, which may differ from the raw API model. Its launch-week baseline is a reference point, and the available data does not establish whether performance has changed.


