Tech

Opus 5 leads but fails to achieve autonomous code maintenance in new benchmark

Claude Opus 5 achieved a 24% strict pass rate on a subset of coding challenges, significantly outperforming peers, yet all models struggled to maintain codebases without human intervention.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
Independent testing on University of Wisconsin-Madison’s SlopCodeBench reveals persistent defects and rising complexity in frontier AI models

A recent independent benchmark conducted by a GitHub user has placed Claude Opus 5 at the top of the leaderboard for SlopCodeBench, a long-horizon coding evaluation developed by a laboratory at the University of Wisconsin-Madison. The test, which ran three frontier models through a subset of 17 checkpoints across three distinct coding challenges, revealed that while Opus 5 technically outperformed its competitors, no model demonstrated the ability to reliably maintain software engineering tasks without human steering.

The benchmark measures a model’s capacity to evolve a codebase over time as new requirements are introduced, requiring the model to pass held-out black-box tests without introducing regressions. On the tested subset, Opus 5 achieved a 24% strict pass rate, securing four out of 17 checkpoints. In contrast, both Opus 4.8 and Sonnet 5 achieved a 6% strict pass rate, each passing only one checkpoint. The results suggest that while Opus 5 is the current leader, the benchmark remains unsaturated, indicating significant room for improvement in the next generation of models.

Despite the higher pass rate, the testing exposed systemic issues with code quality across all models. The analysis showed that every model increased code complexity and verbosity as the challenges progressed. Opus 5 generated five times the number of functions compared to Opus 4.8, with a significant portion of these being single-use callables. The author noted that while small functions can be beneficial, the sheer volume generated by Opus 5 represented expensive verbosity that did not directly translate into superior results or maintainability.

Defect accumulation was a consistent failure point, with no model managing to complete any of the three challenges without introducing errors. The strict pass criteria required that new code remain green while also passing all inherited regression tests from previous checkpoints. Opus 5 maintained a lead in the early stages of the circuit_eval challenge but eventually accumulated defects in subsequent checkpoints. The data indicated that as new requirements conflicted with initial designs, models struggled to refactor effectively, leading to steady increases in code duplication and complexity metrics.

The findings reinforce the conclusion that current frontier models cannot run software engineering workflows "lights-off." The author argued that while metrics such as cyclomatic complexity and duplication provide directional signals, they do not fully capture maintainability. The benchmark highlights a critical gap in AI capability: while models can solve discrete problems, they lack the architectural discipline to maintain a coherent codebase over an extended horizon. Investors and technology leaders should view these results as a signal that human oversight remains essential for autonomous coding agents in the near term.

Continue reading

More from Tech

Read next: Anthropic’s Amodei rejects open-weight bans, warns of Chinese AI threat
Tech
Image unavailable
TechDraft

DConf 2026 to convene in London this September

The annual gathering for the D programming language community will take place from 2–4 September, with organisers advising early travel bookings and strict adherence to official UK government channels for entry requirements.

Tech DeskRead story
Read next: DConf 2026 to convene in London this September
Read next: France Records First Pyrocumulonimbus Cloud as Wildfires Shatter Historical Averages