Tech

Fable 5 leads frontier AI models in new nanoGPT optimisation benchmark

Prime Intellect’s analysis of 153 autonomous runs reveals significant performance gaps among 18 leading AI models, with Fable 5 closing 81.7% of the distance to the human record.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Markets & Technology

Prime Intellect has published a comparative analysis of 153 autonomous runs conducted by 18 frontier AI models on the nanoGPT optimiser speedrun. The study, released on 22 August 2026, evaluates the optimisation capabilities of these models by measuring the proportion of the performance gap between a defined baseline and the human record that each entry closes. The benchmark serves as a rigorous test for the efficiency and problem-solving capacity of current artificial intelligence systems.

The results indicate a clear hierarchy among the tested models, with Fable 5 achieving the highest performance. The model closed 81.7% of the gap to the human record, posting a score of 2,726. This result significantly outperformed the rest of the cohort, establishing a new benchmark for autonomous optimisation tasks in the sector.

Opus 5 and Kimi K3 were identified as the next top performers in the benchmark. Opus 5 closed 53.6% of the gap with a score of 2,920, while Kimi K3 achieved a score of 2,930, closing 52.2% of the gap. Other notable entries included Opus 4.8 and GPT-5.6 Sol, which closed 39.4% and 35.9% of the gap respectively. The data highlights a steep drop-off in performance beyond the top three models, with the majority of entries closing less than 30% of the gap.

The report provides detailed metrics for each entry, including agent time, model harnesses, and resource usage. For instance, Fable 5’s best validated result was achieved using the claude-code harness in high mode, with an agent time of 8.7 days. In contrast, some lower-performing models, such as Grok 4.6 and Muse Spark 1.2, completed their runs in as little as 0.6 days, though with significantly higher scores indicating less optimisation.

The human record for the benchmark is established at 2,600, while the baseline is set at 3,290. The study measures progress based on the share of this gap closed, rather than absolute score alone. This methodology allows for a standardised comparison across different model architectures and resource constraints. The analysis includes 41 curated full agent trajectories, offering insight into the specific tool calls and subagents employed during the runs.

For investors and institutions tracking the rapid evolution of AI capabilities, the benchmark offers a granular view of where frontier models stand in terms of practical optimisation. The disparity between the top performers and the rest of the field suggests that while general capabilities are advancing, specific efficiency gains remain concentrated among a select few models. Prime Intellect’s data will likely serve as a reference point for future evaluations of AI model efficiency.

Continue reading

More from Tech

Read next: Local LLMs: Why quantised models often underperform benchmarks
Read next: Harvard Business School deploys AI avatars in $699 startup bootcamp
Read next: Meta expands Instagram privacy controls to curb off-platform data use for AI and ads