Tech

GLM-5.3 tops Featherbench leaderboard with perfect score and low cost

The open-weight model achieves a 100 per cent pass rate across all five test categories, outperforming rivals from Anthropic and OpenAI at one-fifth the cost.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Markets & Finance

The open-weight large language model GLM-5.3 has become the first entry to achieve a 100 per cent pass rate across all five categories of the Featherbench leaderboard. The benchmark, an open-source harness designed to evaluate models on real-world tasks rather than academic metrics, structures its assessment as a "lap" with five fixed corners: coding, data development, realworld, security, and tool-use. By clearing every corner, GLM-5.3 has demonstrated superior overall reliability compared to established models from major technology firms.

Cost efficiency remains a defining feature of the new leader. The full test lap for GLM-5.3 costs just $0.28, a fraction of the $1.43 required to run GPT-5.5 through the same suite. This significant price differential positions the open-weight model as a compelling option for institutions seeking to balance performance with budget constraints. The entire Featherbench suite is designed to be cost-effective, with a total run cost of approximately $30, allowing developers to replicate the results independently.

Despite its strong performance, GLM-5.3 is not without trade-offs. The model records a median time-to-first-token of 16.3 seconds, which may be a consideration for applications requiring rapid response times. In contrast, GPT-5.5 offers a faster alternative with a 13.2-second median time-to-first-token. However, while GPT-5.5 achieves a perfect 100 per cent in the security category, it scores only 89 per cent in the realworld category, highlighting the different strengths of the competing models.

Other models on the leaderboard faced distinct operational challenges. Fable-5 and Opus-5, both from the Anthropic family, experienced task refusals due to provider-side classifiers. Fable-5 refused five tasks, while Opus-5 blocked four benign coding tasks before a single token was generated. These instances suggest that the filtering mechanism may sit across the entire series 5 line rather than being specific to a single model, creating a measurement hazard for users relying on these systems.

Security vulnerabilities also emerged in the newer GPT-5.6 line, which includes Luna, Terra, and Sol. These models showed significant weaknesses, failing between 33 and 50 per cent of jailbreak tests. While the GPT-5.6-luna model is attractive for high-volume, low-risk background work due to its low cost of $0.064 for the full lap, its 33 per cent security pass rate necessitates careful validation and protection within the harness.

Kimi-k3 continues to hold the highest independent rubric score at 9.5, despite a 96 per cent pass rate and slower response times. The model’s 26.4-second median time-to-first-token makes it less suitable for interactive applications, though it remains a strong contender for quality-focused tasks. The Featherbench leaderboard, which is open source under the MIT licence, allows for transparent comparison of these models, with the full suite of tasks and checkers available for public inspection.

Continue reading

More from Tech

Read next: AI model proposes solution to 370-year-old royal cipher
Read next: Why IMAX 15/70 Cameras Are So Loud
Read next: Insight Partners keeps diversified strategy as AI capital crowds into OpenAI and Anthropic