Tech

Netlify tests 11 AI models on identical coding prompt

A recent benchmark by Netlify using its AXIS evaluation tool highlights the trade-offs between high-cost, high-detail models and budget-friendly alternatives in web generation tasks.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Comparative analysis reveals stark disparities in credit consumption and output quality across major providers

Netlify has expanded its platform capabilities through a new partnership with OpenRouter, enabling developers to access a broader range of artificial intelligence models, including Kimi K3, GLM 5.2, and DeepSeek V4. To assess the practical implications of this expanded choice, the company conducted a comparative test of 11 distinct AI models using an identical prompt to generate a static website for a neighbourhood coffee shop. The exercise, designed to evaluate functional correctness and credit efficiency, exposed significant variations in both output quality and resource consumption across providers such as Anthropic, OpenAI, Google, Kimi, GLM, and DeepSeek.

The testing utilised Netlify’s internal evaluation tool, AXIS, which scores models based on functional accuracy rather than design aesthetics. Results indicated that while premium models like Claude Opus produced highly detailed designs, they consumed substantially more credits than cheaper alternatives. Claude Opus demonstrated high variability in credit usage, with one run consuming 1,055 credits compared to an average of roughly 250 credits for other runs. In contrast, OpenAI’s GPT 5.6 Sol in low effort mode averaged 141 credits, while the lower-tier GPT 5.6 Terra averaged 103 credits.

Google’s Gemini 3.6 Flash averaged 103 credits, significantly outperforming the older Gemini 3.1 Pro, which averaged 53 credits but produced minimal output. Kimi K3 averaged 102 credits but was noted as being better suited for long-horizon agentic tasks rather than design-led web generation. Meanwhile, Kimi K2.7 Code averaged just 19 credits, producing simple results with limited design complexity. GLM 5.2 averaged 27 credits; however, as a text-only model, it cannot process image inputs for design inspiration.

DeepSeek V4 Pro averaged 47 credits, while the newer DeepSeek V4 Flash (0731 revision) set a new low at 2.4 credits on average. The test highlighted that while high-cost models like Claude Opus produced detailed designs, they consumed substantially more credits than cheaper alternatives like DeepSeek V4 Flash, which used only 2.4 credits on average. The findings suggest that for simple static sites, users may achieve adequate results with significantly lower credit expenditure by selecting mid-tier or open-weight models.

Netlify currently runs GPT 5.6 Sol specifically on low effort by default to provide a cost-effective alternative to Opus, though users can now adjust effort settings. The company emphasised that while Opus may offer more clever word games and sleeker designs, it carries a higher-than-average credit cost due to relentless self-validation. The results invite developers to weigh the benefits of turnkey, high-detail solutions against more iterative, budget-conscious approaches when building web applications.

Continue reading

More from Tech

Read next: DeepSeek AI releases open-source agent harness framework in developer preview
Read next: Engadget guide contrasts traditional and AI photo upscaling methods
Read next: DJI launches Mic Mini 2S with internal 32-bit float recording, excluding US market