Tech

Writer cuts enterprise AI costs by 50 per cent with Palmyra X6 launch

Writer CEO May Habib says enterprises are prioritising cost flattening over benchmarks, citing distrust of major labs’ pricing incentives.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · View original source
Writer introduces new AI model and upgraded harness to contain token costs
New flagship model and upgraded harness infrastructure target token price sensitivity

Writer has launched its flagship AI model, Palmyra X6, alongside an upgraded agentic harness infrastructure designed to reduce enterprise deployment costs. Built as a post-training variation on Z.ai’s open-source GLM-5.2 model, the system aims to provide deployment-ready capabilities at a lower price point. The company estimates that the combination of the new model and harness changes will cut customer costs for basic tasks by up to 50 per cent. The features are available to Writer clients immediately.

This launch occurs amidst a broader industry shift where enterprises are prioritising cost flattening over chasing new benchmarks. Writer CEO May Habib cited distrust towards major AI labs regarding their financial incentives to drive up token use, noting that chief information officers are increasingly giving up on these providers due to unprecedented cost explosions.

Writer researchers published a paper suggesting that harness efficiency changes can reduce costs by an average of 40 per cent across multiple models, often more reliably than model choice alone. The new approach emphasises complex, multi-step tasks executed faster and with fewer tokens, positioning the harness as a critical lever for efficiency.

Palmyra X6 will sit alongside other Writer models or outside models imported through Azure or Amazon Bedrock, maintaining a model-agnostic experience for clients. This strategy allows businesses to leverage open-source foundations while retaining the flexibility to integrate external systems without disrupting existing workflows.

The move reflects a growing urgency among users to manage deployment expenses. While open-source models typically offer lower per-token costs, selecting the right model for specific jobs has historically been difficult. Writer’s latest offering attempts to solve this by combining a cost-effective base model with infrastructure improvements that maximise output efficiency.

Continue reading

More from Tech

Read next: GameCube’s library still commands attention 25 years on
Read next: US AI leaders urge restraint as Trump team prioritises China competition
Read next: WIRED names Sonos Arc Ultra its best overall soundbar for 2026