Tech

Quesma benchmark finds RTK does not generally cut AI coding costs

Tests across Claude Code and OpenCode found mixed cost results, with reported token reductions failing to translate consistently into lower bills.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Artificial intelligence

Quesma’s benchmark of RTK, a tool that filters and compresses terminal output for AI coding agents, found that lower reported output volumes did not generally reduce costs. The testing covered Claude Code with Fable 5.0 and OpenCode with DeepSeek V4 Pro 0813 on Terminal-Bench 2.1.

Across 85 Fable tasks and 89 DeepSeek tasks, with five RTK and five baseline attempts per task, costs fell 5% for Fable but rose 5% for DeepSeek. Pass rates were marginally lower with RTK, by 1% for Fable and 2% for DeepSeek.

Quesma’s task-level comparison produced a wider divergence. Fable costs were 1% higher on average, while DeepSeek costs rose 17%. Across 36 DeepSeek tasks where all attempts passed, the increase was still 18%.

RTK reported an 89% reduction in DeepSeek terminal-output tokens across 445 attempts, but Quesma said its measure is based on removed output bytes divided by four, rather than billed tokens or direct monetary savings. In some cases, smaller terminal responses were offset by additional agent turns, cached-input pricing and other changes in the agent’s behaviour.

Quesma also identified a command-rewrite issue in RTK 0.45.0 that caused one DeepSeek attempt to repeat errors and cost about nine times as much as its baseline counterpart. The issue was fixed in RTK 0.46.0 after the testing.

The findings are limited to the tested benchmark, models, software versions and RTK release. Quesma does not recommend RTK as a generic cost-saving tool, describing it instead as a niche optimisation.

Continue reading

More from Tech

Read next: Mecka AI nears US$500 million valuation in Sequoia-led funding round
Read next: Dayzle developer challenges Google Ads over suspected bot-driven installs
Read next: Moss developer Polyarc closes after nearly 12 years