Breakthroughs and research·September 20, 2026, 18:23

Poor setup of AI coding assistant can quintuple the price

AI-generated and checked against the sources listed below.

A new study shows that the software surrounding an AI model during coding tasks can make the solution up to five times more expensive without delivering better results. Simpler setups often perform just as well as complex ones.

AI-generated image

A new study called HarnessTax has taken a closer look at what happens when an AI model is wrapped in different software, a so-called "harness." It is the program that controls how the AI model uses tools such as reading files, writing code and running commands when it solves a coding task.

The researchers tested 21 combinations of seven AI models and three different harnesses: Claude Code, Codex CLI and Pi. Each combination was tried on 30 tasks from two well-known test suites for coding tasks, SWE-bench Lite and Terminal-Bench 2.0. Both how often the tasks were solved correctly and how much it cost in AI usage were measured.

Big price differences for almost the same result

One of the clearest findings is that the choice of harness can make a big difference in price, without delivering correspondingly better results. The model Claude Fable 5 solved 97.8 percent of the tasks correctly with Claude Code at a cost of $1.33 per attempt. With the Pi harness, it solved 96.7 percent, almost the same, but at only $0.67. That is roughly half the price for a difference of just 1.1 percentage points in success rate.

The explanation is that Claude Code uses ten times as much "context" as Pi from the start. Context is the amount of text (instructions and descriptions of tools) that the AI model must read and take into account before it even starts on the task. The more context, the more expensive each attempt becomes.

Simple solutions perform well

The Pi harness uses only four basic tools: read, write, edit and run commands (bash). Even so, it gave the best ratio between price and result on both test suites and achieved the highest accuracy within given budgets.

The study also shows that it is not always best to use the harness the vendor itself recommends for its model. Across six models from Anthropic and OpenAI and the two test suites, a harness other than the default choice gave the highest success rate in nine out of 12 comparisons. The model GPT-5.6 Sol, for example, solved 83.3 percent of the tasks with Pi at $0.42 per attempt, versus 78.9 percent with Codex CLI at $0.76.

The researchers recommend that developers test different combinations of model, harness and task type themselves on a representative sample of tasks, measuring success rate, price and response time, rather than simply relying on general benchmark figures or the settings the vendor sets as default.

Source

More on this topic

Get the week's AI news in your inbox

Choose your level, topics and length. One email a week, unsubscribe at any time.

Subscribe to Promptly Newsletter
PromptlyNewsletterRSSLog in

The news on aijour is AI-generated and checked against the cited sources.