​

FinOps for AI coding agents: Stop counting tokens. Start pricing the fix.

Contents

What does it cost an AI agent to fix one bug? The price of a million tokens won’t tell you. As agentic AI systems take on longer, messier jobs, the useful FinOps for AI question shifts from AI usage and AI cost per token to the business outcome—what the business actually got: an accepted change.

 

Four words. Millions of tokens.

Picture a developer typing four words: “Fix the failing test.”

The agent opens files, edits a function, runs the tests, hits another failure, reads the error, and tries again. To the developer, that is one task. To the meter, it is a chain of model calls—and every call can resend the context the agent has accumulated so far.

That difference is where the economics get strange. In a 2026 study, eight AI models were run as coding agents across 500 real GitHub issues. The average run consumed about 4.2 million tokens, roughly 1,200 times a multi-turn coding chat, and cost about $1.86. The comparison chat cost about two cents.

The instruction is tiny. The work it sets in motion has no fixed size.

 

Seats are the wrong unit

Seat-based budgeting answers a simple question: how many developers have access? It says much less about how much autonomous work those developers will trigger—or whether that work produces anything worth keeping.

GitHub made that gap explicit on June 1, 2026, when it moved Copilot from premium requests to usage-based AI Credits. Under the old model, a quick chat and a multi-hour autonomous coding session could look identical from a licensing perspective. They are not identical AI workloads.

Use GitHub’s own numbers. Copilot Business costs $19 per seat and includes $19 in monthly AI Credits, pooled across the organization. Fifty developers therefore share a $950 monthly pool. At the study’s average of about $1.86 per task—from older models, so treat this as an order-of-magnitude example—that pool covers roughly 500 tasks, or about ten per developer. Whether ten is plenty or nowhere near enough depends on the team. Only measurement will tell.

Seats are only one route for AI spend to hit the P&L. The others are per-token API bills from AI providers and dedicated or self-hosted capacity, where the cost per token is something you calculate rather than read from a price list. Finance needs all three in one cost management view.

 

The tax on memory

An agent cannot fix a bug from the latest instruction alone. It needs repository rules, relevant files, earlier decisions, and the latest test output. Then it carries that material into the next call. And the next. A larger context window changes how much fits into a request; it does not make the session free.

Take a simple 20-call run. The first call carries 20,000 tokens of context. Each pass adds about 4,000 more. By call 20, the context has grown to 96,000 tokens. Output across the run totals 20,000 tokens. See Figure 1 below:

Fig. 1 Why long agent sessions get expensive

At Anthropic’s published Claude Sonnet 5.5 rates, checked October 6, 2026—$2 per million input tokens, $10 for output, $2.50 to write cached input for five minutes, and $0.20 to read it—those 20 calls cost about $2.52 without caching. That is $2.32 for 1.16 million input tokens, twelve times the final context, plus $0.20 for output.

Turn caching on and the same illustration falls to about $0.65, roughly 74 percent less. This is not a forecast, and it excludes non-token charges. But scale the example to 50 developers running 100 such workflows a month each and the gap is about $9,300 a month.

Caching makes memory cheaper. It does not make memory small. In the example, the last call carries nearly five times the context of the first. In the study’s breakdown for Claude Sonnet 4.5, cache reads were the largest cost item in every phase of a task, while spending spiked when the agent pulled fresh content into context. Unless someone controls what the agent reads and keeps, every step can make the next one more expensive.

 

Same job, wildly different bill

Two 2026 studies have started to show where those tokens actually go.

The first, How Do AI Agents Spend Your Money? (Bai et al., April 2026), ran eight models through 500 real GitHub issues, four times each, while carrying the full conversation history into every round.

  • Input drove the bill. Input tokens, not output tokens, were the main cost driver, even with caching.
  • The same task could swing dramatically. Across repeated runs, total token use for one task could differ by as much as 30 times; more typically, the most expensive run cost about twice the cheapest.
  • More spending did not mean more success. Accuracy often peaked at intermediate cost. On tasks no model could solve, models spent more tokens than on tasks all of them solved, suggesting they do not reliably know when to stop.

 

The second study, Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering (Salim et al., January 2026), is a work in progress. It followed 30 small tasks through the ChatDev multi-agent framework with GPT-5. Automated code review consumed an average 59.4 percent of the tokens. Initial coding used just 8.6 percent.

The signal is hard to miss: much of the money goes into iteration and repeated context, not the first draft of the code. A cost model that ends at “generate the code” misses a large part of the process.

There are caveats. These are research setups, not production engineering teams. Token share is not the same as dollar share. The 30-times spread is an extreme. Public data from real organizations remains thin, which is precisely why companies need to measure their own workflows instead of treating one successful trial as a price quote for the next.

The metric that matters: cost per accepted change

A token is a billing unit. A company is buying a business outcome. The FinOps Foundation makes the same distinction in its writing on tokenomics: the useful question is the cost of the result, not the volume of the input. That is AI unit economics in practice: tokens and model calls are accounting units; FinOps guidance lists cost per successful outcome as a recommended unit metric.

For coding agents, that result is an accepted change. Count every attempt—including failures—and include the human effort required to review and repair the output. See Figure 2 below:

Fig. 2 The cost per change formula

“Accepted” needs a written definition before the meter starts: relevant tests pass, security requirements hold, and a reviewer judges the change ready. Failed runs stay in the numerator because the company paid for them.

Now watch what happens when quality moves. If an average attempt costs $1.86 and six in ten attempts are accepted, the AI cost per accepted change is $3.10. If only three in ten make it, the figure doubles to $6.20. The token price did not change. The success rate did.

Two cautions matter. First, inference is not the whole cost. When a run costs a few dollars, engineering time spent reviewing and fixing its output can outweigh the model bill. Some charges are not token charges at all: Copilot code review, for example, also consumes GitHub Actions minutes.

Second, count what you actually pay, not simply the list-price value of actual usage. On a pooled plan, a developer whose usage would be worth $60 at list prices can add $0 in marginal spend while the organization still has credits left.

None of this makes expensive runs bad by definition. A costly session that resolves a hard problem can deliver excellent business value. A cheap run that produces code nobody will accept is not efficient.

Who owns the bill?

AI spend is already landing on FinOps teams’ desks. In the State of FinOps 2026 survey of 1,192 practitioners, 98 percent said they manage AI spend, up from 63 percent in 2025. Seventy-eight percent of FinOps teams report to the CTO or CIO. The respondents are the FinOps community, so read those numbers as evidence of attention, not proof of organizational maturity.

For leadership teams, the practical decisions are straightforward—and they cut across finance and engineering.

Budget outcomes, not seats—CEO and CFO. Ask for one monthly number for a defined workflow: cost per accepted change. Budget in ranges, and plan against something like the 90th percentile rather than the average, because a bad run can be radically more expensive than a typical one. Fund a pilot before scaling licenses. Once the workflow has enough cost data and usage data, the same history can feed forecasting models that combine past AI spend and usage patterns with planned changes, giving finance a forward view of AI costs rather than a post-invoice explanation.

Make AI spend visible and owned—CIO and CTO. Tie each cost line item to a project, repository, and task type so cost tracking and cost reporting stay connected to the work that created the spend. Put one accountable owner around the table with engineering, finance, and the business. Claude Code can export token and cost metrics through OpenTelemetry, but its documentation calls the figures estimates, so a cost dashboard still needs to be reconciled with invoices.

Keep context on a leash—CIO and CTO. Limit what an agent may read and retain. Trim bulky tool output. Summarize or restart long sessions. Split large jobs into smaller ones. Then check whether cost per accepted change actually falls.

Build brakes into the loop—CIO and CTO. An alert after an expensive run is evidence; a limit that stops the loop is control. Cap cost or calls per task, retries, and parallel agents, then hand work to a person when a run crosses a threshold—twice the median cost for that task type, for example. GitHub already lets administrators set budgets by enterprise, cost center, and user, and cap spend when the pool runs out. The same telemetry can also automate anomaly detection and faster alerts, surfacing unusual jumps in usage or spend before they become a month-end surprise.

Test before announcing savings—everyone. For cost optimization, a smaller model may be enough for routine edits. A stronger model may earn its higher price on difficult changes. Tighter context may beat both. Compare options against the same acceptance criteria, and claim cost savings only after the numbers confirm them.

 

Back to the developer who typed “Fix the failing test.”

The organization should be able to see the entire run, failed attempts included; connect the bill to an accepted change; and know when further retries stopped being worth the money.

That is a better operating goal than buying fewer tokens.

This week, pick one recurring workflow. Define what “accepted” means. Start recording the cost of every run.

The token is the billing unit. The fixed bug is the reason to spend.

Sign up for the newsletter and other marketing communication

You may also find interesting:

Book a free 15-minute discovery call

Looking for support with your IT project?

Let’s talk to see how we can help.

The controllers of the personal data are companies of FABRITY Group (hereinafter referred to as “Fabrity”) with its mother company Fabrity SA seated in Warsaw, Poland, National Court Register number 0000059690; the data is processed for the purpose of marketing Fabrity’s products or services; the legal basis for processing is the controller's legitimate interest. Individuals whose data is processed have the following rights: access to the content of your data and the right to rectification, erasure, restriction of processing, the right to object if the processing of personal data is based on consent and the right to data portability. You also have a right to lodge a complaint with PUODO. Personal data in this form will be processed according to our privacy policy.

dormakaba 400
frontex 400
pepsico 400
bayer-logo-2
kisspng-carrefour-online-marketing-business-hypermarket-carrefour-5b3302807dc0f9.6236099615300696325151
ABB_logo
Fabrity
Privacy overview

Cookies are small text files that are stored on your device using the browser. They do no harm and do not allow any conclusions to be drawn about your identity. We use cookies to make our offer user-friendly. You can find more information under our data protection notice.