The AI Margin: Pricing the Token and the Task Inside TBM

Helping Businesses Thrive Through Exceptional Branding and Website Solutions Tailored to Achieve Growth.

Abstract

Internal AI platforms bill their business units at cost. Charge equals cost, so the difference is zero every period, for every workload, regardless of what the AI produced. The margin is not measured. It is set to zero by construction, and the finance system records what AI consumed without recording whether the consumption was worth paying for.

TBM (Technology Business Management) already defines profit margin as the difference between the price of a service delivered to business consumers and the fully-burdened cost of delivering it. Its rates management mechanism already permits a rate set away from cost, and the sanctioned purpose of that departure is shaping demand. What is absent is a rate set away from cost in order to value the service, declared against a named basis, with the resulting spread read as a margin on a unit of AI work.

Three bases are available for that declaration: cost, market, and value, the last anchored to the cost of the human work the AI displaced. The output is two spreads rather than a ratio. The sourcing margin, market price minus fully-loaded cost, is the observable floor. The value margin, value anchor minus fully-loaded cost, is the wider range and the location of the business case. Both are read per unit, per period, against a cost that includes failures and the cost of undoing them.

The mechanism adds nothing to the taxonomy. The allocated cost and the charged rate are already booked for every consumer. The margin is the difference between them, and three attributes on the existing tagging mechanism make it interpretable: the declared basis, the anchor, and whether the margin has left the cost base.

The method is worked on one customer-service resolution, then run again against an offshore wage geography where the value margin compresses to near zero and turns negative in the harder workload segment. Both readings are correct. They describe different operations.

The value basis carries a temporal limit. It works because a human still does the displaced work, and that human's cost is the yardstick. As AI takes over the work the yardstick disappears, and the market basis does not escape it, because vendors price their agents against the same labor cost. The measurement erases its own basis by succeeding.


1. Introduction: the resolution nobody can price


A shared support platform sits inside a large enterprise and answers customer questions for three business units. It runs on a large language model. Every month it resolves tens of thousands of customer issues. At the end of the month it sends each business unit a bill. It adds up everything the platform costs to run, divides by the number of resolutions, and charges each unit for what it consumed.

The bill is accurate. It is also silent on the one thing the business units want to know. It records what a resolution cost. It does not record whether the cost was worth paying.

Two prices for the same work sit in plain view, and the ledger ignores both. Intercom's Fin agent resolves a customer issue for $0.99. Salesforce charges $2.00 for an Agentforce conversation. A business unit could buy this work on the open market for roughly a dollar, or roughly two, depending on which vendor it called.

The two outside prices do not agree on the unit. One is priced per resolution. The other is priced per conversation, and a conversation is not a resolution, because not every conversation resolves. The market has not settled on what the unit of AI work is. The internal ledger has not noticed there is a question.

Behind both prices stands a third number. Before the platform existed, a human agent resolved these issues. What that agent cost the company, wage plus benefits plus the overhead that carries a seat, is what a resolution used to cost. It is the price the company was already paying for the same outcome. The ledger does not carry it either.

So the platform books AI at cost, and three numbers that would tell the company whether the cost is worth paying are left off the books:

- the price to buy the resolution from a vendor

- the price to sell it to another business

- the price the company used to pay a person to produce it

This is the ordinary condition of AI spend inside enterprises today. Internal AI platforms charge their business units at cost, and the finance system records what AI consumed. The AI cost management tools reviewed for this work all operate the same way (Section 7). The result is a technology spend that finance can measure precisely and cannot evaluate at all. Other technology spend does not sit in this position. A data center has a market rate to compare against. A SaaS seat has a list price. AI has vendor prices in public view and a labor cost it was bought to displace, and the ledger connects to neither.

Two facts explain why this matters now rather than later. AI spend is no longer marginal. The FinOps Foundation's 2026 practitioner survey reports that 98 percent of responding teams now manage AI spend, against 31 percent two years earlier (FinOps Foundation, State of FinOps 2026). And the value side is not landing. A 2025 MIT study of enterprise generative AI deployments reported that the large majority of pilots produced no measurable effect on profit and loss, a finding widely cited and fairly criticized as a survey rather than a ledger fact (MIT NANDA, The GenAI Divide, 2025). The spend is real, the value is asserted, and nothing in the accounting connects the two.

What already exists, and what it leaves open

None of this requires a new idea. Margin is old. Cost accounting has read the gap between a price and a cost for a century.

Technology Business Management, or TBM, is the standard enterprise framework for tracing technology cost from the general ledger through to the business units that consume it. It is maintained by the TBM Council and is in use across large enterprises and governments. It already carries cost to a consumer. It already defines a rate. It already defines profit margin, and it defines it as the difference between the price of a service delivered to business consumers and the fully-burdened cost of delivering it (TBM Council, Book of TBM, glossary). The definition has existed for years. It has never been computed for a unit of AI work.

The reason is structural rather than technical. The rate is set at cost, so the difference is zero, so there is nothing to read. The Book of TBM notes in passing that most organizations avoid the word price because it connotes a profit margin. The omission is not an oversight in the framework. It is a habit in the practice.

The work surrounding this problem sits at two poles, and neither closes it. On one pole, the supply side optimizes the spend. Recent token-economics research models the token as a factor of production and minimizes computational cost against a quality threshold (Chen et al., arXiv 2605.09104). On the other pole, the internal cost side records the spend. TBM's own AI cost guidance, the FinOps framework, the FOCUS billing-data standard, and the commercial AI chargeback tools attribute AI cost to a consumer on a cost-recovery basis. One pole makes the spend cheaper. The other makes it visible. Neither says what the spend is worth, or how the ledger would carry that.

The gap between those poles is where this work sits. The internal price of an AI unit is set against a declared basis, and the distance between that price and the cost is read as a margin, using the rate mechanism the framework already has. No new cost pool. No new object in the taxonomy.

The unit that carries the argument

The customer-service resolution is the unit. It runs through each basis in turn, and the same resolution returns in the sections that follow. It is chosen because it is one of the few AI workloads with public outside prices on both sides: what vendors charge to do the work, and what the displaced human labor costs.

Where a number is public it is named and sourced. The Intercom and Salesforce rates are published vendor prices, current as of mid-2026, in a market that reprices fast. The human-agent cost is built from published wage and benefits data and presented as a constructed range, not a quoted figure, with its derivation shown where the value basis is worked through. Where no honest anchor exists, the number is left unstated rather than invented.


2. From cost recovery to a declared basis


An internal platform that charges its business units at cost has already answered the pricing question. It has answered it with zero.

The arithmetic is not subtle. Charge equals cost. Margin is the difference between them. Set the charge to the cost and the difference is zero every month, in every business unit, for every workload, regardless of what the AI produced. The number is not measured. It is assumed.

This is why cost recovery feels safe. It cannot be wrong, because it never makes a claim. A cost-recovery bill says the platform spent this much and here is your share. It never says the share was worth having. The finance system records a complete and accurate account of AI spending and produces no evidence about whether the spending should continue.

For most technology this is fine. A shared network exists because the enterprise needs a network. Nobody asks whether the network cleared a hurdle rate, because the alternative is no network. AI is different in one specific way. Nearly every enterprise AI workload was justified by a comparison. It would be cheaper than the current way, or faster, or better. That comparison was made once, in a business case, before the system was built. Then the system went live and the comparison was never carried into the accounting. Cost recovery is where the comparison goes to die.

The basis is a decision, not a default

A price needs something to be set against. That something is the basis. Three are available for an internal AI unit: cost, market, and value. Section 3 runs each of them over a single resolution.

The point here is narrower. Every internal AI platform already has a basis. It is cost, and it was chosen by nobody. It arrived as the default behavior of the chargeback system, not as a governance decision. Naming the basis converts a technical default into a management choice with an owner and a rationale.

Two definitions this argument needs

Two terms carry weight from here forward. Both are used with a specific meaning, and both come from earlier work in this series.

Fully-loaded cost of an AI unit is the gross cost of producing it plus the downstream cost of undoing what it got wrong. The gross cost is the visible bill: inference charges, or the infrastructure and licensing that produce the tokens on owned hardware. The downstream cost is what the enterprise spends cleaning up incorrect output, which for a customer-service resolution includes the human agent who handles the escalation after the AI answers wrongly, the repeat contact, and the remediation when a wrong answer reaches a customer. That downstream cost is anti-value, and it is developed at length in The Anti-Value of AI. A system that resolves cheaply and fails often is not cheap. The gross bill will say otherwise.

Cost per successful task is the fully-loaded cost divided by the number of tasks that actually succeeded, not the number attempted. The metric is introduced in Decoding the Cost Genome of Agentic AI. The denominator is what matters. An AI system that attempts a thousand resolutions and completes six hundred has a cost per successful resolution roughly two-thirds higher than its cost per attempt. Attempts are what the vendor bills. Successes are what the business receives. Pricing against attempts prices the wrong thing.

Both definitions exist because AI spend behaves differently from the technology cost finance is used to reading. A server that runs produces what it was asked to produce. A model that runs produces an answer that may be wrong, and the wrongness has a price that lands somewhere else in the organization, on a different budget line, usually in the business unit rather than in technology.

What TBM already permits, and what it does not

Charging a business unit a rate is not new, and the framework is more specific than it is usually given credit for.

TBM defines Rate as the unit cost or set price for delivering a technology or service, and notes that on rare occasions it is called a price. It defines Rates Management as setting chargeable rates by understanding unit costs and establishing prices that sufficiently recover the average unit cost and costs in total. It defines Profit Margin, generically, as the difference between the price of a service delivered to business consumers and the fully-burdened cost of delivering that service (TBM Council, Book of TBM, glossary).

Two things follow, and the second one is the important one.

Cost recovery is written into the definition. Rates Management describes recovering unit cost and total cost. The default this section opened with is not an accident of implementation. It is in the definition of the mechanism.

And the mechanism already permits departure from cost. Rates Management explicitly allows discounts to encourage consumption and premiums to discourage it. So a rate set away from cost is not novel and is not unsanctioned. It is available today, and the sanctioned purpose is shaping demand.

That draws the boundary precisely. TBM permits a rate above or below cost in order to change how much of a service gets consumed. What it does not describe is a rate set away from cost in order to value the service, declared against a named basis, with the resulting spread read as a margin. The gap is not the rate. It is the purpose of the rate, the declaration of what it was set against, and the reading of the difference.

That is one decision: declare the basis. Everything that follows is a consequence of making it explicitly rather than accepting it silently.


3. The three bases


The three bases are easier to see on one unit than in the abstract. So take one customer-service resolution and price it three ways.

The resolution is the same in all three cases. A customer asks a question, the AI system answers it, and the issue closes without a human touching it. What changes is what the internal price is set against.

Cost basis

The cost basis charges what the resolution cost to produce.

For a platform running on a commercial model API, that is the inference charge for the tokens consumed by the resolution, plus the platform's share of orchestration, retrieval, storage, and the engineering that keeps it running. Add the downstream cost of the resolutions that went wrong, then divide by the resolutions that succeeded. The result is the cost per successful resolution.

Assume it comes to $0.60. This figure is illustrative and every number downstream of it inherits from it. The platform charges each business unit $0.60 per resolution. At year end the platform has recovered its costs exactly and shown no margin, which is what it was designed to do.

The cost basis answers one question: did the platform get its money back. That is a real question and worth answering. It is also the only question a cost basis can answer, and it says nothing about whether $0.60 was a good price.

This is the classic transfer-pricing default. Where no external market exists for an internally transferred good, the accounting literature holds that the transfer should happen at the supplying division's marginal cost, because any markup distorts the buying division's decisions (Hirshleifer 1956). The rule is sound. It also assumes no external market exists. For AI resolutions, one does.

Market basis

The market basis charges what the same resolution would cost from an outside vendor.

The market is public. Intercom's Fin agent charges $0.99 per resolution. Salesforce charges $2.00 per Agentforce conversation. Several other vendors price on resolutions or conversations under enterprise contracts that are not published, which is itself worth noting: much of this market prices per outcome, and most of it prices privately.

Charge the business unit $0.99 per resolution, and the internal platform is now priced against what the business unit's alternative would cost. The platform produces at $0.60 and charges $0.99. The difference is $0.39 per resolution, and that difference is information. It says the enterprise is producing this work for less than it would pay to buy it. Build was the right call, at this volume, at these prices, this month.

The market basis answers the build-or-buy question, and it answers it continuously rather than once in a business case. If the vendor price falls below the internal cost, the basis says so on the next bill.

Transfer-pricing theory prefers this basis where a competitive external market exists, because a market price is objective, it cannot be argued internally, and it forces the producing division to stay competitive (Hirshleifer 1956, and the standard cost-accounting treatment that followed). Whether the AI agent market is competitive enough to qualify is a fair challenge, and Section 10 takes it up.

Value basis

The value basis charges what the resolution is worth to the business unit consuming it.

Worth is the word that gets frameworks into trouble, so it needs an anchor. The anchor here is what the same outcome cost before AI produced it. A human agent resolved these issues. That agent's cost to the company, wage plus benefits plus the overhead that carries a seat, divided across the resolutions that agent handled, is a real number built from published data rather than an estimate of benefit. Section 4 derives it. Assume for now that it lands between $6.00 and $8.60 per resolution, and take the bottom of that range.

Charge the business unit $6.00 per resolution and the picture changes completely. The platform produces at $0.60 and the work is worth at least $6.00 to the business that receives it. The difference is $5.40 per resolution at the conservative end of the anchor, and $8.00 at the top. That is not a recovery number and not a sourcing number. It is the answer to whether the enterprise should be doing this work with AI at all.

The value basis comes from a different discipline. Value-based pricing sits in the marketing and pricing-strategy literature, where price is anchored to the buyer's next-best alternative plus whatever differentiates the offer (Nagle and Holden). It is not one of the transfer-pricing methods. The three classic transfer-pricing methods are market-based, cost-based, and negotiated. Value-based pricing is absent from that list, and pretending otherwise would misattribute the idea.

This matters for more than citation hygiene. The two literatures were built for different purposes. Transfer pricing exists to make divisions behave correctly toward each other. Value-based pricing exists to capture what a buyer will pay. Borrowing the second one for an internal ledger is a deliberate move and carries risks the transfer-pricing bases do not, which Section 10 sets out.

Outside vendors already reason this way. Several AI agent vendors explicitly benchmark their pricing against the cost of a human agent. The displaced-labor anchor is not a novel invention. What is missing is that the anchor exists in the vendor's pricing model and not in the buyer's ledger.

One unit per workload

Three prices for one resolution: $0.60 to produce, $0.99 to buy, $6.00 or more in displaced labor. Before generalizing, one structural rule has to be fixed, or the arithmetic stops working.

Price one unit per workload. A token or a task. Never both.

A task contains its tokens. A customer-service resolution consumes some number of input and output tokens, and those tokens have a cost that is already inside the cost of the resolution. Pricing the resolution and separately pricing its tokens counts the same money twice.

Which unit to choose depends on what the workload delivers to the business:

- Workloads with a discrete business outcome are priced per task. A resolution, a claim adjudicated, a document reviewed, a ticket triaged. The business consumes outcomes, so the outcome is the unit.

- Workloads that supply raw model capacity are priced per token. An internal inference platform serving many downstream applications has no single task to point at. It manufactures tokens, and tokens are what its consumers draw.

One consequence needs stating plainly. Whether the tokens were bought from a vendor or manufactured on owned infrastructure is a property of the cost, not a fourth basis. A bought token and a made token both produce a cost per token. That cost feeds the cost basis identically. 

Build-versus-buy shows up in the spread between the cost basis and the market basis, which is exactly where it belongs, and not as a separate way of pricing.

The same rule prevents a second error. For a given workload, the consumption cost of buying inference and the production cost of manufacturing tokens are alternatives, not addends. A platform either buys the capacity or makes it. Where it does both, the units are split by volume and priced separately, never summed into one blended figure that describes neither.

With the unit fixed and the bases named, the distances between the three prices become measurable. Section 4 measures them.


4. The spreads, with the worked example


Three prices sit on the same resolution: $0.60 to produce, $0.99 to buy, and $6.00 or more in displaced labor. The prices themselves are not the finding. The distances between them are.

Call those distances spreads. Each one is a separate signal and answers a separate question.

Sourcing margin is the market price minus the fully-loaded cost. It says whether producing this work internally beats buying it.

Value margin is the value anchor minus the fully-loaded cost. It says whether the work is worth doing at all.

They are read together and never summed. Adding them would double-count the cost, which appears in both.

The primary example: one customer-service resolution

The platform serves three business units. Assume it produces at a fully-loaded cost of $0.60 per successful resolution.

Sourcing margin. Intercom Fin charges $0.99 per resolution. Salesforce charges $2.00 per Agentforce conversation. The two published rates are not comparable until normalized, and once normalized they differ by a factor of three. Intercom's $0.99 is already per resolution. Salesforce's $2.00 per conversation, at a seventy percent resolution rate, is $2.86 per resolution. Take the lower of the two, because it produces the smaller and more defensible margin, and record that choice in the anchor reference. The sourcing margin is $0.99 minus $0.60, or $0.39 per resolution. At forty thousand resolutions a month the platform is producing about $15,600 a month of work below the price of buying it.

A market price stated in a different denominator is not a market price for your unit until it is converted, and the conversion carries an assumption about resolution rates that has to be stated alongside it. A three-times spread between the only two published rates is also a fact about the market rather than noise in it. Section 10 takes up what that means for the market basis.

Value margin. The anchor is what the same resolution cost when a human produced it. Building that number honestly takes four inputs:

- Median hourly wage for a customer service representative, from BLS occupational wage data, most recent release.

- Benefits load, from BLS Employer Costs for Employee Compensation, which puts total compensation at roughly 1.42 to 1.46 times the wage base.

- Overhead beyond benefits: supervision, facilities, telephony, training, quality assurance. Unlike wages and benefits, no public series covers these. They have to be identified from the enterprise's own cost model and stated as separate components rather than folded into a multiplier.

- Resolutions handled per agent per year.

Three of those four inputs are public. The fourth is not, and it is the one the answer is most sensitive to. Throughput per agent is operational and enterprise-specific, it is not published anywhere reliable, and halving it doubles the anchor. It is also a number every contact centre already holds, in the workforce management system that schedules the agents. So the instruction is to use your own figure rather than a borrowed one.

Run those together and a fully-burdened agent cost lands near $60,000 to $86,000 a year. At ten thousand contacts a year, an illustrative figure used here and not a benchmark, the cost per human-resolved contact falls between $6.00 and $8.60.

That is a range and it should be published as a range. Every enterprise's figure differs, because complexity, handle time, and utilization differ. The method is what transfers: take a published wage, apply a published benefits load, add separately identified overhead, divide by your own throughput.

The value margin is the anchor minus the cost: $5.40 at the bottom of the range, $8.00 at the top. At forty thousand resolutions a month, that is $216,000 to $320,000 a month of value margin, against $24,000 of production cost.

Reading the two together. The sourcing margin of $0.39 is small and certain. The value margin of $5.40 to $8.00 is large and soft. Both are true. They say different things. The first says building beat buying, modestly. The second says the workload should exist, decisively. A platform that reported only the first would look marginal. A platform that reported only the second would look unbelievable.

That is why the sourcing margin is the floor and the value margin is the range. The floor is defensible in a room full of skeptics because both of its inputs are observable. The range is where the actual business case lives.

The same resolution, a different wage geography

The anchor above uses a United States wage. Most enterprise customer service does not run on one. Contact work was offshored long before AI touched it, and a large share of the resolutions this paper describes are handled today from Manila, Bengaluru, Cairo, or Kraków.

Run the same method against that operation and one input moves. A fully burdened agent in a major offshore delivery center runs closer to $8,000 to $14,000 a year than to $60,000 to $86,000. At the same ten thousand contacts a year, the anchor lands between $0.80 and $1.40 per resolution.

Nothing else changes. The model charges the same. The orchestration costs the same. The failure rate is a property of the workload, not the wage. So the fully-loaded cost stays at $0.60 per successful resolution, and the vendor rate stays at $0.99.

The two spreads now say something different.


US operation

Offshore operation

Fully-loaded cost per success

$0.60

$0.60

Market price per resolution

$0.99

$0.99

Value anchor per resolution

$6.00 to $8.60

$0.80 to $1.40

Sourcing margin

$0.39

$0.39

Value margin

$5.40 to $8.00

$0.20 to $0.80

The sourcing margin is unchanged, because neither of its inputs depends on where the displaced work was done. Building still beats buying by the same $0.39. That number was never about labor.

The value margin has almost disappeared. At the bottom of the range it is $0.20 per resolution, which is less than the sourcing margin. That ordering is itself a finding. It says the vendor price of $0.99 sits above the cost of the human work it replaces, so an enterprise in this position that bought the resolution rather than building it would be paying more for automation than it was paying for people.

Push it down to the segment level and the two business units separate. Retail, at a true cost of $0.26 per success, carries a value margin of $0.54 to $1.14 and is comfortably worth running. Disputes, at $0.99, carries a value margin of negative $0.19 to positive $0.41. It straddles zero. In the US operation both segments were clearly worth running and the segment split was a fairness question. Here it is a continuation question, and the answer for one of the two segments depends entirely on where inside the anchor range the real number sits.

That is the metric working. Two enterprises, the same platform, the same model, the same failure rates, and two different correct answers about whether the automation is worth having. The first is told to keep going. The second is told to firm up its anchor before it decides anything, and to look hard at whether complex disputes belong in the automated path at all.

It also shows what the value basis actually measures. It is not a measure of how good the AI is. It is a measure of the distance between what the AI costs and what the displaced work cost, and an enterprise that already reduced the second number has less distance left to capture. Offshoring was the first pass at this. AI is the second, running against a baseline the first pass already lowered.

Margin computed is not margin realized

One question sits underneath every value margin and the paper has not answered it yet. The platform displaced $5.40 of human cost per resolution. Where did the money go.

Often nowhere. The agents whose cost anchors the number are still on payroll, still handling the contacts the AI escalates, still occupying seats the enterprise is still paying for. The value margin says what the displaced work was worth. It does not say the enterprise stopped paying for it.

This matters because a CFO will ask, and because the honest answer is uncomfortable. A value margin computed while the displaced capacity is still funded is a gross figure. It becomes money only where that capacity was actually removed, redeployed to work the enterprise wanted done and was not doing, or absorbed as growth the enterprise would otherwise have had to hire for. Absent one of those three, the margin is a measure of what could be captured rather than what was.

The distinction has a name in ordinary finance. Cost avoidance is not cost reduction. An enterprise that reports one as the other loses the argument the first time someone reconciles it against the general ledger.

So the value margin carries a third piece of information: whether it landed.

Realized. The displaced capacity left the cost base. Headcount reduced, contract volume cut, a seat licence not renewed. There is a corresponding line somewhere that got smaller.

Partially realized. Some capacity left and some did not, or capacity was redeployed rather than removed. Redeployment counts, but only where the receiving work has its own justification. Moving an agent to a queue nobody was asking for converts a labor cost into a different labor cost.

Unrealized. The capacity is still funded and still there. The margin is real as a measurement and has not touched the P&L.

None of these is a failure state. Unrealized is the normal condition in the first year of a deployment, and it is often the correct condition, because capacity is usually removed on a slower cycle than automation is deployed. What is not acceptable is reporting the figure without saying which of the three it is, because the same number means three different things.

The reporting rule follows. Report the value margin with its realization state attached, and never sum an unrealized margin into a savings total. The gross figure informs whether the workload should continue. The realized portion is the only part that belongs in a financial claim.

The secondary example: the self-hosted factory

A second platform manufactures tokens on owned GPU infrastructure and serves them to internal applications. There is no single business task to point at, so the unit is the token. The cost mechanics of that platform, absorption and yield, are worked through in Decoding the Token BOM of Enterprise AI.

This platform carries a sourcing margin only. Market price per million tokens minus fully-loaded manufactured cost per million tokens. No value margin, because the platform does not produce a business outcome. Its consumers do, and their outcomes get valued at their own level.

Two states are worth naming. When utilization is high and the hardware is well loaded, the manufactured cost per token falls below the commercial API price and the sourcing margin is rich. When utilization drops, the fixed cost of the hardware spreads across fewer tokens, the manufactured cost rises, and the margin goes tight or negative. Same infrastructure, same market price, opposite conclusion, and the only variable is load.

The cost side has to be computed against practical capacity rather than actual utilization, or the number lies in a specific direction. Allocating idle capacity cost into the unit cost drives the per-token price up exactly when volume is falling, which makes the platform look uncompetitive, which drives volume down further. Cost accounting has a name for this and a fix: measure the rate at practical capacity and report unused capacity as its own line rather than burying it in the unit (Cooper and Kaplan 1992). For AI infrastructure, where utilization swings hard, this is not a refinement. It decides whether the margin is real.

Failure cost, and who pays for it

The fully-loaded cost includes failures. Every attempt costs money whether or not it succeeds, and failures carry a downstream cost when someone has to undo them. Dividing that total by successes only is what makes the cost per successful task honest.

An example. One thousand attempts in a month, nine hundred succeed. The AI bill for successes is $162. The AI bill for failures is $40, because failed attempts tend to retry and run longer. The downstream cost of undoing those hundred failures, mostly human escalation, is $338. Fully-loaded cost is $540. Divided by nine hundred successes, $0.60 per successful resolution.

The vendor invoice for that month reads about $0.20 per attempt. The real figure is three times higher, and the difference is almost entirely the undo cost, which never appears on any AI invoice.

Now split that month across two business units with different workloads. Retail runs password resets. Disputes runs complex billing arguments where failures escalate to a specialist.

                                   Retail      Disputes

Attempts                          500         500

Successes                         480         420

Undo cost per failure             $1.50       $3.85

True fully-loaded cost            $124        $416

True cost per success             $0.26       $0.99

Charged at the blended $0.60      $288        $252

Retail pays $288 for work that cost $124. Disputes pays $252 for work that cost $416. About $164 a month moves from the reliable workload to the unreliable one, and neither bill records that it happened.

The incentive runs the wrong way. Disputes buys work at $0.60 that costs the enterprise $0.99, so it has every reason to send more hard cases. Retail pays a premium for having a workload the system handles well. A single blended rate across consumers with different failure profiles does not just misallocate cost. It rewards the behavior that generates the cost.

This is not a defect of the cost basis. A single market-basis rate across the same two consumers distributes the identical distortion, and hides it better, because a rate anchored to a published vendor price looks externally validated. The problem is the blending, not the basis.

The practical fix is to set the rate per consumer segment where failure rates differ materially, and to publish the failure rate alongside the rate so the cross-subsidy is visible rather than silent. The deeper question, how to fairly divide the cost of shared AI infrastructure among consumers who impose different loads on it, is a distinct problem in allocation theory and is not solved here.

There is a second place this cost hides. The gross AI spend sits on the platform's ledger. The undo cost does not. When the AI fails a dispute, the specialist who handles the escalation is already funded inside the business unit's labor budget. The cost is being paid today. It is simply not tagged as an AI cost, which is why the AI bill looks cheap and the business unit cannot explain why its headcount is not falling.

Bill on one basis, report on another

The value margin invites a fight. Charging a business unit $6.00 for a resolution that cost $0.60 to produce takes $5.40 out of that unit's budget and moves it to the platform, and the unit's leadership will contest it. They will be partly right to. Internal profits are artificial. A transfer price above cost distorts the buying unit's own numbers and can push it to curtail demand for work the enterprise wants it to consume. Transfer-pricing literature has documented these effects for decades, and Section 10 sets them out.

So separate the invoice from the report.

Bill on cost or market. Both have observable inputs, both survive audit, and neither transfers a large sum on the strength of a modeled assumption. A cost-basis bill recovers the platform. A market-basis bill also disciplines it.

Report the value margin as an overlay. Compute it, publish it, review it in the governance forum where investment decisions get made. Do not put it on the invoice.

This keeps the political argument off the monthly bill and keeps the analytical rigor in the review. It also protects the value margin from the pressure that would otherwise corrupt it. A number that moves budget gets negotiated. A number that informs a decision can stay honest.

One constraint modifies this for a large share of readers. Many chargeback IT departments operate under a mandate to recover one hundred percent of costs, no more and no less (TBM Council, Book of TBM). Under that mandate a market-basis invoice is not available either, because a market rate above cost over-recovers and a market rate below cost under-recovers. Section 9 gives both variants.


5. The metric: the AI margin and its spread


The metric is one line.

AI margin equals internal price minus fully-loaded cost, per unit of AI work.

The unit is a token or a successful task, fixed per workload as set out in Section 3. The internal price is set against a declared basis. The fully-loaded cost includes failures and the cost of undoing them. Nothing else enters.

It is stated as an amount per unit, not as a ratio, and it is reported as two figures rather than one:

- The floor. Sourcing margin, market price minus fully-loaded cost. A single number, both inputs observable.

- The range. Value margin, value anchor minus fully-loaded cost. A range, because the anchor is a range.

For the customer-service resolution, that reads: floor of $0.39 per resolution, range of $5.40 to $8.00. One line, two figures, and a reader can see immediately what is certain and what is estimated.

Why the spread and not the ratio

The obvious alternative is a ratio. Value divided by cost. For this resolution, $6.00 over $0.60, a ten-to-one return. It is a tempting headline and it should not be the centerpiece.

Three reasons.

A ratio hides the scale. Two workloads both returning ten to one look identical, and one might be producing $5.40 a unit across forty thousand units while the other produces four cents a unit across two hundred. The first funds a department. The second is a rounding error. The spread carries the money and the ratio throws it away.

A ratio becomes unstable as the cost falls. Inference prices have dropped steeply and continue to. As the denominator approaches zero the ratio approaches infinity, and a metric that trends toward infinity as the technology improves is not measuring the thing anyone cares about. The spread converges on the value anchor instead, which is the correct behavior. It says the enterprise now captures nearly the whole value of the displaced work.

And a ratio is already occupied. The FinOps framework carries a Cost Performance Indicator, borrowed from earned value management, that compares value delivered against cost spent and reads above one, at one, or below one. It also describes an AI breakeven measure comparing the cost of doing a function with AI against doing it another way, including with labor. That is the same conceptual comparison in ratio form, and it exists. Claiming it here would be claiming someone else's work.

The ratio survives as a derived figure, useful for one narrow job: comparing workloads of very different sizes across different enterprises, where absolute amounts are not comparable. It is a cross-system comparable, computed from the spread, never the headline.

Value is a distribution, and the metric reports its point

One complication has to be handled and then contained.

An AI system does not return the same value every time it runs. Sometimes the resolution is complete and the customer leaves satisfied. Sometimes it is technically closed and the customer calls back in a week. Sometimes it is wrong and creates work. The value of a single run is drawn from a distribution, not read off a constant.

The stochastic character of model output is not a new observation, and the academic work models it directly. Token-economics research treats output quality as a stochastic variable and optimizes cost against a quality threshold (Chen et al., arXiv 2605.09104). That is established, and the point is not claimed here.

What matters for a ledger is narrower. A ledger books an amount, not a distribution. So the metric reports the expected value per successful task, and the distribution shows up in two places rather than as a third figure.

The failure rate sits in the denominator. Tasks that did not succeed do not count as successes, and their cost stays in the numerator. That is the distribution's left tail, already priced.

The value anchor is published as a range. The range absorbs the variation in what a successful resolution is worth across different contact types.

On the resolution: the platform succeeds on ninety percent of attempts, the fully-loaded cost per success is $0.60, and the anchor spans $6.00 to $8.60. The metric reports a margin of $5.40 to $8.00 per successful resolution. Anyone who wants the variance can read the failure rate, which is published beside the rate.

The claim being made here is specific and limited. Not that AI output is stochastic, which is known. Not that quality should be optimized against cost, which is modeled elsewhere. Only that the resulting economics can be denominated in money, per unit, against a declared basis, and carried in the enterprise's own financial system rather than in a model.

What the metric is not

Three clarifications, because each one is a place a reader could file this incorrectly.

It is not ROI. Return on investment is computed once, over a project, against an invested amount, and it answers whether the investment should be made. The AI margin is computed every period, per unit, on running operations, and it answers what the running work is worth right now. A project can clear its ROI hurdle and still run at a negative margin per unit, which is precisely the condition worth catching.

It is not chargeback. Chargeback moves cost to the consumer who caused it. That is allocation, and it is settled practice. The margin is what appears once the charge is set against something other than cost. Chargeback answers who pays. The margin answers whether what they paid for was worth it.

It is not a valuation. The metric does not estimate what AI is worth to the enterprise. It reports the distance between a price and a cost on one unit of work, where the price is anchored to something observable. Every figure traces to a vendor rate, a wage table, or a ledger entry. Nothing is modeled forward.

One number per workload, per period

The reporting discipline is plain. Each workload publishes, each period:

- the unit, token or successful task

- the declared basis

- the fully-loaded cost per unit

- the internal price per unit

- the sourcing margin, as a single figure

- the value margin, as a range, or marked unestimated

- the realization state of the value margin: realized, partially realized, or unrealized

- the failure rate

Eight lines. Six come from data the enterprise already holds. The other two, the declared basis and the realization state, are governance judgments rather than measurements, and Section 6 puts both on the ledger.


6. Booking it in TBM


The metric has to live somewhere. This section puts it in the enterprise's existing cost model without adding anything to that model's structure.

The claim is deliberately small. No new cost pool. No new tower. No new object in the taxonomy. Not even a new definition, because TBM already carries one.

The definition already exists

TBM's glossary defines Profit Margin generically as the difference between the price of a service, including an IT service delivered to business consumers, and the fully-burdened cost of delivering that service (TBM Council, Book of TBM).

That is the AI margin. Price minus fully-burdened cost, on a service delivered to a consumer. The definition has been in the framework for years and has never been computed for a unit of AI work, for one reason: the price has been set at the cost, so the difference is zero, so there has been nothing to report.

What this section adds is not a definition and not an object. It is the instruction that makes the existing definition produce a non-zero number, and the two attributes that make that number interpretable.

How the cost gets there

TBM traces technology cost through four layers. Cost pools hold the financial inputs as they arrive from the general ledger: labor, hardware, software, outside services, facilities. Technology resource towers group those inputs into the technical capabilities they fund, such as compute, storage, and network. Technology solutions are what the enterprise actually delivers, the applications and services. Consumers are the business units and functions that use them (TBM Taxonomy 5.0.1). Cost moves upward through these layers by allocation, from the ledger entry to the business unit that consumed it.

An AI platform maps onto this without difficulty. Model API spend and GPU infrastructure land in the cost pools. They aggregate into the compute and platform towers. The support platform is a solution. The three business units are consumers. The allocation carries the platform's cost to those three consumers, and the monthly bill is the result.

That flow is already in production at enterprises running TBM. Nothing in this section changes it.

Margin is a view, not an object

Here is the whole mechanism.

The allocation produces a cost per consumer. Rates Management produces a rate the consumer is billed. Two numbers therefore already exist in the model, for every consumer, every period: the allocated cost and the charged rate.

The margin is the difference between them, which is what TBM's own glossary says a profit margin is.

It requires no new object because both operands are already booked. A rate set at cost produces a difference of zero, which is why the margin has been invisible rather than absent. Set the rate against a declared basis other than cost, and the difference becomes a number with meaning. The reporting layer subtracts one existing field from another.

Rates Management already permits a rate away from cost. It allows discounts to encourage consumption and premiums to discourage it. So the mechanical capability is present and sanctioned, and its sanctioned purpose is demand shaping.

That is also why the basis attribute is not optional. A rate set above cost to suppress demand and a rate set above cost to reflect market value look identical on the invoice. They produce the same number and mean opposite things. Without a declared basis, the spread cannot be read at all, because nobody can tell which of the two it is.

The three attributes

Three attributes travel with the rate.

Pricing basis. Which basis the rate was set against: cost, market, value, or demand shaping. Without it, no consumer of the report can interpret the gap.

Value anchor reference. What the price was anchored to, identified specifically enough to audit. Not market rate but the vendor, the published price, the unit, and the date. Not human cost but the wage source, the load factors applied, and the throughput assumption. The anchor is the part most likely to be challenged, so it carries its provenance.

Realization state. Whether the value margin has left the cost base: realized, partially realized, or unrealized. This is the attribute a finance reader will look for first, because it separates a measurement from a financial claim. It is set by the same function that owns the capacity, not by the platform, for the same reason the basis is not set by the platform.

These three ride the tagging mechanism the taxonomy already provides. TBM 5.0.1 supports optional metadata tags on towers and solutions for reporting, modeling, and filtering, with published examples including an AI-enabled flag and a project identifier. These two attributes are tags of the same kind. They are enterprise-populated, they do not alter the taxonomy, and they need no ratification to start using.

The practical consequence is that an enterprise can begin next quarter. Declare the basis, tag it, tag the anchor, and let the reporting layer subtract. Nothing waits on a standards body.

Where the failure rate sits

One more field belongs beside the rate, for the reason Section 4 gave. Two consumers with different failure profiles charged one blended rate are cross-subsidizing each other silently, and the direction of the subsidy runs from the reliable workload to the unreliable one.

Publishing the failure rate alongside the rate makes the subsidy visible. Setting the rate per consumer segment where failure rates differ materially removes it. Both are reporting choices, not structural ones, and both sit at the same layer as the two attributes.

How this differs from the existing AI guidance

The TBM Council has published guidance on AI value realization, covering cost pools and taxonomy mapping for simple and complex AI architectures, showback and chargeback, total cost of ownership forecasting, cost optimization levers, and AI financial risks (TBM for AI Value Realization, TBM Council, 2024).

Its consumption treatment stops at allocation and cost recovery. It sets no per-unit rate and names no basis. The closest it comes to a unit is a recommendation to set token budgets and usage quotas by department, which is budgeting rather than rating. It also states plainly that implementing cost recovery mechanisms requires additional tools, policies, and processes beyond TBM itself.

That guidance and this work sit on the same rail and stop at different points. The guidance establishes allocation and recovery. This work takes the rate that recovery implies and asks what it should be set against, then reads the resulting difference. Recovery is the starting condition here, not the conclusion.

The relationship is extension, not correction. The rate mechanism is the one the framework already describes. What is added is the basis declaration on top of it.

One transfer price, and what happens next

Everything above assumes a single transfer: one platform charging one business unit for one unit of work. That assumption holds for the majority of enterprise AI deployments today, and it is what makes the mechanism simple enough to implement immediately.

It will not hold indefinitely. Agents call other agents. A resolution agent invokes a retrieval agent, which invokes a summarization agent, which calls the model. Each of those hops is a consumption event, and if each one carries an internal price, the cost flows sideways and sometimes in loops rather than upward through the layers. TBM's allocation is one-way by design. A priced unit that circulates does not fit that shape.

That is a real problem and it is out of scope here. It is also why this work comes first. Pricing a unit against a basis is the input a chained model needs. Build the unit once, correctly, and the chain problem becomes tractable. Attempt the chain first and there is no priced unit to flow.


7. Where this sits against existing work


Seven bodies of work touch this problem. Each contributes something, and each stops before the specific move made here.

Transfer pricing

The cost and market bases come from transfer pricing, the branch of management accounting governing how one division charges another.

The foundational rule is Hirshleifer's. Where a competitive external market exists for the transferred good, the transfer price should be the market price. Where none exists, it should be the supplying division's marginal cost, because a markup above cost distorts the buying division's decisions and can reduce total firm profit (Hirshleifer 1956). The standard methods that followed are three: market-based, cost-based, and negotiated (Horngren, Datar and Rajan; Anthony and Govindarajan).

Two things follow. Cost and market are not inventions here, they are the two canonical bases applied to a new unit. And the existing rule already argues for the market basis when a market exists, which is the position an AI platform is now in and was not five years ago.

Value-based pricing is not among the three methods. The third is negotiated. Nothing in the transfer-pricing literature supports a value basis, and this work does not claim otherwise.

Value-based pricing

The value basis comes from a different literature. Value-based pricing sits in marketing and pricing strategy, where price is anchored to the buyer's next-best alternative plus quantified differentiation (Nagle and Holden). The reference value is the price of the alternative the buyer would otherwise choose.

Applied internally, the buyer is a business unit and the next-best alternative is the way the work was done before AI, which for the customer-service resolution is a human agent. The anchor is that agent's cost.

Two literatures, kept apart deliberately. Transfer pricing exists to make divisions behave correctly toward one another. Value-based pricing exists to capture what a buyer will pay. Borrowing the second for an internal ledger is a deliberate crossing, and it brings risks the first two bases do not, set out in Section 10.

TBM

TBM supplies the rail and, as Section 6 sets out, the definition. It carries technology cost from the general ledger to the consuming business units, it defines Rate and Rates Management, and it defines Profit Margin as price minus fully-burdened cost on a service delivered to consumers.

This is worth stating precisely, because it narrows what is being claimed. Margin is not new to TBM. Rates set away from cost are not new to TBM either, and they are explicitly sanctioned for demand shaping through discounts and premiums. Three things are absent: a rate set away from cost for the purpose of valuation rather than demand management, a declared basis recorded against that rate, and the resulting spread read as a margin on a unit of AI work.

The Council's AI value realization guidance is the direct extension target. It establishes allocation and cost recovery for AI consumption and does not set a per-unit rate or name a basis (TBM for AI Value Realization, 2024).

FinOps

The FinOps framework is the closest adjacent practice, and it holds the nearest prior art to this metric.

Its Quantify Business Value domain contains a Unit Economics capability that ties technology spend to a business denominator, with documented examples including cost per transaction, cost to serve, and cost per case resolved. That is the cost-side unit this work builds on. The same domain carries a Cost Performance Indicator, adapted from earned value management, that compares value delivered against cost spent. It also describes an AI breakeven measure comparing the cost of performing a function with AI against performing it another way, including with labor.

That last item is the conceptual neighbor of the value basis and deserves to be stated plainly rather than buried. The breakeven measure already compares AI cost against labor cost. What differs is not the comparison but where it lives, and the difference is durability rather than sophistication.

A rate is carried per consumer, per period, inside the system finance already reconciles. It accrues a time series without anyone deciding to keep one. It attaches to a named consumer's bill, which means it has an owner who will contest it if it is wrong, and it inherits the controls that apply to everything else in that ledger. A KPI sits in a dashboard. It is computed by whoever was asked to compute it, on inputs chosen at the time, and it is recomputed when someone asks again. One survives an audit, a tooling migration, and a change of leadership. The other usually does not survive the third of those.

That is the argument for booking rather than indicating. The comparison is not claimed here. The booking is.

FOCUS, the billing-data specification, is treated in Section 8.

The token-economics literature

Recent academic work models the economics of tokens directly. A 2026 survey conceptualizes tokens as production factors, exchange media, and units of account, and works the problem across single-agent, multi-agent, and ecosystem levels using firm theory and transaction-cost theory, treating output quality as stochastic and optimizing cost against a quality threshold (Chen et al., arXiv 2605.09104).

That is the supply side. It asks how to produce tokens efficiently under a budget. This work is its counterpart on the demand and governance side, asking what the produced work is worth to the business consuming it and how the finance system should carry that. The two do not overlap. The survey does not address enterprise financial governance, internal chargeback, or a margin booked in a corporate ledger.

Adjacent to it sits seller-side pricing research on how AI providers should price their models to customers, including markup and tariff design. That is external pricing by a vendor to a market. This is internal pricing by a platform to its own business units, under management control rather than commercial negotiation. Different buyer, different purpose, different constraint.

Commercial tooling, and the vendor pricing rationale

Cloud cost management platforms, LLM gateways, and open-source metering services reviewed for this work in mid-2026 attribute AI spend to consumers on a cost-recovery basis. They allocate AI and cloud cost to teams, features, and customers, attribute token consumption, and enforce budgets by team. Some allow a custom unit price to be applied to token usage, used to monitor erosion against external vendor pricing rather than to set an internal basis and read a margin. None of the tools reviewed sets a non-cost internal price against a declared basis and books the difference.

Separately, several AI agent vendors benchmark their own prices against human labor cost, stating openly that agents should be priced against what a human agent costs. The displaced-labor anchor is therefore not novel as a pricing rationale. It is in commercial use. What is different is location. The anchor lives in the vendor's pricing model, facing the customer. It does not live in the buyer's ledger, facing the buyer's own decisions. Moving it inside, declaring it, and booking the resulting spread is the contribution.

Tax transfer pricing

Tax authorities apply cost, market, and margin methods to AI-related intangibles and intercompany transactions, under the arm's length principle. That work is cross-border tax compliance. Its purpose is to allocate taxable profit between jurisdictions and satisfy a regulator. It has no bearing on internal management control, and a price built for tax compliance is a poor instrument for steering internal decisions. The two are kept separate here.

Three things this is not

Not generic chargeback. Chargeback allocates cost to the consumer who caused it and is settled practice. The margin is what becomes visible only once the charge is set against something other than cost.

Not ROI. ROI is computed once, over a project, against an invested amount. The margin is computed every period, per unit, on running operations.

Not one-time build-versus-buy TCO. A TCO comparison is a decision input evaluated before commitment. The sourcing margin is the same comparison converted into a standing periodic reading that changes as vendor prices and internal costs move.


8. The pricing bridge


The mechanism in Section 6 needs three things in the data: a rate, a declared basis, and a count of what actually succeeded. The obvious question is whether the standards already carry them, so an enterprise could inherit rather than populate. They do not, and the gaps are specific.

FOCUS is the FinOps Foundation's specification for technology billing data, giving cost and usage from every provider a common schema. Version 1.4, ratified in June 2026, added invoice and billing-period datasets and expanded commitment detail. The cost side is thorough: list, contracted, effective, and billed cost, alongside pricing and consumed quantities. An enterprise can source the cost operand of every figure in Section 4 from FOCUS data.

Four things it does not carry.

No standardized AI unit. The specification does support pricing in non-monetary units such as credits or tokens, introduced in version 1.2. That support is for virtual currency: prepaid scrip a vendor sells and the customer draws down, with the stated use case being credit and token purchase patterns, burn-down tracking, and forecasting exhaustion. It is a currency denomination, not a unit of AI work. An inference token would land instead in the generic consumed-quantity and consumed-unit fields, where nothing defines it. No normative definition separates input tokens from output tokens or cached tokens, and nothing obliges two providers to mean the same thing by the word. The word appears in the specification. The unit does not.

No internal price basis. FOCUS describes what a provider charged a customer. There is no field for what an internal platform charges its own business units, and none declaring what that charge was set against. This is by design and not an oversight. FOCUS standardizes supplier billing, and an internal transfer price is not supplier billing.

No marked-up internal rate. The specification models cost as billed and discounted, not as internally re-priced. So the two operands of the margin are split. The cost operand is standard and portable. The price operand does not exist in the schema at all.

No output reliability measure. Nothing records whether the consumed units produced a successful outcome. A thousand tokens spent on a failed resolution and a thousand spent on a successful one are identical rows. Cost per successful task cannot be computed from billing data, because billing data has no concept of success.

OpenTelemetry covers the side FOCUS cannot. Its semantic conventions for generative AI capture model calls, token counts, and operation outcomes at the trace level, which is where success and failure are actually observable. That makes it the natural source for the denominator. It is not a financial standard and does not attempt to be. It carries no rate, no basis, and no price.

Between the two, the enterprise has cost and it has telemetry. What neither supplies is the join: a declared basis, an internal rate, and a per-task success signal reconciled to the same unit. That join is enterprise work today.

This is a timing statement, not a complaint. Both standards are moving fast and the gap is closing from both directions. FOCUS has shipped four versions in three years and names AI spend as target scope. The generative AI observability conventions are actively developing. A basis field or an internal-rate extension is a reasonable thing to expect eventually, and the enterprise that has been tagging its own from the start will map to it in an afternoon.

Which is the practical point. The two attributes in Section 6 are enterprise-populated because they have to be, not because the standards should be worked around. Populate them locally now, using the tagging mechanism the taxonomy already provides, and keep the definitions clean enough to map when the standards arrive.


9. What leaders should do


Seven moves. Each one is available now, with the systems already in place.

Declare the basis, and name the owner

Pick the basis for each AI workload deliberately: cost, market, or value. Write it down. Where a rate is set away from cost to shape demand rather than to reflect value, record that too, because on the invoice the two are indistinguishable.

Name the owner rather than leaving it to whoever moves first. The TBM office owns the basis declaration, because the basis is a governance decision about how technology cost is represented to the business, and arbitrating that is what the TBM office is for. The FinOps practice owns the cost operand and the failure rate, because those are measurement. The platform owner owns neither, and the declaration exists partly to constrain them, because they are the party with an incentive to pick the flattering basis.

Where a credible external price exists for the same unit, the market basis is usually the better default. The accounting literature has held for seventy years that a competitive external price beats an internal cost figure, because it is objective, it cannot be negotiated internally, and it keeps the producing platform honest about whether it should still be producing. Enterprise AI now has such prices for a growing number of workloads. That was not true three years ago.

Bill on one basis, report on the others

Two variants, depending on how the technology function is funded.

Where rate-setting is discretionary, bill on cost or market and report the value margin as an overlay in the governance forum. Section 4 gives the reasoning.

Where the mandate is to recover one hundred percent of cost, no more and no less, bill on cost only. A market rate above cost over-recovers and a market rate below cost under-recovers, so neither is available on the invoice. Report both the market anchor and the value anchor as overlays instead. The margin is then entirely an analytical reading, which loses the pricing discipline a market-basis invoice would impose and keeps every diagnostic the metric provides. Most enterprises will be in this second position.

Set the rate annually, re-derive the anchor quarterly

Mature TBM practices set rates once a year and hold them, because business units need predictable internal prices to budget against. The AI agent market reprices in months. Those two facts pull against each other and both are right.

Resolve it by separating the billed rate from the reported anchor. Hold the billed rate for the fiscal year. Re-derive the market anchor quarterly and report the spread against the held rate. The invoice stays predictable. The signal stays current. If the quarterly re-derivation shows the market has moved past the held rate materially, that is the input to next year's rate-setting, and in a severe case the trigger for an off-cycle change.

Carry the three attributes from day one

Tag every AI rate with the pricing basis and the value anchor reference. Do it when the first rate is set, not after the first argument about what a margin means.

The anchor reference needs enough detail to audit: the vendor, the published price, the unit, and the date for a market anchor, and the wage source, the load factors, and the throughput assumption for a value anchor. An anchor without provenance will be challenged in its first review and will not survive.

Retrofitting is expensive. Historical rates cannot be re-tagged with a basis nobody recorded, which means the first year of data is uninterpretable. The cost of doing it from the start is two fields.

Report the floor and the range

Publish the sourcing margin as a single figure, the value margin as a range, and the failure rate beside both. Every period, per workload.

Reporting only the floor makes a valuable workload look marginal. Reporting only the range invites the reader to distrust the whole thing. Reporting the margin without the failure rate leaves the reader unable to tell whether the cost figure reflects a system that works or one that is quietly expensive.

Act on the reading

A number nobody acts on is a dashboard. These are the readings and what each one means.

Sourcing margin persistently negative. The market produces this work more cheaply than the platform does. Buy rather than build, and treat the internal platform's remaining justification as something other than cost.

Sourcing margin thin and volatile. Hold. Re-derive quarterly. Do not restructure on a single period's reading in a market that reprices this fast.

Value margin negative across the whole range. The work costs the enterprise more than the outcome is worth to it. Stop the workload, or change it until the cost falls below the anchor.

Value margin negative only at the bottom of the range. This is usually a statement about the anchor rather than the workload. A range wide enough to straddle zero has a soft input in it, most often throughput. Rebuild the anchor before making any decision on it.

Value margin unestimated. The workload cannot currently be evaluated. That is itself a finding and should be reported as one, not left blank. An unevaluable workload running at scale is a governance exposure regardless of what it costs.

One guardrail. These are operating decisions, read every period on running work. They are not investment decisions and they do not replace the business case. A workload can be worth continuing this quarter and worth replacing next year, and the metric is designed to say which without being asked to say both.

Name the dysfunctions before they arrive

Internal pricing above cost has known failure modes. They are documented, they are old, and they will show up.

Demand suppression. A business unit charged above cost for AI may cut its usage below what the enterprise wants, because the internal price makes its own numbers look worse. The work the enterprise wanted done does not get done. Watch consumption after any basis change and treat a sharp drop as a signal about the price, not about demand.

Basis gaming. Once the basis is known to affect reported margin, there is pressure to pick the basis that flatters. A platform that switches to a value basis in a strong quarter and back to cost in a weak one is not measuring anything. Fix the basis for a defined period and require a documented case to change it.

Sandbagging the anchor. The value anchor is the softest input, and whoever benefits from a large margin has an incentive to inflate it. Bias the other way. Conservative anchoring understates the margin, which is recoverable. Over-attribution discredits the number permanently and takes the rest of the reporting with it.

Political resistance. A business unit that sees its costs rise will contest the basis, and it will have a real argument, because internal margin is an accounting construct rather than money earned. The invoice-and-report split defuses most of this. What remains is handled by naming the purpose plainly: the number exists to inform whether the work should continue, not to move budget between organizations.

None of these are reasons to avoid setting a basis. They are the reasons cost recovery became the default, and they are worth accepting, because the alternative is a technology spend that no one can evaluate.


10. Honest limits, and the closing window


Several things about this approach are weaker than the mechanism makes them look. They are worth stating before someone else states them.

The anchors are local

Every figure in Section 4 is enterprise-specific. Contact complexity, handle time, agent utilization, wage geography, and failure rates all vary enough that two companies running the same vendor model on the same workload will produce different numbers, correctly.

That is not a defect in the method. It is what an internal price is. But it means the numbers here are not benchmarks and should not be quoted as such. The method transfers. The figures do not.

The constants are not public

The hardest input is the value anchor, and the pieces that would make it precise are the pieces nobody publishes.

Wage data is public. Benefits loads are public. What is not public is the fully-burdened cost per resolved contact in a comparable operation, because contact centers do not publish it and the figures that circulate are vendor marketing. Throughput per agent is not public either, though every operation holds its own. The number built in Section 4 is a construction from public inputs plus one operational input, not a benchmark drawn from anywhere.

Report a range. Show every input. Let the reader substitute their own. A single confident figure would be more useful and less true.

The market basis assumes a market

Transfer-pricing theory prefers a market price where a competitive external market exists. The AI agent market of 2026 is only partly that.

Two vendors publish per-unit rates. Most do not, pricing instead through custom enterprise contracts with no public rate at all. The two published prices are not stated in the same unit, and once normalized they differ by a factor of three, which is a wide band to call a market price. And the market is repricing constantly, with at least one major vendor having changed its pricing model three times in eighteen months.

A market basis built on two published rates is a thinner foundation than the theory assumes. It is still better than no external reference, and it improves as more vendors publish. But an enterprise should know that its market anchor rests on a small number of quotes in a volatile market, and it should re-derive the rate on a schedule rather than setting it once.

The value basis is borrowed across a boundary

Cost and market bases come from transfer pricing, a literature built to make divisions behave correctly toward one another. The value basis comes from value-based pricing, a literature built to capture what an external buyer will pay. Moving it inside is a deliberate crossing, and the destination was not designed for it.

The consequences are the dysfunctions in Section 9, and they are real rather than theoretical. This is the main reason the value margin belongs in the report and not on the invoice. It is a strong analytical instrument and a poor billing instrument, and treating it as the second would damage both.

It measures displacement, not creation

The value anchor is the cost of the work being replaced. That works for AI applied to work a human was already doing, which is most enterprise AI today.

It does not work for AI doing something nobody was doing before. A system that reviews every contract instead of a sample has no displaced cost to anchor against, because the unreviewed contracts had no cost. A system producing genuinely new capability sits outside this method entirely.

Where that happens, report the sourcing margin and mark the value unestimated. An unestimated value margin costs nothing. An invented one costs the credibility of every other number on the page. The method covers displacement. It does not cover creation, and stretching it to would mean inventing the anchor.

The closing window

The deepest limit is not methodological. It is temporal.

The value basis works because a human still does the displaced work. That human's cost is the yardstick. The whole measurement depends on being able to point at what the outcome cost before AI produced it.

As AI takes over the work, the yardstick disappears.

When the last human resolution agent is reassigned, the wage line that anchored the value of a resolution stops being a live figure and becomes a historical one. It ages. Wages move, roles change, and the operation that produced the throughput assumption no longer exists. Within a few years the anchor is an artifact rather than a measurement, and there is no way to rebuild it, because nobody is doing the work anymore.

The measurement erases its own basis by succeeding.

The obvious escape is to switch to the market basis and let vendor prices carry the reference. That escape does not hold, and the reason is in Section 7. Vendors price their agents against the cost of a human agent. That is the stated rationale, and it is how a per-resolution price gets set in the first place. So when the buyer-side anchor disappears, the vendor-side anchor is resting on the same vanished labor cost, and the buyer is pricing against a vendor who is pricing against a workforce neither party can observe any longer. The reference does not survive by moving outside the firm. It becomes circular, with a fossil at the centre.

This is not a reason to avoid the value basis. It is a reason to establish it now, while the human operation still exists and its costs are still being incurred and recorded. An enterprise that measures the value of its AI resolutions in 2026, against a human operation still running, has a defensible anchor and a baseline it can carry forward. An enterprise that waits until the transition is complete has nothing to anchor to and will be reduced to asserting value it cannot demonstrate.

The practical instruction is narrow. Capture the anchor while the displaced operation is still live. Record it with its inputs and its date. Carry it forward as a stated historical baseline, marked as such, rather than pretending it is current. And accept that the value basis has a shelf life measured in a few years for any workload where displacement is running to completion.

Which is the uncomfortable part. The instrument that shows AI is worth what it costs is the same instrument AI is in the process of destroying. Enterprises have a window in which the question can still be answered with evidence. It is open now. It closes workload by workload, quietly, as the last person who used to do the work moves on to something else.


11. Conclusion


An internal AI platform bills its business units at cost. The bill is accurate, and it settles the pricing question by default, with an answer nobody chose. Margin is zero because the charge equals the cost, every month, regardless of what the AI produced.

Setting the charge against something else changes what the ledger can say.

Against fully-loaded cost, the platform learns whether it recovered its money. Against a published vendor rate, it learns whether building beat buying, and it learns it again every period as prices move. Against the cost of the human work that used to produce the outcome, it learns whether the work should be done at all.

The distances between those prices are the finding. The sourcing margin is the floor, built from two observable numbers. The value margin is the range, built from published inputs and reported with them shown. Both are read per unit and per period, on the workload's own terms.

None of this requires a new object, or even a new definition. TBM defines profit margin as the price of a service to business consumers minus the fully-burdened cost of delivering it. The cost is already allocated. The rate is already charged. The difference has been zero only because the rate was set at the cost, and rates set away from cost have been reserved for shaping demand rather than reading value. Two attributes, a declared basis and an anchor reference, carried on the tagging mechanism the taxonomy already provides, are enough to make the difference interpretable. The instruction is a governance decision, not a schema change, which means an enterprise can start next quarter.

Two bodies of work sit on either side of this and neither closes it. The supply side is optimizing the spend, modeling the token as a factor of production and driving computational cost down against a quality threshold. The cost-management side is recording the spend, with the frameworks, the billing standards, and the commercial tools attributing AI cost to consumers on a cost-recovery basis. Cheaper on one side, visible on the other. Neither says what the spend is worth.

That question is answerable today with evidence, and it is answerable because a human still does the work in most places where AI is being deployed. The human's cost is the yardstick. It exists in payroll systems, it is being paid this month, and it can be tied to a unit of output.

That will not hold. The yardstick is what AI is displacing. Every workload where the transition runs to completion loses the reference against which its value could have been measured, and the loss is quiet. No system flags it. The wage line simply stops being current, and the last honest anchor becomes a number in an old file.

So the window is not a rhetorical device. It is the interval between AI being deployed and the displaced operation being gone, and it is different for every workload and closing in all of them. Enterprises that measure inside it will carry a defensible baseline forward. Enterprises that wait will assert value they cannot demonstrate, against a comparison that no longer exists.

The mechanism is available now. The anchor is available now. Only one of those stays true.


Glossary

AI margin. The internal price of a unit of AI work minus its fully-loaded cost, expressed per unit and per period. A specialisation of TBM's Profit Margin, applied to a token or a successful task.

Anti-value. Cost created by an AI system's incorrect output, incurred downstream of the system that produced it. Developed in The Anti-Value of AI.

Basis. What an internal price is set against: cost, market, or value. Declared per workload.

Cost basis. An internal price set at the fully-loaded cost of producing the unit.

Cost per successful task. Fully-loaded cost divided by the number of tasks that succeeded, not the number attempted. Introduced in Decoding the Cost Genome of Agentic AI.

Fully-loaded cost. The gross cost of producing a unit of AI work plus the downstream cost of undoing incorrect output.

Market basis. An internal price set at the published rate an outside vendor charges for the same unit.

Practical capacity. The output an asset can sustain under normal operating conditions, used as the denominator for unit cost so that idle capacity is reported separately rather than inflating the unit.

Profit Margin (TBM). Defined in TBM's glossary, generically, as the difference between the price of a service delivered to business consumers and the fully-burdened cost of delivering that service.

Rate (TBM). Defined in TBM's glossary as the unit cost or set price for delivering a technology or service.

Rates Management (TBM). Defined in TBM's glossary as setting chargeable rates by understanding unit costs and establishing prices that recover average unit cost and total cost. Permits discounts to encourage consumption and premiums to discourage it.

Realization state. Whether a computed value margin has left the enterprise's cost base. Realized where the displaced capacity was removed or redeployed to justified work, partially realized where some was, unrealized where the capacity is still funded.

Sourcing margin. Market price minus fully-loaded cost. The observable floor.

Transfer price. The price one internal division charges another for a good or service.

Value anchor. The observable substitution price a value basis is set against, normally the fully-burdened cost of the displaced human work.

Value basis. An internal price set at the value of the unit to the consuming business, anchored to the value anchor.

Value margin. Value anchor minus fully-loaded cost, reported as a range.


Companion papers

Decoding the Cost Genome of Agentic AI. Consumption cost and the seven-layer agent transaction cost framework. Source of cost per successful task.

Decoding the Token BOM of Enterprise AI. Production cost of self-hosted token manufacturing, absorption costing and yield.

The Anti-Value of AI. Harm cost, the disposition flag, and the definition of fully-loaded cost as gross spend plus downstream undo.


References

  • Hirshleifer, J. (1956). On the Economics of Transfer Pricing. The Journal of Business, 29, 172-184.

  • Horngren, Datar and Rajan. Cost Accounting: A Managerial Emphasis. For the three transfer-pricing methods.

  • Anthony and Govindarajan. Management Control Systems.

  • Nagle, T. and Holden, R. The Strategy and Tactics of Pricing.

  • Cooper, R. and Kaplan, R. S. (1992). Activity-Based Systems: Measuring the Costs of Resource Usage. Accounting Horizons.

  • TBM Council. TBM Taxonomy Version 5.0.1, July 2025.

  • TBM Council. Book of TBM. 

  • TBM Council. TBM for AI Value Realization, 2024.

  • FinOps Foundation. FinOps Framework: Quantify Business Value, Unit Economics, Planning and Estimating.

  • FinOps Foundation. FOCUS Specification v1.4, June 2026.

  • FinOps Foundation. State of FinOps 2026.

  • Chen, Y. et al. (2026). Token Economics for LLM Agents: A Dual-View Study from Computing and Economics. arXiv:2605.09104.

  • MIT NANDA. The GenAI Divide: State of AI in Business 2025.

  • US Bureau of Labor Statistics. Occupational Employment and Wage Statistics, SOC 43-4051, most recent release.

  • US Bureau of Labor Statistics. Employer Costs for Employee Compensation, most recent release.

  • Intercom. Fin pricing, retrieved June 2026.

  • Salesforce. Agentforce pricing, retrieved June 2026.

  • OpenTelemetry. Semantic conventions for generative AI systems.