How to track AI spend, and what an AI governance layer actually is
You can't govern AI spend until every call to a model carries a team and a workflow on it. Here is what to build, what to measure, and how to compare an AI workflow against a hire without fooling yourself.
Tag every AI API call with the team and the workflow that made it. Do that before you buy a tool, write a policy or argue about ROI, because until it's done you have a bill and no way to read it, and every question about whether AI is paying for itself turns into a matter of opinion.
That's the shift we're watching inside our own accounts. A few months ago most of our engineering clients were approving unlimited API budgets for OpenAI and Anthropic with no tracking on where the tokens went. One had three teams running up five-figure token bills before a single production agent shipped. Today the ask is a governance layer and a monitor on every dollar.
Why unlimited budgets happened
They were rational at the time. The spend was small, the models were improving month to month, and nobody wanted to be the company that made its engineers file a purchase order while a competitor shipped. A handful of developers on a coding assistant cost about what a few SaaS seats cost, and approving that was cheaper than a procurement cycle.
Two things ended that. The first is scale. An assistant that answers a developer is one kind of bill. An agent that runs in a loop against a seven-year-old codebase, re-reading files and retrying tool calls, is a different kind of bill, and it arrives without anyone deciding to spend the money. The second is the audit question. Once model output reaches production, finance asks what it cost and engineering asks which part of the product the model wrote. A flat invoice from a provider answers neither.
What a governance layer actually is
In plain words, it's the thin piece of your own infrastructure that every AI call passes through, so you can say who made the call, why, and what it cost.
It's usually a gateway, a key policy and a place the records land.
The gateway is yours. No team calls a provider directly with a raw key. They call your endpoint, which attaches your own labels to the request before forwarding it: team, service, workflow, environment, and the person or agent that triggered the run. The provider's response comes back with its token counts, and you store those next to your labels.
Keys get issued per team and per workflow. One shared company key is what makes the cost question unanswerable, because the provider's dashboard can only report what the key told it, and a shared key tells it nothing.
The records land in your database, next to your own identifiers: the repository, the ticket, the customer, the job. A provider console is fine for a rough monthly total. It can't give you cost per closed ticket, because it doesn't know what a ticket is in your company.
Then limits. A monthly ceiling per team, an alert well under it, and a hard stop for runaway loops. The hard stop matters more than the ceiling. An agent that retries a failing tool call a few thousand times overnight spends real money and ships nothing.
None of this needs a vendor. It's a few days of plumbing and then a decision about who owns the numbers once they exist.
What to measure
Spend per team and per workflow. This is the one that changes behavior, because it turns a company-wide number nobody owns into a number with a name on it. Workflow matters as much as team: knowing the platform group's total tells you less than knowing which of its four automations spent it.
What share of shipped code was AI-generated. Mark it when the code is written, not months later from a guess. A trailer on the commit, a label on the pull request, something the tooling sets without a developer remembering. Reconstructing authorship later from style or diff size produces a number you can't defend, and the first time someone senior challenges it you lose the whole measurement program.
Cost per task. Divide the spend for a workflow by the units it produced: tickets closed, documents processed, pull requests merged. Raw token totals move when prices change and when context windows grow, while cost per unit of output tells you whether the workflow is getting better or just getting used more.
Cycle time against the same work done by hand. This needs a baseline, which means measuring the manual version before you automate it, or keeping one team on the manual path long enough to compare. Teams that skip this step end up with an AI workflow that's clearly faster than nothing and unprovable against the old way.
How many agents actually reached production. Count them against how many were started, and count only the ones serving real work for a month or more. If the gap is wide, that's useful rather than embarrassing, because a wide gap usually shows the constraint isn't model quality. It's evaluation, data access, error handling and who gets paged when the agent is wrong.
Comparing an AI workflow against a hire
Compare fully loaded cost to fully loaded cost, over the same window, on the same scope of work. Anything else is a sales deck.
Putting API spend next to a salary is wrong on both sides. On the AI side, add the engineering time to build the workflow and to keep it running when a provider changes a model, plus review time for the output, the evaluation harness, and the failures that reach a customer. On the hiring side, a salary isn't the cost either: add benefits and payroll tax, recruiting, the ramp before the person is productive, and the senior time spent on that ramp.
Then be honest about what each one gives you. A person absorbs ambiguity, carries context between projects and can be pointed at a new problem on a Monday. A workflow does one job cheaply at any hour and nothing else. Both are worth paying for, and they aren't substitutes. A comparison that pretends otherwise lands wherever its author already wanted it to.
One test keeps the math honest. If the case only works when you assume the workflow needs no maintenance, the case doesn't work.
If nothing is tracked today
Start by finding every place a provider key exists. You'll find them in CI, in a developer's local environment file, in one function nobody owns. Rotate them into per-team keys behind a single endpoint. That one step gives you attribution before you've built any reporting at all.
Then pick one workflow and instrument it end to end until you have its cost per unit of output. One real number beats a dashboard of estimates, and it gives every other team a format to copy.
Set ceilings and alerts low at the same time. Early limits exist to catch loops, not to ration engineers.
Next, agree with the engineers on how AI-written code gets marked. A convention imposed from outside gets skipped quietly.
And take a baseline for the next workflow before you automate it, while the manual version still exists to measure.
The trap of judging AI by token spend
Token spend is an input. On its own it tells you nothing about whether the work was worth doing.
A falling bill can mean a workflow got efficient. It can also mean people stopped using it and went back to doing the job by hand, which is the explanation nobody checks. A rising bill can mean waste, or it can mean a workflow that returns more than it costs and deserves a bigger budget.
So pair every spend number with the output it bought. If you can't name the output, that's the finding.
If you want to see what this would look like on your stack, talk to the team. Bring last month's provider invoice and the list of teams holding keys.