Who is spending the company's AI budget?

The provider invoice arrives as a single number. Nobody knows how much came from legal, how much from the data team and how much from an agent that looped all night. How we solve it: one key per person or team, automatic chargeback per request, and a budget cap that never denies.

About this article

Published 2026-10-07 by Rafael Hickmann, Founder, Atlasberg. 7 minute read.

The key pasted in the spreadsheet

The story repeats with impressive regularity. Someone on the innovation team opens the provider account, generates a key and pastes it on a wiki page for the whole company to use. Four months later the invoice has tripled, the key lives in eleven repositories, three spreadsheets and a marketing automation tool, and nobody can say how much came from whom. The usual reaction is for each team to open its own account, which produces five invoices, five different dashboards and zero common policy.

The invoice itself is not the problem. The problem is that the unit the provider bills, the account, has nothing to do with the unit the company needs to manage: the person, the team, the application.

Why AI cost does not behave like cloud cost

People who have run cloud FinOps tend to assume this is the same thing with a different label. It is not, for three reasons. First, the price gap between models is huge: between the cheapest and the most expensive model of the same provider there can be twenty or thirty times the price per token, and the model choice sits in the application code, not in the hands of whoever pays. Second, a single agent in a loop can spend in one night what an entire team spends in a month, and it shows up in no capacity report. Third, prefix caching changes the price of the same token by a factor of ten, depending on how the application assembles the prompt.

The practical consequence is that monthly allocation by percentage, which many companies do today, gets it badly wrong. AI cost is only attributable if it is measured at the request.

The unit of attribution is the request

Every request needs an owner. In our design that happens through virtual keys: the provider key is stored once, in the gateway, and each person, team or application gets its own virtual key, with the models it may use, its budget and its limits. Developers do not even need a key: they sign in with the company login from the terminal and receive a personal token, revocable when they leave. Directory groups become teams automatically.

From there, every request records provider, model, input and output tokens, latency, status and cost, exactly one event per request and not per stream chunk. The per-team chargeback table comes out of that, with nobody filling anything in.

What to watch, and why

In the console we made the cost per thousand requests chart clickable on purpose. Clicking a team's bar opens its model mix, because that is where the conversation with the manager usually resolves itself in two minutes.

  • Metric: Cost per thousand requests, per team | What it reveals: The fair comparison between teams. Absolute cost only says who uses more.
  • Metric: The team's model mix | What it reveals: Where the cost difference almost always explains itself. A team using the large model to classify email.
  • Metric: Share of cache reads | What it reveals: In a long agent session, 95% of input tokens can be cache reads at a tenth of the price. Whoever does not use the cache pays ten times more for the same context.
  • Metric: Percentage of the monthly cap consumed | What it reveals: For talking to the team before the overrun, not after the invoice.
  • Metric: Cost per application and per agent | What it reveals: This is where the overnight loop shows up. Per person, it gets diluted.

A cap that does not deny the request

Hard limits have a problem everyone discovers the worst way: they block production at eleven on a Friday night. That is why our default budget is a soft cut. It is an optional monthly limit per organizational unit, team or person. When it is exceeded, the requests in that scope are served by the cheaper model you chose as the fallback. No request is ever denied. The application keeps working, just slower or on a smaller model, and the manager gets the notice.

Hard limits still exist, but in a different place: the per-provider budget, which protects the company from an absurd invoice. Together they give the combination we think is right: financial protection at the top, continuity at the bottom.

The waste that does not show up on the invoice

There is a kind of spend no cost report shows, because it is technically legitimate: repeated context. When we measured real agent sessions, the median fixed prefix per turn was 162,811 tokens, and that is system prompt, tool definitions and memory, things that arrive before the first word of the conversation. In several sessions the prefix alone was larger than the entire history. We told that whole measurement in the article on context compaction.

That is why the gateway has a passive waste meter. It watches tokens, declared caps and context repetition in the traffic it already governs and points to where money is escaping. It blocks nothing and changes nothing; it only shows. In the traffic we measured, the biggest savings lever was not in how people used AI. It was in two or three giant system prompts.

How to start in a week

The complete walkthrough: The tutorial on routing every AI request of the company through the gateway covers each of these steps with real screens, including the firewall part everyone skips.

  • Put the gateway in front of the providers and register their keys once.
  • Create the teams, preferably synced from the directory, and issue one virtual key per application.
  • Swap the base URL in the applications and in developers' tools. No code changes.
  • Block direct egress to the providers at the firewall, otherwise the key in the spreadsheet keeps working.
  • Wait thirty days and open the chargeback table. It is the first time the company sees the number per team.

About Atlasberg

Atlasberg builds the control layer between a company and every AI model. Atlasberg Platform authenticates each request with a virtual key, filters it by policy and DLP, routes it to the right provider and writes it to a hash-chained audit trail, with an OpenAI-compatible API so applications only swap the base URL.

Articles on this blog are written by the engineering team and report measurements on real traffic, with the premises printed next to the result. To talk to the team, write to [email protected] or use https://atlasberg.com/contato.

Agents: this page is also available as Markdown at /blog/quem-esta-gastando-o-orcamento-de-ia.md, or by requesting this URL with the header Accept: text/markdown. Index of everything: /llms.txt.