Context compaction saves 17%, not 60%

We measured it on real traffic. The numbers came out far less pretty than the promise, and far more useful. Context compaction is available in Atlasberg Platform starting today.

About this article

Published 2026-09-23 by Rafael Hickmann, Founder, Atlasberg. 9 minute read.

What we tested

The idea behind context compaction is to prune a conversation's history before sending the whole thing to the model. Prune, not summarise. Each tool call is paired with its result, a fast classifier decides call by call whether that still needs to be there, and whatever does not need to be there goes. Human text and assistant text are never touched.

We built this inside Atlasberg Platform and ran it against real working sessions, not a synthetic benchmark.

The sample: 30 real agent sessions, 6,892 turns, 13,867 messages and 30,831 lines of transcript. The classifier was consulted for real, 250 times, on 1,025 decisions. Nothing simulated.

Who decides is Jev, not an LLM

The classifier making these decisions is Jev, a model by TypeSafe that does not generate text. It receives the whole conversation as state, up to 32 thousand tokens, with every tool result replaced by a short note, and answers yes-or-no questions with a calibrated probability. For each tool call we ask two: does the call still matter? Does the result still need to stay in full? Whatever falls below the threshold goes, or keeps only its head plus a note saying it was truncated.

That changes the economics of the problem. Asking an LLM to summarise the history costs a meaningful slice of what you were trying to save, and returns a summary that can drop exactly the file path or the precise error the agent will need later. Jev costs US$ 0.042 per million input tokens, 1/119 of what a frontier model charges for input, and does not charge for output. Across the 30 sessions, 250 consultations and 1,025 decisions cost US$ 0.26. And because it sees the whole conversation, it can answer the question that actually matters: will this result still be needed further down?

[Figure: Three numbers about Jev: US$ 0.042 per million input tokens, 32 thousand tokens of state holding the whole conversation, and 1,025 decisions in 250 consultations for US$ 0.26. - Jev answers yes or no with a calibrated probability. It does not generate text, summarise or rewrite.]

The algorithm came from fast-jev-compaction, an open library under MIT that replaces the compaction summary with Jev decisions. We ported the idea into the gateway, in Go, with two changes: the threshold trigger, which the cache section explains, and a second engine.

The second engine is Laya, by Convai Innovations, with open weights under Apache 2.0. It is an encoder of 322 to 421 million parameters that runs on your own machine with about 1 GB of memory and, according to the model's documentation, answers a question in 7 ms when batched. It speaks the same protocol as Jev, so switching engines means changing the address. What changes is the reach: Laya sees about 768 tokens at a time, so it gets one state per call, with the signal that matters most (was this file mentioned again later?) computed deterministically by the gateway. It is the path for anyone who cannot let an excerpt of a conversation leave the network. Every number in this article comes from Jev.

[Figure: Comparison between Jev, the remote engine, and Laya, the local engine: visible state, where it runs, whether the prompt leaves the network, cost and latency per question. - Two engines, one protocol. Switching engines means changing the address in the configuration.]

The number

17.01% saved, or US$ 101.58 out of a modelled input cost of US$ 597.22. Jev, the classifier that made those decisions, cost US$ 0.26.

Before measuring, our own projection said 60%. We were wrong, and the reason is worth telling.

Why the projection was wrong

We picked the first sessions by file size, which seemed reasonable at the time. But the largest transcripts were 38 MB with only 90 KB of prose. Everything else was base64 images. That pushed the "tool result" share up to 96% of characters and produced a projection that was beautiful and completely false.

We rebuilt the sample using the context the provider actually reported, and the composition is a different story:

[Figure: Stacked bar with the composition of the history: 46.9% tool call inputs, 44.0% tool results, 9.2% prose. Only the results slice can be pruned. - What fills the history of an agent session, measured on the context the provider reported. Only the gold slice can be pruned.]

In other words: you do not prune what you sent, only what came back.

  • What fills the history: Tool call inputs | Share: 46.9% | Can it be pruned?: No. If the call stays, its input stays with it
  • What fills the history: Tool results | Share: 44.0% | Can it be pruned?: Yes, and it is the only target
  • What fills the history: Prose, including reasoning | Share: 9.2% | Can it be pruned?: No, by invariant

The finding that changes the problem

The median fixed prefix per turn is 162,811 tokens. That is the system prompt, tool definitions and memory, stuff that arrives before the conversation and that compaction cannot reach at all. In several sessions that prefix alone is bigger than the entire history available for pruning.

[Figure: Three numbers: 162,811 tokens of median fixed prefix per turn, 95% of input tokens read from cache at 0.1x, and 83 of 346 triggers that became an actual compaction. - Where the cost lives. The fixed prefix arrives before the first word of the conversation, and no threshold changes that.]

Anyone trying to cut agent cost by looking only at the conversation is aiming at the smaller half of the problem.

The second finding: the cache had already solved most of it

In a long session, 95% of input tokens are not billed at full price. They are cache reads, at 0.1x.

That flips the intuition. Compaction rewrites the prefix, and a new prefix costs 1.25x to enter the cache. In practice every compaction starts in the red and needs enough turns after it to pay for itself.

That is why compacting on every request is the worst possible policy. Ours fires on a threshold and gives up on its own when the reduction does not justify the rewrite. Of 346 triggers, only 83 became an actual compaction. The other 263 were discarded because they were not worth the price of their own cut.

And quality? The question nobody publishes

Savings without that measurement mean nothing. If you delete 40% of the context and the agent starts failing, you did not save anything, you broke it.

So we measured the other side as well, and here is the ugly number: 40.6% of the removals deleted material the conversation later mentioned again.

It sounds terrible. It is, in fact, a ceiling on damage, and the difference matters for two reasons.

The first is that what goes out is recoverable. 97.6% of everything removed came from Bash and Read, deterministic tools. The file is still on disk, the command runs again. The algorithm also leaves the first 300 characters of the result plus a note, in plain text, saying that it was truncated and can be fetched again. Only 1.7% of removals came from web search and web reads, the only ones where the content may have changed in the meantime. Nothing is lost for good; the worst cost of a classifier mistake is one repeated tool call.

The second is that you can put a price on it. Assuming the worst possible case, where every one of those 40.6% had to be fetched again and came back in full, paying a cache write:

[Figure: Bars comparing the measured gain of US$ 101.58 with the worst-case loss of US$ 4.32 and the classifier cost of US$ 0.26. Balance of US$ 97.00, gain 22 times the loss. - Measured gain against the worst case we could build: everything mentioned later comes back in full, paying a cache write.]

The gain beats the loss 22 times over in the most pessimistic scenario we could construct. In the real scenario it is larger, because "mentioned later" does not mean "had to be fetched again".

Latency has the same shape. Each compaction cost 330 ms on average, and since only 83 of the 6,892 turns compacted, 1.20% of requests waited for it. The other 98.8% paid nothing.

  • Line: Measured gain | Value: US$ 101.58
  • Line: Worst-case loss (fetch again everything mentioned later) | Value: US$ 4.32
  • Line: Jev cost | Value: US$ 0.26
  • Line: Balance | Value: US$ 97.00

What we did not measure, and you should ask about

We did not measure task completion with and without compaction. To claim that quality did not drop, you would have to run the same agents in both configurations and compare the final result, and we did not do that. What we have is the damage ceiling, the recoverable nature of what goes out, and the economic balance. Anyone claiming more than that without that experiment is selling something.

We also assume the cost model sits 60.2% below the sample's actual bill, because it models only the input side of the conversation and ignores the fixed prefix, images and redacted reasoning. Both numbers are printed side by side in the tool, on purpose. Tuning the model to match the invoice would make the savings percentage prettier and less true.

What remains

Context compaction works. It does not deliver 60%, it delivers 17%, it costs 26 cents to save 101 dollars, and it charges latency on a little over 1% of requests.

But the most useful result of the measurement is not the percentage. It is finding out where the cost actually lives: in 162 thousand tokens of fixed prefix that arrive before the first word of the conversation, and in the call inputs, which are larger than everything you are allowed to prune.

The conversation was never the main problem. It was just the easiest part to look at.

Available starting today

Context compaction ships today in Atlasberg Platform as a gateway feature. It runs on the outbound path of every request, before the history reaches the provider, and works with any model that goes through the platform. Teams running agents at scale will see the input bill drop without touching a single line of the agent.

The 17% in this article came from our own traffic, and agent traffic varies a lot from one shop to the next. A workload of short conversations with few tool calls has almost nothing to prune; a fleet of coding agents running all day has plenty. What we want now is the customers' number, measured on their traffic, with the same premises printed next to it. That is what we will publish next.

If you already run Atlasberg Platform: Turn on measurement mode for a week and tell us what came out. The measurement tool, the cost model and the safety metric are part of the platform and reproducible: same sample, same premises printed next to the result.

  • It ships off. Turning it on is a deliberate act, and every failure of the decision engine ends in the same place: the original history is sent. No request is ever broken to save tokens.
  • It has a measurement mode. With ATLASBERG_COMPACTACAO_MEDIR=1 the gateway computes what compaction would save on your traffic without compacting anything. The numbers in this article came out of exactly that mode.
  • Two engines, one protocol. Jev, remote, sees the whole conversation and requires explicit consent, recorded in the audit trail. Laya, local, runs on your own machine and no excerpt of any conversation leaves the network. Switching one for the other means changing the address.
  • It decides per call, not by age. Two results of the same tool in the same session can get different verdicts.

Context compaction documentation

How it works, the two engines, the threshold, measurement mode and every configuration field.

See Atlasberg Platform live

A demo with real traffic, to watch the gateway decide request by request.

About Atlasberg

Atlasberg builds the control layer between a company and every AI model. Atlasberg Platform authenticates each request with a virtual key, filters it by policy and DLP, routes it to the right provider and writes it to a hash-chained audit trail, with an OpenAI-compatible API so applications only swap the base URL.

Articles on this blog are written by the engineering team and report measurements on real traffic, with the premises printed next to the result. To talk to the team, write to [email protected] or use https://atlasberg.com/contato.

Agents: this page is also available as Markdown at /blog/compactacao-de-contexto-17-por-cento.md, or by requesting this URL with the header Accept: text/markdown. Index of everything: /llms.txt.