Context compaction

The gateway trims a request's conversation history before it reaches the provider, removing tool results a fast classifier judges no longer necessary. It never summarises: user and assistant text comes out byte-identical. It ships off.

What it does, and what it never does

On the way out, before the request reaches the provider, the gateway reads the conversation history and pairs each tool call with its result. For every pair a classifier answers two questions: does this call still need to be here, and does its full result still need to be here. What the classifier marks as no longer necessary is removed. Everything else goes through untouched.

Client (Agent, CLI or application) -> Gateway (Pairs calls and results, consults the engine, rewrites) -> Provider (Receives a shorter history, with the same prose). Compaction happens on the outbound path of a single request. Nothing is stored, nothing is rewritten on the client side.

The guarantee: If you compare the request that left the gateway with the one that came in, the difference is always a subtraction. Nothing in the output was written by the compaction engine.

  • Prose is never touched. User messages and assistant messages come out byte-identical, character for character.
  • Nothing is rewritten. There is no summary, no paraphrase, no condensed version of anything. The only operation is removal.
  • A tool call and its result are judged as a pair. A result can be removed while its call stays, but a call never leaves without its result leaving too.
  • The decision is per call, not per conversation. Two results of the same tool in the same session can get different verdicts.

Two engines, one protocol

Who decides can be a model the customer runs on their own machine (the local engine) or a hosted decision service (the remote engine). Both speak the same wire protocol, so the only difference in configuration is the endpoint. Nothing else in the block changes.

  • The model runs on the customer's own machine. No excerpt of any conversation leaves the network.
  • The local engine refuses any endpoint that is not a local address. The configured address is resolved at boot, and if it resolves to a public address the boot fails. There is no flag to relax this.
  • This is the recommended setup for regulated environments and the only one available in an air-gapped install.
  • Requires an explicit consent flag in the configuration. Without it the gateway refuses to start the engine, because excerpts of the conversation do leave the network to be judged.
  • Announces its destination at boot, in the log, so no one discovers after the fact where the excerpts went.
  • Records the consent, dated, in the signed audit trail. The record survives a later change to the configuration file.

Local engine

Remote engine

It ships off, and it fails open

Compaction is off in a fresh install. With no configuration block and no environment variable turning it on, the gateway never calls any engine and the history goes to the provider exactly as it arrived. Turning it on is a deliberate act.

Once on, every failure path ends in the same place: the ORIGINAL history is sent.

The rule: No request is ever broken to save tokens. Compaction is an optimisation, so any doubt resolves in favour of sending everything.

  • Engine unreachable or returning an error: original history.
  • Deadline exceeded: original history. The request does not wait for a slow engine.
  • A reduction too small to be worth the rewrite: original history. Compaction gives up on its own.
  • A malformed or unparseable answer: original history.

The credential is the name of a variable

The configuration field for the engine credential holds the NAME of an environment variable, never a pasted value. The gateway reads the variable at startup. A configuration file that carries the secret itself is refused, with a message naming the field, so a credential never reaches a backup, a repository or a support ticket by accident.

"api_key_env": "ATLASBERG_COMPACTACAO_API_KEY"   ok: the name of the variable
"api_key_env": "sk-live-9f2c..."                 refused: a value in the file

Why the threshold matters

This is the part that decides whether compaction is worth anything at all, and it is not intuitive. In a long agent session, most input tokens are not read at full price: they are cache reads, billed at 0.1x. The prefix of the conversation stays stable turn after turn, so it stays in the cache and costs almost nothing to send again.

Compaction rewrites that prefix. A new prefix is a cache miss, and entering the cache costs 1.25x. So every compaction starts in the negative: it pays 1.25x once, on a smaller history, to then read at 0.1x for as long as the new prefix survives. It only turns a profit if enough turns happen after it.

  • A compaction on a short history, or one followed by two more turns, loses money even though it removed tokens.
  • That is why it triggers on a threshold: below a certain history size the rewrite cannot pay for itself, so nothing is attempted.
  • And that is why it gives up on its own when the reduction does not justify the rewrite, even after the engine has already answered.
  • Compacting more often is not the same as saving more. Beyond a point, it is the opposite.

Measurement mode

Before turning compaction on, the gateway can measure what compaction WOULD save on the customer's own traffic without compacting anything. It runs the engine alongside real requests, records what it would have removed and what that would have been worth, and sends every request untouched. No behaviour changes while it is measuring.

Start here. Traffic profiles differ enormously: a workload of short conversations with few tool calls has almost nothing to remove, and a measurement week costs a few cents to find that out. The numbers below came out of exactly this mode.

The variable works with the feature still off. What it produces are counts, never text: how many calls would have become questions, how many tokens would have left the history and what that would be worth. No prompt content reaches the metric or the trail.

ATLASBERG_COMPACTACAO_MEDIR=1

Measured results, with the premises stated

From 30 real agent sessions, 6,892 turns and 250 real consultations to the engine:

The two numbers at the bottom belong together. 40.6% of the removals deleted material the conversation later mentioned again, which sounds alarming until you look at what was removed: 97.6% of everything removed came from deterministic tools, a file still on disk, a command that runs again with the same output. The material is recoverable by repeating the call. That is why the maximum cost of a classifier error is one repeated tool call, and why the worst case above is US$ 4.32 and not a broken session.

Two limits, stated plainly: The cost model covers only the input side of the conversation. And task-completion quality with and without compaction was NOT measured: nothing here says the agent finished the same work equally well.

  • Measure: Saving | Result: 17.01% (US$ 101.58 of a modelled US$ 597.22 input cost)
  • Measure: Engine cost | Result: US$ 0.26
  • Measure: Worst-case loss | Result: US$ 4.32, if everything removed had to be fetched again
  • Measure: Balance | Result: US$ 97.00, a ratio of 22x
  • Measure: Latency | Result: 330 ms, paid by 1.20% of requests
  • Measure: Removals later mentioned again | Result: 40.6%
  • Measure: Removals from deterministic tools | Result: 97.6%

What compaction cannot reach

Compaction only ever addresses part of the bill, and it is worth knowing which part before expecting a number from it.

  • The fixed prefix. The median is 162,811 tokens per turn of system prompt, tool definitions and memory, which arrives before the conversation even starts. Compaction never touches it, and no threshold changes that.
  • Tool call inputs. They are 46.9% of the history, against 44.0% for the results. An input stays as long as its call stays, so the larger half of the history is mostly out of reach.
  • What is left is the part compaction works on, and 17.01% of the input cost is what came out of it on this traffic.

Configuration

Recommended order: measure first on your own traffic, read the reduction against the fixed prefix you cannot touch, and only then set enabled to true, starting with the local engine.

  • Field: enabled | What moving it does: The master switch. False in a fresh install; with false, no engine is ever called.
  • Field: engine | What moving it does: Chooses who decides: the local engine or the remote engine.
  • Field: endpoint | What moving it does: Address of the decision engine. The local engine refuses anything that is not a local address and fails the boot if it resolves to a public one.
  • Field: model | What moving it does: Which decision model the endpoint should load. A heavier model decides better, costs more and adds latency.
  • Field: api_key_env | What moving it does: The NAME of the environment variable holding the credential. A pasted value here is refused.
  • Field: limiar_tokens | What moving it does: History size from which compaction is attempted. Raising it makes compaction rarer and each one more likely to pay for itself.
  • Field: piso_bytes | What moving it does: Minimum result size to be considered at all. Raising it stops the engine from spending a decision on results too small to matter.
  • Field: preservar_recentes | What moving it does: How many recent turns are never touched. Raising it protects more of the tail and reduces the saving.
  • Field: limiar_decisao | What moving it does: Confidence the classifier needs before removing. Raising it makes it more conservative and removes less.
  • Field: truncar_cabeca_chars | What moving it does: How many characters of the head of each result go to the classifier. Raising it gives a better basis for the decision and raises both the engine bill and the latency.
  • Field: timeout_ms | What moving it does: Deadline for the engine's answer. Exceeding it sends the original history.
  • Field: min_reducao | What moving it does: Minimum reduction accepted for the rewrite to happen. Below it, the original history goes.
  • Field: consentimento_exportar_prompt | What moving it does: Explicit consent to send excerpts of the conversation to the remote engine. Without it, the remote engine does not start. It is recorded, dated, in the audit trail.
"compactacao": {
  "enabled": true,
  "engine": "local",
  "endpoint": "http://127.0.0.1:8080",
  "model": "<decision model>",
  "api_key_env": "ATLASBERG_COMPACTACAO_API_KEY",
  "limiar_tokens": 40000,
  "piso_bytes": 2048,
  "preservar_recentes": 6,
  "limiar_decisao": 0.7,
  "truncar_cabeca_chars": 800,
  "timeout_ms": 1500,
  "min_reducao": 0.1,
  "consentimento_exportar_prompt": false
}

About Atlasberg Platform

Atlasberg Platform is the control layer between a company and every AI model: each request is authenticated with a virtual key, filtered by policy and DLP, routed to the right provider and written to a hash-chained audit trail. It exposes an OpenAI-compatible API, so applications only swap the base URL.

The same artifact runs in Atlasberg Cloud, in your VPC, on-premises or fully air-gapped, and is priced by capacity and modules, never per seat.

This page is part of the official documentation. To talk to the engineering team, write to [email protected] or use https://atlasberg.com/contato. Answers come within one business day.

Agents: this page is also available as Markdown at /docs/console/compactacao-de-contexto.md, or by requesting this URL with the header Accept: text/markdown. Index of everything: /llms.txt.