# Every AI request of the company through your gateway

> Make all LLM traffic of the company pass through Atlasberg Platform: publish the gateway with TLS, register provider keys once, block direct egress at the firewall, set a default model, create teams, virtual keys and soft-cut budgets. With a deployment checklist.

Workstations stop talking to Anthropic, OpenAI or Google. They talk to an internal name, and the gateway does the rest: identity, budget, DLP, audit trail and routing, without the developer noticing.

## What you are going to build

Developer workstation (Claude Code, Cursor, SDKs, scripts and agents. Send a virtual key.) -> llm.company.com (Atlasberg Platform. Governance, cost, DLP, audit trail, routing.) -> Providers (api.anthropic.com, api.openai.com, generativelanguage. Receive the real key.). The firewall blocks the dotted path: workstations cannot reach the providers directly. Only the gateway can.

- The real key lives only in the gateway. The Anthropic, OpenAI or any other provider key is registered once, under Models > Providers. Nobody has a provider key. What a person receives is a virtual key, which is only valid against the gateway.
- The firewall blocks direct egress. Provider domains are blocked for the whole network, except for the gateway. Without this rule, "all traffic passes through us" is an intention. With it, it is a fact, and any misconfigured tool fails loudly instead of leaking.
- Clients only change the address. Claude Code, Cursor and SDKs keep speaking their native protocol. The only difference is the base URL, which points to the matching drop-in route on the gateway.
- Cost per person, team and customer, with budgets and usage limits enforced before the invoice arrives.
- DLP, guardrails and an audit trail on every request, without relying on the developer remembering anything.
- A default model and routing: one global rule decides which model serves whoever does not choose, and the complexity router sends simple tasks to a cheap model.
- The Economometer: a passive meter that points at giant prompts, repeated context, exaggerated token ceilings and expensive models on trivial tasks, with the estimated saving.

### The three rules

### What the company gains in the same move

## Five concepts that matter

- Header: x-atlasberg-vk | Rule: Accepted verbatim | Who sends it this way: Own integrations, scripts
- Header: Authorization: Bearer | Rule: Only with the sk-atlas- prefix | Who sends it this way: OpenAI SDK, Codex, Claude Code via ANTHROPIC_AUTH_TOKEN
- Header: x-api-key | Rule: Only with the sk-atlas- prefix | Who sends it this way: Anthropic SDK, Claude Code via ANTHROPIC_API_KEY
- Header: x-goog-api-key | Rule: Only with the sk-atlas- prefix | Who sends it this way: Google GenAI SDK, Gemini CLI
- Header: api-key | Rule: Only with the sk-atlas- prefix | Who sends it this way: Azure OpenAI SDK
- Route: /anthropic | Speaks the protocol of: Anthropic Messages, count_tokens, models, batches, files | Used by: Claude Code, Anthropic SDK
- Route: /openai | Speaks the protocol of: OpenAI chat, responses, embeddings, files, batches | Used by: OpenAI SDK, Codex, Continue, Cline, Aider
- Route: /cursor | Speaks the protocol of: OpenAI, with the hybrid payload Cursor sends | Used by: Cursor
- Route: /genai | Speaks the protocol of: Google GenAI | Used by: Google SDK, Gemini CLI
- Route: /bedrock, /cohere | Speaks the protocol of: Bedrock and Cohere | Used by: Matching SDKs
- Route: /litellm, /langchain, /pydanticai | Speaks the protocol of: Those frameworks' formats | Used by: Internal applications
- Route: /v1/chat/completions, /v1/messages | Speaks the protocol of: Unified API at the root | Used by: Any client

### Virtual key

The credential a person, an agent or a pipeline uses against the gateway. It always starts with the sk-atlas- prefix. Each key carries the allowed providers and models (empty list denies everything; the default is deny), a budget in dollars with a reset period, a usage limit in tokens and requests per window, an optional expiry, an active or inactive state, and the MCP tools it may run.

### Customer, team and key

The hierarchy is customer > team > virtual key. A key belongs to a team, and a team may belong to a customer. Each level has its own budget and limits, and evaluation follows the order key, customer, team. There is no "user" entity that owns a key: to give a key to a person, you create a key with their name inside their team. The key becomes the person's identifier in cost, trail, Economometer and chargeback. Tutorial 3 shows how single sign-on replaces that per-person key with a personal token.

### Where the gateway reads the key

The prefix is mandatory in SDK headers by design: a provider key is never mistaken for a virtual key.

### Drop-in routes

### Model name

The client may send the model without a prefix, such as claude-sonnet-5. The model catalog resolves the provider. With a prefix, such as anthropic/claude-sonnet-5, the choice is explicit. If the catalog does not know the model and there is no prefix, the gateway answers 400 asking for the provider.

## Part 1: prepare the gateway

Infrastructure work, done once. At the end of this part the gateway is published on the network, holds the real provider keys and has a default model.

About internal TLS: If the certificate comes from the corporate CA and Claude Code complains about it, point NODE_EXTRA_CA_CERTS at the CA's PEM bundle in the managed settings file. That applies to any Node tool.

### 1. Publish an internal name with valid TLS

Choose a name such as llm.company.com. The certificate may come from the corporate CA, as long as that CA is installed on every workstation. Claude Code, Cursor and SDKs validate the certificate like any HTTPS client. If you prefer, a public certificate solves it without depending on the MDM.

### 2. Register the provider keys

In the console, under Models > Providers, add the Anthropic, OpenAI and any other provider key the company uses. Prefer credentials via environment variable, which stay out of the appliance state file. Mark the enabled models. This feeds the catalog that resolves names without a prefix.

### 3. Allow the gateway out and block the workstations

At the firewall, the gateway may reach api.anthropic.com, api.openai.com, generativelanguage.googleapis.com and the rest. The rest of the network may not. Put this rule in place before announcing the gateway, so nobody depends on a path that is about to close.

### 4. Keep the key requirement on

The gateway ships requiring a virtual key on every inference (enforce_auth_on_inference). Keep it that way. The read is fail-closed: a missing value is treated as "require". You can check the effective value in GET /api/config, in the enforcement block.

### 5. Keep direct keys off

The allow_direct_keys flag lets a client bring its own provider key with the x-atlasberg-direct-key: true header and bypass the registered pool. That defeats rule number one. Leave it off.

### 6. Set the default model with a global rule

Under Models > Routing rules, create a rule with global scope, the CEL expression true and a target with the desired provider and model. Every request now lands on that model, regardless of what the client asked for. If you want something finer, use an expression that only matches requests without a preference, or a rule per team.

### 7. Optional: turn on the complexity router

The gateway classifies each prompt into a complexity tier. A rule that references complexity_tier sends simple tasks to a cheap model and heavy tasks to a strong one. Combine it with the default model: the complexity rule comes first, the default closes the rest.

```
curl -s https://llm.company.com/api/config | jq .enforcement
# { "enforce_auth_on_inference": true, "allow_direct_keys": false }
```

## Part 2: build the governance

This is where attribution is born. Without it the gateway still governs, but cannot say who spent what.

Self-service: Anyone can check their own quota without admin access, presenting only the key: GET /api/governance/virtual-keys/quota with the x-atlasberg-vk header. Good for a status script in the terminal bar.

### 1. Create customers, if it makes sense

Under Governance > Customers, register business units, subsidiaries or accounts that need a separate invoice. A small company can skip this step.

### 2. Create teams with a budget

Under Governance > Teams, create one team per area (Engineering, Product, Marketing). Give each team a monthly budget and, if you want, a daily token limit. Team budgets are what the Costs screen and the Executive panel add up.

### 3. Create one key per person inside the team

Under Governance > Virtual keys, create the key with the person's name (eng-maria), assign it to the team, mark the allowed providers and models, set a budget and an expiry. Copy the sk-atlas-... value: it is shown once. Agents, CI pipelines and bots get their own key, never a person's key. If the allowed models list is empty, the key denies everything; mark at least one model per provider.

### 4. Turn on soft cut in budgets

Under Costs > Budgets, set a monthly budget per unit, team or person that, when exceeded, degrades to a cheaper model instead of denying the request. Precedence is person, then team, then unit. Careful: the cheap model must be in the key's allowed list. If it is not, the gateway does not rewrite, records the budget.degrade_blocked event and warns in the panel.

### 5. Deliver the key through infrastructure, not through the person

The right way for the key to reach the machine is provisioning: MDM, onboarding script or a secret in the corporate password manager. The person should not paste a key anywhere. Tutorial 2 shows how Claude Code reads the key without anyone seeing it, and Tutorial 3 replaces the key with single sign-on.

## Part 3: verify and troubleshoot

Four calls with curl prove the deployment end to end. Then, the errors that show up most and what each one means.

- Response: 401 virtual_key_required | Meaning: The request arrived without a key and the requirement is on | What to do: Check the variable or the apiKeyHelper; check that the key has the sk-atlas- prefix in the header used
- Response: 401 virtual_key_not_found | Meaning: The key exists on the machine but not on the gateway: revoked, expired or mistyped | What to do: Look at the key under Governance; reissue if needed
- Response: 403 model or provider not allowed | Meaning: The key does not have that model in its allowed list | What to do: Add the model to the key or the team; remember an empty list denies everything
- Response: 400 "could not auto resolve a provider" | Meaning: The model is not in the catalog and came without a prefix | What to do: Enable the model under Providers or use provider/model
- Response: 429 budget or limit | Meaning: Key, team or customer exceeded the budget or the rate limit | What to do: Check the quota; adjust the budget; turn on soft cut to degrade instead of deny
- Response: Certificate error in the client | Meaning: The corporate CA is not on the workstation or is not seen by Node | What to do: Install the CA via MDM; set NODE_EXTRA_CA_CERTS in the managed file
- Response: Claude Code slow to open or with network warnings | Meaning: Telemetry and update calls hitting the firewall | What to do: Set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
- Response: Panel shows zero cost with traffic | Meaning: The provider did not report a price, or the model has no price in the catalog | What to do: Check the catalog and the custom price; the gateway never invents a value

### Common errors

```
curl -s https://llm.company.com/health
# {"status":"ok",...}
```

```
curl -s https://llm.company.com/anthropic/v1/messages \
  -H "x-api-key: sk-atlas-..." \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","max_tokens":32,
       "messages":[{"role":"user","content":"answer only ok"}]}'
```

```
curl -s https://llm.company.com/api/governance/virtual-keys/quota \
  -H "x-atlasberg-vk: sk-atlas-..."
```

```
curl -s --max-time 5 https://api.anthropic.com/v1/models
# curl: (28) Connection timed out
```

## Deployment checklist

- Internal name published with TLS valid on the workstations
- Provider keys registered via environment variable
- Models enabled in the catalog, with a price
- Firewall: gateway goes out, workstations do not
- enforce_auth_on_inference on and checked in /api/config
- allow_direct_keys off
- Global default model rule created
- Customers and teams created with a monthly budget
- One key per person, one per agent or pipeline, all inside a team
- Allowed models marked on each key
- Soft cut configured, with the cheap model in the allowed list
- The four tests above pass from a regular workstation
- The test request shows up under Costs with the person's and the team's name
- A request without a key receives 401
- A direct call to the provider fails by network

### Gateway

### Governance

### Validation

## Do it all with an AI agent

If you would rather delegate, paste the prompt below into Claude Code or another AI agent with terminal access. It is written to ask for everything it needs before touching anything, show a plan, wait for your confirmation and only then execute. It never asks for the value of a provider key and never writes a secret into a repository file.

### Next: connect Claude Code, Cursor and SDKs

Two variables, a managed settings file and the credential helper that keeps the key invisible.

### Or skip the key entirely

Configure SSO and device login so people sign in with the corporate identity, from the terminal.

```text
You are my infrastructure engineer and you are going to configure Atlasberg Platform as the company's single AI gateway, following the tutorial at https://atlasberg.com/docs/tutoriais/gateway-da-empresa and the management API reference at https://atlasberg.com/docs/api/gestao. Work through the gateway API with curl. Never invent values: anything not listed below, ask.

Before doing anything, ask me the following, one question at a time, and wait for the answer:
1. The gateway URL (example: https://llm.company.com) and how to authenticate to the management API (administrator username and password, or a token). Never echo the password back into the conversation and never write it to a file.
2. Which providers the company uses (Anthropic, OpenAI, Google, others) and, for each one, the NAME of the appliance environment variable that holds the real key. Do not ask me for the key value. If I paste a key by mistake, warn me and ask me to rotate it.
3. The models that must be enabled per provider and which one is the company default.
4. Whether there are business units or customers that need a separate invoice, with name and monthly budget in dollars. If none, skip.
5. The teams: name, monthly budget in dollars, token limit per hour (optional) and which unit they belong to.
6. The people, agents and pipelines in each team that need their own key, and the allowed models for each. One key per person, one per agent or pipeline, never shared.
7. Whether I want soft cut (when the budget is exceeded, degrade to a cheaper model instead of denying) and which cheap model to use. It must be in the allowed list of every key.
8. Whether the gateway uses a certificate from a corporate CA.

When you have everything, show a short plan as a table (providers, units, teams, keys, default model rule) and ask for my confirmation. Only then execute, in this order, checking the status code of every response and stopping at the first error:
a) GET /api/config: confirm enforcement.enforce_auth_on_inference is true and allow_direct_keys is false. If not, tell me before changing anything.
b) Providers: GET /api/providers to see what exists; POST /api/providers with {"provider":"<name>"} for the missing ones. For the key, first read the response of GET /api/providers/<name> and the API reference to learn the exact shape of the value field, then register the key by reference to the environment variable, never by value, with POST /api/providers/<name>/keys, the enabled models and weight 1.
c) Units: POST /api/governance/customers with {"name":"...","budgets":[{"max_limit":<usd>,"reset_duration":"1M"}],"calendar_aligned":true}.
d) Teams: POST /api/governance/teams with {"name":"...","customer_id":"<unit id>","budgets":[{"max_limit":<usd>,"reset_duration":"1M"}],"calendar_aligned":true} and, if there is a limit, "rate_limit":{"token_max_limit":<n>,"token_reset_duration":"1h"}.
e) Virtual keys: POST /api/governance/virtual-keys with {"name":"eng-maria","description":"...","team_id":"<team id>","provider_configs":[{"provider":"anthropic","allowed_models":["claude-sonnet-5"],"key_ids":["*"]}],"is_active":true} and, if any, "budgets":[...]. Keep the value field of each response and hand me all the keys ONCE, in a separate block at the end, warning that they will not be shown again. Do not write the keys to any file.
f) Default model: POST /api/governance/routing-rules with {"name":"Company default model","description":"Requests without a model choice land here","cel_expression":"true","targets":[{"provider":"<provider>","model":"<model>","weight":1}],"scope":"global","priority":100,"enabled":true,"chain_rule":false}.
g) Soft cut, if I asked for it: configure it under Costs > Budgets in the console, or through the matching API in the reference; confirm the cheap model is allowed on every key.
h) Tests: GET /health; POST /anthropic/v1/messages with one of the keys in the x-api-key header and a 32-token request; GET /api/governance/virtual-keys/quota with the x-atlasberg-vk header. Show the responses, summarized.
i) Deliver the tutorial's deployment checklist filled in, marking what is still pending outside the gateway: internal name with TLS, the firewall rule that blocks direct egress from workstations to the providers, key distribution through MDM or an onboarding script, and NODE_EXTRA_CA_CERTS if the CA is corporate.

Rules: if an endpoint answers differently from what is expected, show the response and ask before working around it. Do not create a key without a team. Do not enable direct provider keys. Do not turn off the virtual key requirement.
```

## About Atlasberg Platform

Atlasberg Platform is the control layer between a company and every AI model: each request is authenticated with a virtual key, filtered by policy and DLP, routed to the right provider and written to a hash-chained audit trail. It exposes an OpenAI-compatible API, so applications only swap the base URL.

The same artifact runs in Atlasberg Cloud, in your VPC, on-premises or fully air-gapped, and is priced by capacity and modules, never per seat.

This page is part of the official documentation. To talk to the engineering team, write to contact@atlasberg.com or use https://atlasberg.com/contato. Answers come within one business day.

## More from Atlasberg

- [Documentation index](https://atlasberg.com/docs)
- [Platform overview](https://atlasberg.com/)
- [Talk to engineering](https://atlasberg.com/contato)

---
Source: https://atlasberg.com/docs/tutoriais/gateway-da-empresa
Company: Sobimann Tecnologia da Informacao LTDA (Atlasberg), Porto Alegre, RS, Brazil.
Contact: contact@atlasberg.com - answered within one business day.
Index for agents: https://atlasberg.com/llms.txt
