# How to keep a CPF from leaving the company inside a prompt

> Where personal identifiers leak when a company uses ChatGPT, Claude or a model API, why training and site blocking fail, and how to mask or block personal data at the gateway with a verifiable record for LGPD and GDPR.

Training does not hold, browser extensions cannot see the API, and banning the site just moves usage to personal phones. The only place a personal identifier can be stopped is on the network path, before it reaches the provider. This is how we do it, and this is what remains as evidence.

## About this article

Published 2026-10-07 by Rafael Hickmann, Founder, Atlasberg. 7 minute read.

## Where the identifier actually leaks

It is almost never bad faith. It is an analyst in finance pasting a spreadsheet of overdue accounts into the chat to get a summary. It is a lawyer sending the whole contract so the model can find the termination clause. It is a developer writing correct code that happens to send the complete customer record to the model when the order history would have been enough. In none of these cases did the person see the CPF, the Brazilian tax ID. It was in the middle of forty lines.

Whenever we map this inside a company, the same three doors show up: the chat in the browser, the agent or IDE on a developer's machine, and the company's own applications calling the provider API. The second and third are the ones nobody is watching.

## What companies try first, and why it does not hold

Internal policy and training are necessary, and we recommend them, but they are not controls. Nobody reads a 2,000 character email looking for an eleven digit number before pasting. Blocking the ChatGPT site at the firewall works for a week; then usage moves to personal phones and you lose even the visibility you had. Browser extensions and endpoint DLP see the browser, and only the browser. They do not see the SDK call an application makes, they do not see the agent running in a terminal, and people install another browser.

There is also the option of negotiating a zero retention agreement with the provider. It helps, genuinely, and it is worth asking for. But it does not change the fact that the data left the country and was processed by a third party. Under LGPD, and under GDPR, that is still processing and still an international transfer.

## The one point every path crosses

All three doors end in the same place: an HTTPS call to the provider's address. If the company routes those calls through its own gateway and blocks direct egress at the firewall, there is one place, and only one, that every prompt crosses before leaving. That is where to look.

What surprises people who have never done it is that this requires no code changes. The gateway speaks the same API as the provider. The application swaps the base URL and the key and keeps working. Claude Code, Cursor and the official SDKs do the same with an environment variable. Our tutorial on routing every AI request of the company through the gateway walks through it, with the firewall checklist at the end.

## What to do when the gateway finds an identifier

There are three possible actions, and the choice is per data category, not per person. Masking replaces the identifier with a placeholder before the prompt leaves, and the model keeps working for summaries, classification or drafting, because it never needed the number. Blocking returns an error with the reason, for categories that should never leave. Approval holds the request in a queue until someone decides.

Approval has one detail that matters a lot to privacy teams: the approver never sees the prompt. They see the codes for what the rules found and the hash of the content. The identity of whoever decided is mandatory and recorded. It is segregation of duties applied to AI usage.

- Category: Tax ID, national ID, email, phone | Action we usually recommend: Mask | Why: The model does not need the number to do the job. The text goes, the data does not.
- Category: Card number, password, API key | Action we usually recommend: Block | Why: There is no legitimate use case for these leaving in a prompt.
- Category: Confidential term (project, client, deal name) | Action we usually recommend: Human approval | Why: It depends on context. Someone at the company decides, and the decision is recorded.
- Category: Attached file | Action we usually recommend: Same policy, with a size limit | Why: Spreadsheets and PDFs are where identifiers show up most. Binary content without analysis is blocked.

## Test before switching it on

The policy screen has a test button that runs the real engine over a sample text without writing anything to the trail and without creating an approval. It shows the decision, the masked text exactly as the provider would receive it, the detected categories and the guardrail signals. Our advice is to run it with snippets that resemble the company's real documents, with fictitious data, before turning the key. You can adjust the mask style and the action per category while looking at the result.

## What remains as evidence

Every request produces an event in the audit trail: who made it (virtual key, person, team), which model, the policy decision, the count of personal data masked per category, in the form cpf:2,email:1, and a pseudonymized hash of the content. The prompt text never enters the trail. The trail is hash-chained and signed, so editing one event breaks the chain.

The difference between this and a policy in a PDF is the sentence you can say to the auditor or the regulator. Instead of we have an AI usage policy, you say something like: in September, this many prompts contained a tax ID, none reached the provider in clear text, and here is the verifiable record. The compliance panel organizes that evidence article by article for LGPD, and the same records serve a GDPR article 32 conversation.

## Where this touches LGPD

- Article 5, item I: a tax ID is personal data, because it identifies the person. There is no debate here.
- Article 6, item III, necessity: if the model does the job without the number, sending the number is processing beyond what is needed.
- Article 46: the controller must adopt technical measures able to protect the data. Masking at the network edge is a technical measure; training is an administrative one. Both count, but only the first is verifiable.
- Article 48 and ANPD Resolution 15/2024: a relevant incident must be reported within three business days. If you do not know what left, you cannot even assess whether there was an incident.
- Article 33: a prompt going to a server outside Brazil is an international transfer. Masking first reduces what is being transferred.

## What this does not solve

Pattern detection catches tax IDs with and without punctuation, validating the check digits, and the same goes for company IDs, card numbers and emails. It does not catch the client from the São Paulo contract we discussed yesterday. Semantic guardrails and a local judge exist for that, and they help, but nobody should promise 100% on free text. What we can promise is the combination: minimize what leaves, record everything that passed and, for what cannot leave at all, run the model inside the network.

To see it running: The demo at atlasberg.com/demo plays a live DLP scene: a prompt with a tax ID goes in, what the provider would receive shows up masked, and the event lands in the trail. To talk about your company's case, the path is atlasberg.com/contato.

## About Atlasberg

Atlasberg builds the control layer between a company and every AI model. Atlasberg Platform authenticates each request with a virtual key, filters it by policy and DLP, routes it to the right provider and writes it to a hash-chained audit trail, with an OpenAI-compatible API so applications only swap the base URL.

Articles on this blog are written by the engineering team and report measurements on real traffic, with the premises printed next to the result. To talk to the team, write to contact@atlasberg.com or use https://atlasberg.com/contato.

## More from Atlasberg

- [All articles](https://atlasberg.com/blog)
- [Tutorial: every AI request through your gateway](https://atlasberg.com/docs/tutoriais/gateway-da-empresa)
- [Platform overview](https://atlasberg.com/)
- [Talk to engineering](https://atlasberg.com/contato)

---
Source: https://atlasberg.com/blog/como-impedir-que-um-cpf-saia-no-prompt
Company: Sobimann Tecnologia da Informacao LTDA (Atlasberg), Porto Alegre, RS, Brazil.
Contact: contact@atlasberg.com - answered within one business day.
Index for agents: https://atlasberg.com/llms.txt
