Tokenizing PII at ingress in four steps
The pattern runs in four steps. Order matters here.
- Detect. Before you build a prompt, scan the text for private spans. Microsoft Presidio is the standard open-source tool for this. Its Analyzer pairs spaCy or Hugging Face NER. It adds regex checks and context scoring. It finds names, card numbers, emails, and the rest.
- Replace. Swap each found value for a surrogate token. Presidio's Anonymizer ships operators built for this (
replace,redact,mask,hash, and a reversibleencrypt). A value scrubbed on the way in can come back. It returns later, on the way out. Write-ups on agent design call this same swap step the PII tokenization pattern. - Build the prompt from tokens. Everything downstream runs on tokens, not plain text. The model. The RAG index. The tool calls. The logs. Plain text never travels with any of them.
- Detokenize on the way out, and only sometimes. Say an allowed output truly needs a real value. Resolve the token behind a tightly controlled wall. Most outputs never need it.
The idea under all four steps is simple. A model cannot leak a secret it was never handed. Keep the model away from data worth stealing. Then its trustworthiness matters far less.
One honest limit, before you get comfortable. The whole scheme is only as strong as step one. Detection is not a solved problem. Presidio's own analyzer docs say its checks are probabilistic. They will miss things. Anything the detector fails to catch reaches the model in plain text. Ingress tokenization is bounded by detector recall, full stop. FPE handles structured fields cleanly. The weak spot is loose, unstructured text. Detection is least reliable there. That's also where most of your real risk lives. So treat detection as the main risk surface. Not a box you check once. Watch your recall rate. Put your highest-risk fields on an allowlist. Then the values you can't afford to leak get tokenized by rule. Not by a model's best guess.
What tokenization actually does to a sensitive value
Tokenization swaps a private value for a surrogate. That could be a payment PAN or a PII field. The surrogate means nothing on its own. You get back the source value only one way. Detokenize it against a protected map. That's the definition the PCI SSC Tokenization Product Security Guidelines work from. It sorts tokens into three kinds.
- Irreversible. No path back to the source value. Great when you never need the real value again. Think analytics, or joins. Useless when you do.
- Reversible vaulted. The map lives in a lookup table or vault. You detokenize by asking it. Easy to reason about. But the vault is a hot target. It is a read-time dependency you now own.
- Reversible vaultless. The token is a cipher output. It's computed from the plain text plus a key. There's no stored table to break into. The value gets rebuilt from the token, plus the key. That happens only when you're allowed to see it.
For LLM pipelines, reversible vaultless is often the sweet spot. You keep the ability to restore values. You skip parking a giant target table next to your model's path.
FF1, FF3-1, and why format-preserving encryption fits prompts
One trick makes vaultless tokenization behave inside a prompt: format-preserving encryption (FPE). A 16-digit number stays 16 digits. An email still looks like an email. That matters. You can drop a token into a prompt or a structured field. It won't trip length checks. Nor schema checks, nor the model's guess at what it's looking at.
The standard is NIST SP 800-38G, published March 2016 and updated 4 August 2016. It sets out two AES-based Feistel modes, FF1 and FF3. Each takes a non-secret "tweak" that scopes the encryption to one context.
The tweak does two jobs. Most people miss the second one. Hold it steady, and FPE becomes deterministic. The same plain text always lands on the same token. That's not a footnote. It's what keeps your pipeline working. Deterministic tokens, under a stable tweak, mean joins still join. Coreference still resolves. RAG retrieval still finds the right chunk. One name becomes the same surrogate everywhere it shows up. Vary the tweak per context, and you split domains apart. A token from one system can't be matched against another. Hold it steady within one context, and you keep that link intact.
Format-preserving does not mean meaning-preserving. That's the catch nobody prints on the box. FPE keeps the shape and wipes out the meaning. Tokenize a field the model must reason over. Its reasoning breaks. A tokenized birthdate can't do age math. A tokenized ZIP can't tell the model which region it's in. A tokenized dollar amount can't be compared or added up. Tokenizing everything has a real cost. The safe targets are names, card numbers, and emails. Account IDs, SSNs, and other standalone PII fit too. Fields the model must compute on are better left in the clear. Or generalized: an age bucket instead of a raw birth date. Not turned into noise the model then reasons over badly.
One more catch you don't get to skip. After codebreakers cracked FF3, NIST cut the tweak from 64 to 56 bits. It renamed the fixed mode FF3-1. Draft SP 800-38G Revision 1 was out for public comment in 2019. It proposed a floor too. FF1 and FF3-1 need a domain of at least 1,000,000 values. So use FF1 or FF3-1, never plain FF3. Treat a million possible values as the floor. Stay above it before you tokenize a field. Run FPE over a two-letter US state code, and you're asking for trouble.
Prompt injection and the lethal trifecta
Tokenization matters, because the model cannot tell your orders from a hacker's. The OWASP Top 10 for LLM Applications 2025 came out as v2.0. It was announced 17 November 2024. It keeps LLM01 Prompt Injection at #1. It keeps LLM02 Sensitive Information Disclosure at #2, too. The root cause OWASP names runs deep. LLMs read orders and data through the same channel. There is no clean split between them. So crafted input gets read as a fresh order. It comes in two flavors. Direct injection is typed into the prompt. Indirect injection hides in a document, web page, or email. The model reads it later, and trusts it. OWASP is blunt that neither RAG nor fine-tuning fully fixes it. So you need layers of defense.
Simon Willison sharpened the agent version of this. He did it on 16 June 2025, with the lethal trifecta. Give an agent three things. Access to private data. A path to unsafe content. And a way to reach the outside world. Prompt injection can then drive it to leak data. No exploit code needed. His take on guardrail products that block "95% of attacks": 95% is a failing grade. The hacker just aims at the other 5%.
This is where tokenization earns its place. By removing the private-data leg, it does one thing. It leaves a hijacked agent that only ever sees tokens. Nothing worth stealing. You never have to win an impossible fight. Sorting each input perfectly into trusted or unsafe is not required. You have already made the payload worthless.
What this looks like in the wild: Samsung and EchoLeak
Keep two buckets apart. Things that truly happened. And things researchers proved could happen.
In the wild. Samsung Semiconductor lifted its own ChatGPT ban in March 2023. Within about 20 days, three engineers pasted secret material straight in. Source code from a faulty facility-measurement database. Code for spotting bad equipment. A recorded staff meeting they wanted turned into notes. Samsung then banned generative-AI tools on company devices, a move reported 2 May 2023. The specific use cases trace back to Korean-press reporting. Treat the details as attributed, not gospel. This is the textbook case of PII and IP strolling into an LLM. Ingress tokenization is built to catch just this.
Proof of concept. In June 2025, Aim Labs revealed EchoLeak (CVE-2025-32711) in Microsoft 365 Copilot. One crafted email, zero clicks from the user. It could make Copilot reach into OneDrive, SharePoint, Teams, and chat. It could then send the contents out. The researchers named the trick "LLM Scope Violation" and rated it CVSS 9.3. Microsoft fixed it server-side. By its own account, it found no sign of use in the wild. That keeps it firmly in the research column. It is a clean demo of one chain. Unsafe content, then private data, then a leak. Data-plane tokenization blunts that middle link. Had that data been tokenized, the story changes. The leaked payload would be a bag of surrogates.
PII tokenization before the LLM: best practices
If you're building this, here's the short list I'd hold a design to.
- Tokenize at ingest, before the model ever sees it. The earlier you swap values, the better. Fewer copies of plain text survive in logs, caches, and vector stores. Presidio's reversible
EncryptandDecryptoperators give you the restore path. This also matters for another reason. One tokenization vendor notes this. A third-party API logs and keeps raw PII you send it. You want it gone before it's sent. - Use FF1 or FF3-1, and mind the domain size. Prefer FF1 or FF3-1 over plain FF3. Follow the draft million-value floor proposed in SP 800-38G Revision 1. Never run FPE over small domains, like state codes.
- Tokenize identifiers, not the fields you need to reason over. FPE keeps format, not meaning. Tokenizing dates, geography, or number ranges quietly hurts the model's output. Tokenize standalone PII and identifiers instead. Leave computable fields in the clear, or generalize them.
- Keep the map out of the model's reach. The token-to-value store lives behind a wall. The model, and its tools, can't touch it. If the inference path can query it freely, you have a problem. You've rebuilt the weak central store you were trying to avoid.
- Detokenize sparingly, and only with approval. Apply MFA, least privilege, and rate limits on the resolve path. Most outputs never need real values. Keep plain text as an exception you explicitly approve.
- Assume pseudonymized data is still regulated. Under GDPR Recital 26, data stays personal data. That holds even when you can restore it with a separate key. And NIST SP 800-122 treats it as PII short of true de-identification. Tokenizing before the LLM shrinks risk. It does not erase your duties on the mapping store. Govern that store with care. Treat it like the data it stands in for.
One honest caveat. This is a technical threat model. It is not a full legal program. Ask your own counsel what "de-identified" means for your data.
How to secure the detokenization path
A detokenization feature is the riskiest surface in this design. You may expose it as a tool an agent can call. The moment you do, things change. It becomes just what hackers aim at. So it inherits every control you would put on the tool itself.
Say you expose detokenize over the Model Context Protocol. Then the MCP Security Best Practices aren't optional. The server MUST check each inbound request. It MUST NOT use sessions for login. Token passthrough is forbidden. An MCP server must reject any token not issued to it directly. That is what stops confused-deputy attacks. In those attacks, one service gets tricked into acting with another's rights. Add least-privilege scope limits, per-client consent, and non-deterministic session IDs. Also block SSRF: private and link-local ranges, plus the cloud metadata address 169.254.169.254. These are the same MCP data-plane risks that recent security surveys keep flagging. A detokenization tool sitting behind anything weaker is a data-leak tool in disguise. If you're standing up MCP servers, our MCP server security walkthrough goes deeper.
Govern the data plane, not just the model
Most "AI security" money goes to shaping what a model will do. Alignment. Refusals. Output filters. Useful, but it's the wrong plane to bet on alone. Prompt injection keeps winning. The durable move is different. PII tokenization before the LLM works by governing what agents can reach.
That's the idea behind DataShield. Tokenize private fields at ingest. A hijacked agent then finds tokens, not raw PII. Authorize each detokenization per tool call, with mid-session revocation. A grant you regret dies at once. It does not wait for the session to end. Each call gets sealed into a tamper-evident audit chain. You can check it yourself at /verify. The ontology decides which fields count as private. The architecture shows where the trust boundary sits. And /auth is where the per-call authorization lives.
It's the tokenize-on-ingress, detokenize-on-egress pattern. The map sits behind a wall the model, and its tools, can't reach. It lines up with OWASP's LLM02 fixes. Input and output cleanup. Least-privilege data access. And it removes the private-data leg of the lethal trifecta. You stop arguing with the model about which inputs to trust. You make the data itself boring to steal. Want to scope it to your stack? Start a quote.
Watch: related explainers
How to keep private data out of the prompt, plus the MCP-security context.
Frequently asked questions
How do you tokenize PII before it reaches an LLM or AI agent?
Detect private spans at ingress. Microsoft Presidio's Analyzer, with NER plus regex, is the standard tool. Replace each value with a format-preserving surrogate token. Compute it with a keyed FPE mode (NIST FF1 or FF3-1). Build the prompt from tokens. Keep the token-to-value map behind a wall. The model and its tools can't reach it. Detokenize only allowed outputs, under MFA and least privilege. One limit to design around: detection is a guessing game. Anything it misses reaches the model in plain text. Tokenization is bounded by detector recall. So you watch recall and allowlist high-risk fields.
Is tokenized or pseudonymized data still considered personal data?
Yes. Under GDPR Recital 26, this still counts as personal data. That's true even so. A separate key can still bring the source value back. NIST SP 800-122 treats it as PII unless it meets full de-identification rules. Tokenizing before the LLM cuts risk. It does not remove your duties on the mapping store. So govern that store tightly.
Does tokenization stop prompt injection?
Tokenization leaves the injection itself possible. But it takes away the payoff. OWASP ranks prompt injection as LLM01 for a reason. Models read orders and data in one channel, with no clean split. Neither RAG nor fine-tuning fully fixes it. Tokenization removes the private-data leg of Simon Willison's lethal trifecta. So a hijacked agent only ever sees tokens. It finds nothing worth stealing.
Should I use FF1 or FF3 for format-preserving encryption?
Use FF1 or FF3-1, never plain FF3. Codebreakers cracked FF3. In response, NIST cut its tweak from 64 to 56 bits. It renamed the fixed mode FF3-1. Draft SP 800-38G Revision 1 also set a minimum domain size of 1,000,000. So don't apply FPE to small domains, like state codes. Keep the tweak stable within one context, so tokens stay deterministic. Your joins, coreference, and RAG retrieval still work.
What controls does a detokenization tool need if I expose it over MCP?
Per the MCP Security Best Practices, the server checks each inbound request. It must not use sessions for login. It must not accept token passthrough. Add least-privilege scope limits. Add per-client consent, to stop confused-deputy attacks. Block SSRF, too: private and link-local ranges, including 169.254.169.254, plus non-deterministic session IDs.