Should you mask or tokenize before an agent reads the data?
Tokenize before ingest, when the reader is an agent. One that acts on its own. Mask if the reader is a human on a BI dashboard. One whose role you truly trust. The rest of this piece walks through why.
Snowflake Dynamic Data Masking is a query-time, role-based control. The plain text still lives in the table, unmasked. It comes back in full to the roles you allow. Masking is reversible inside the platform, by design. The raw identifier stays within reach in the query path. That's fine when you trust the role.
An agent is a role you can't fully trust. The OWASP Top 10 for Agentic Applications came out December 9, 2025. It carries the 2026 label. It ranks ASI01 Agent Goal Hijack as risk number one. Bad content can hide in a document or a RAG result. It can rewrite what the agent is trying to do. Say the hijacked agent's role can read plain text. Then it can leak plain text. Tokenization pulls the private value out of that path first. A hijack then leaks tokens, not the people behind them.
How Snowflake dynamic data masking actually works
Dynamic Data Masking is a column-level control. You attach a masking policy to a column. Per Snowflake's docs, that policy runs at query time. It fires at each spot the column shows up. Projections, WHERE, JOIN, ORDER BY, GROUP BY, all of it. It rewrites the result set, not the stored bytes. The plain text in the table never moves.
What comes back depends on who is asking. Again from the docs. The masking rules play a part. So does the SQL context, and the role hierarchy. A query user may see the plain value. Or a part-masked value. Or a fully masked value. Allowed roles see the clear value in full.
Two things follow from that. One, masking is cheap and built-in. It's the right way to keep a low-rights human analyst from seeing SSNs. Plain and simple. Two, it reverses by role, by design. So it does nothing to shrink what a special reader can pull. An agent connects under a special service role. That agent is a special reader too. That is just where the risk sits.
One more wrinkle worth naming. DDM masks single column values. It does nothing about re-identification through joins or totals. So a masked set can still leak who someone is through linkage. Snowflake pairs DDM with aggregation and projection policies for that exact reason. Masking alone does not stop a linkage attack. Treating it like it does is a common mistake.
How tokenization is different, and what Snowflake tokenization is not
Tokenization swaps a private value for a surrogate. That surrogate means nothing on its own. The mapping lives in its own vault. Only whoever holds that vault can get back the real value. Masking changes a value for display, while the source stays put. Tokenization takes the real value out of the dataset. It parks the key somewhere else. That move is the whole difference that matters.
Vaulted tokens are not the only shape here. Format-preserving shapes keep joins intact too. Format-preserving encryption (NIST SP 800-38G) keeps a field's format and type intact. Column joins and schemas survive, just as they do with deterministic tokens. The trade-off is where the secret sits. FPE builds the surrogate from a key, not a lookup vault. So you give up the key-move edge. The secret that reverses each value lives with the math itself. It does not sit in an outside store. One you could lock down, rotate, or destroy on its own. Moving the key is why you tokenize. A vault gives you that edge. FPE cannot match it.
Snowflake does ship a tokenization feature. It pays to be exact about what it is. Per the docs, External Tokenization lets accounts tokenize data first. This happens before loading it into Snowflake. It can then detokenize the data at query time. Detokenization runs through a masking policy. That policy calls an outside function, to a tokenization vendor. Wiring up a vendor needs Enterprise Edition or higher. The round trip happens at query time. So it adds delay. It also ties query uptime to the vendor's uptime. That is a big reason. The detokenize path stays inside the query engine.
There's a catch. Detokenization happens at query time, for allowed roles. Both DDM and External Tokenization are built as column-level masking policies. So the clear value can be rebuilt inside Snowflake's query path. For the roles you allow. That beats raw plain text sitting at rest. But it is still the same trust question the moment someone reads.
Why the choice matters more for agents than for humans
A human analyst who gets phished is a bad afternoon. An agent that gets goal-hijacked is a program. One that follows a hacker's orders at machine speed. Against each row it can touch. The blast radius is not in the same league.
The published standards back this up. The OWASP Top 10 for LLM Applications 2025 came out in November 2024. It puts LLM01 Prompt Injection at number one. It splits direct injection from indirect injection. That is where the model reads outside content. It should not trust that content. It treats that content as orders. The agentic list then moves this. Goal hijacking, driven by prompt injection, gets its own top slot. It's ASI01 Agent Goal Hijack. One small point worth keeping honest. OWASP kept LLM01 on the LLM list. It reframed the agent version as ASI01. It's a reframe across two lists, not a merger.
Taken together, the picture is plain. An agent treats fetched content as orders. It can be talked into leaking whatever it can read. Masking that reverses for the agent's role has a cost. It leaves raw PII in the query path. So a hijacked agent walks it out the front door. Tokenize before the data reaches the agent. Then the leak channel carries surrogates instead. Same attack, much smaller hole.
EchoLeak: the incident that makes this concrete
If this still reads as theory, meet EchoLeak (CVE-2025-32711). Aim Labs found it. Microsoft rated it severe, at CVSS 9.3. It shipped a server-side fix in May 2025. The bug went public on June 11, 2025.
The trick is simple. One crafted email carries hidden orders. Microsoft 365 Copilot pulls that email into its context. It runs the orders. It then sends private data out. Chat, OneDrive, SharePoint, Teams, all to a hacker's server. Zero clicks from the victim. Aim Labs named this class of bug LLM Scope Violation. That's their own term for it. Other reporting backed the report. A second study calls EchoLeak the first real-world zero-click prompt-injection exploit. It happened in a live LLM system.
Keep the label right, because it matters. This was research, shared the right way. Microsoft found no sign it was used in the wild. It's still the cleanest proof of the point here. An agent can hold reversible read rights to private data. It can get talked into a leak. Had that data been tokenized at ingest, the story changes. The same channel then carries only surrogate tokens. And yes, this is a technical threat model. It is not a full security program. Treat EchoLeak as one control's worth of lesson. Not the whole course.
Snowflake dynamic data masking vs tokenization: a quick decision table
- Human BI, trusted role, needs the real value sometimes: use Dynamic Data Masking. It's cheap and built-in, right for the job.
- An agent reading at scale, on its own: tokenize first. Do it before the agent sees it. Don't lean on a control that reverses by role.
- You need equality joins and
GROUP BY, but never the raw value: use deterministic tokens instead. Same input, same token. Joins and grouping survive. - You need
LIKEsearch? Or range and sort order on the real value? Neither control gives you that. Deterministic tokens hold equality, not substring or order. Design the query pattern first. - Regulated erasure (GDPR right to be forgotten): hold the key yourself. You can crypto-shred. Destroy the key, and the token can't turn back.
One legal note, so nobody gets it wrong. NIST SP 800-188, finalized September 2023, is the federal map for de-identification. It covers pseudonymization, quasi-identifier generalization, quasi-identifiers, and re-identification risk. Under GDPR Article 4(5), this still counts as pseudonymized data. That's true for reversible tokenization with a vault held apart. It is still personal data. The vault mapping is that extra information, kept apart. It is the piece the rule points at. Tokenizing is not the same as anonymizing. Don't let anyone tell you otherwise.
How to secure PII before an agent reads it: best practices
In practice, the steps look like this:
- Classify at ingest, not at query time. Decide which fields are identifiers before anything lands in your data store. This is where quasi-identifier generalization earns its keep. Coarsen quasi-identifiers: birthdate to birth year, ZIP to region. Then a token plus nearby context can't quietly re-identify someone.
- Tokenize the direct identifiers with deterministic, join-preserving tokens. Hold the vault outside the query path. The agent should never have a code path to the clear value.
- Keep the vault on your side. Say detokenization is reachable from the same query engine the agent uses. Then the clear value is back within the agent's reach.
- Authorize per tool call, not per session. A long-lived agent session is a long-lived hole. Scope each call. Keep the power to revoke mid-session.
- Govern the MCP boundary. The MCP Security Best Practices sit alongside the June 2025 authorization spec. They are blunt: "MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server," and token passthrough is flatly forbidden. It also covers the confused-deputy attack on proxies. Plus SSRF during OAuth discovery, and session hijacking. It wants per-client consent, exact
redirect_urimatching, and least-privilege scopes. Injection does not only ride in on user input, either. Tool poisoning hides orders in tool descriptions. In practice, the bad content often arrives in the tool result itself. - Log every read into a tamper-evident chain. Then, after an incident, you can prove just what the agent touched.
A sibling piece walks the ingest pipeline in more detail. See tokenizing PII before it reaches an LLM.
Govern the data plane, not just the prompt
Most agent-security tooling watches what the model says. Prompt filters. Output guards. Jailbreak checks. Useful, sure, but all of it is a guess at intent. The sturdier move is to govern what the agent can reach. Then a fully hijacked agent still finds nothing worth stealing.
That's the idea behind DataShield's Ontology: govern at ingest. PII gets tokenized into Parquet. It's served over MCP, through the DataShield Analytical DB. A vault you hold sits behind it. The tokens are deterministic and join-preserving. Quasi-identifiers get coarsened. Crypto-shred erasure covers the right-to-be-forgotten case. A hijacked agent reading that store sees only surrogates.
Here's the plain framing, so buyers don't get surprised. DataShield does not rewrite your Snowflake tables in place. It isn't a Snowflake masking policy. It isn't a prompt or proxy filter sitting in front of the model. It governs at ingest. Or at the MCP boundary, where per-call authorization and mid-session revocation live. Each call seals into a tamper-evident audit chain you can check at /verify. And no, there's no SOC 2 attestation yet. When there is, we'll say so plainly.
The honest limits
Tokenization isn't a free lunch. Any pitch that claims otherwise is overselling it.
Deterministic tokens preserve equality joins and GROUP BY. They don't preserve LIKE/substring search. Nor range and order queries on the source value. Say your agent workload truly needs fuzzy search over raw names. Then tokenization at ingest will fight you. Better to plan around that up front. Don't find out in production at 2am. Masking keeps those query behaviors. The real value still sits right there. Which is just why it's also still within reach.
Snowflake's surface keeps moving. Check the current edition rules. Check whether Horizon governance features have shifted things. This comparison may differ by the time you see it. The core point outlasts the product renames. A control can hand back clear values to an allowed role. That has one gap. It cannot protect you from a reader you can't trust. An agent is that reader. Pick the control that matches the threat. If the reader acts on its own, act on it. Get the secret out of the query path. Want to scope this against your own data? The architecture overview and a quote are the next stops.
Watch: related explainers
Snowflake dynamic masking in practice. Plus the agent-side risks that make tokenization worth it.
Frequently asked questions
Does Snowflake dynamic data masking change the stored data?
No. Per Snowflake's docs, the masking policy runs on the column at query time. It fires at each spot the column appears. It rewrites the result set. The plain text still lives in the table, unmasked. It comes back in full to allowed roles. Masking is a display-time control that reverses by role. It never changes the stored value.
Is tokenization reversible?
It depends on the vault. Reversible tokenization keeps a mapping in a vault, held apart from the data. An allowed party can use it to get back the real value. Under GDPR Article 4(5), that makes the data pseudonymized, not anonymized. Destroy the key or the vault entry, and the token can't turn back. That crypto-shred is how tokens support a right-to-erasure request.
Can a masked column still leak PII to an AI agent?
Yes, if the agent's role is allowed to see the clear value. Masking hands back full plain text to allowed roles. So a hijacked agent, reading under a special service role, can leak it. The raw ID. OWASP ASI01 Agent Goal Hijack (December 2025) makes this point. So does the EchoLeak proof of concept. Both show agents talked into leaking whatever they can read.
Does deterministic tokenization break my SQL joins?
No. Deterministic tokenization gives the same token for the same input. So equality joins and GROUP BY on the tokenized column still work. What it does not keep: LIKE/substring search. Or range and order queries on the source value. Confirm your query patterns before you tokenize a field.
Is Snowflake External Tokenization the same as masking?
They are built the same way, as column-level masking policies. But the goal differs. External Tokenization lets you tokenize before loading. You then detokenize at query time, through an outside vendor. It needs Enterprise Edition or higher. Allowed roles can detokenize at query time. So the clear value can still be rebuilt inside Snowflake's query path. Same as with masking.