Do Purview labels and DLP protect PII before a Fabric Copilot reads it?
They protect access to it, not the value itself. That gap is worth spelling out.
A Purview sensitivity label is metadata. Microsoft says so plainly. Labels sit in clear text. That's in the metadata of files and emails. The label travels with the content, no matter where it lands. It describes the item. The credit card number underneath it is still a credit card number.
DLP for Microsoft 365 Copilot is the enforcement arm. Point the Content contains > Sensitivity labels condition at something. It then excludes matching items from processing. Per Microsoft, Copilot skips that item's content. It won't process it, or use it in its answer. The item can still show up in citations. It is a detect-and-block gate. It transforms nothing and removes nothing. Each Copilot prompt runs as the user who typed it. So an allowed user can pull the raw value. Copilot just acts on their behalf. Unless a rule kicks that item out first.
So here's the honest answer. Labels and DLP gate reads. They can encrypt or block whole items. They leave the raw PII right where it was.
What a Fabric data agent can actually read
A Fabric data agent is generally available, read only. It runs under the requesting user's credentials and permissions. That's a least-privilege basis. It respects Purview governance on the underlying sources. DLP and access restriction policies count too. That same documentation notes DLP for Fabric Data Warehouse is generally available. Access restriction policies cover the Fabric KQL Database, SQL Database, and Data Warehouse. Those were still in preview. That was true when this was written. Microsoft has stated a path here. Restrict-access policies reach GA, rolling out through late 2026. Microsoft ships fast here. So check the exact status first. GA or preview, before you bet a control on it. That preview label is time bound. The same data can show up through Copilot in Fabric. Those Purview policies still apply there too.
Microsoft's own guidance gives a warning here. Sensitive assets might be out of reach for the agent. That can produce incomplete answers. Look at what that quietly assumes. The raw IDs live in OneLake. The control only decides whether the agent can reach the asset. Block the asset, and you block the answer. Allow it, and the agent reads the real values. No version of this works without the raw PII present. Not if you want a useful answer. The Purview-over-Fabric model is Entra identity plus label-based and authorization-based policy. It sits on top of data that lives, in clear form, in OneLake.
Microsoft Fabric sensitivity labels vs tokenization: the mechanism difference
A label and a token do different things to a value. A label is metadata added on top of a value. That value stays exactly where it is. Tokenizing takes the value out. It drops a stand-in token where the value used to sit. Everything downstream follows from that difference.
Tokenize PII at ingest. The raw ID then never lands in OneLake at all. A token swaps in for it instead. The real value stays in a customer-controlled vault, on your own instance. The agent queries governed data and gets back tokens. It can still join customer_9f83 across the warehouse. That's because deterministic tokens preserve joins. It cannot read the person's name, because the name isn't there. Say you need the raw value later, for some legitimate downstream flow. You resolve the token through the vault, under its own authorization. That happens outside the agent's blast radius.
Microsoft's closest native answer here is dynamic data masking in Fabric Warehouse. Add OneLake column-level and row-level security too. Worth knowing, worth using. But understand the mechanism. DDM masks the value at read time for unprivileged callers. The raw value stays in place. It's reversible, in-place masking. So anyone with the right permission still gets the real value. Straight out of OneLake. So does an agent running as them. Same shape of gap as a label. It just sits at the column instead of the item.
So labels, DLP, and native masking change who may read the ID. Tokenization changes something else. It changes whether the ID exists at all. Does it sit in the plane the agent can reach?
One honest caveat on tokens. I don't like specs that oversell. Deterministic tokens preserve equality joins, not LIKE or substring matching. Partial-match queries against tokenized columns change behavior. They won't work the way they do on cleartext. That's a design tradeoff, not a bug. You plan around it.
Key custody decides the whole thing. Say the token keys sit with the same cloud tenant as the data. Then you've moved the problem, not solved it. Customer-held keys are what let you crypto-shred. Destroy the key, and each token from it becomes unresolvable for good. That happens to be a clean answer to an erasure request.
The threat model: goal hijack and EchoLeak
Why does this matter more for an agent than for a dashboard? Agents can be steered. A dashboard cannot.
The OWASP Top 10 for Agentic Applications treats this as a top-tier risk. Agent goal hijack, named plainly. It's formally coded Agent Goal Hijack (ASI01), in the December 2025 list (the 2026 edition). That list folds in classic prompt injection. That was long the top LLM risk, known as LLM01. It adds excessive autonomy too. Both fold into one agent-specific type. A single bad input can redirect the agent's goals. It can hijack its planning too, and its multi-step tool use. Not just one corrupted reply. Want the exact ordering and type names? Pull them straight from the OWASP page. Not from any summary, including this one.
The canonical demo is EchoLeak, CVE-2025-32711. It's a critical (CVSS 9.3) zero-click indirect prompt injection in Microsoft 365 Copilot. Its discoverers at Aim Labs classed it an LLM Scope Violation. A single crafted email tricked Copilot into leaking data from its own context. The attack needed no click. The type matters here. It was a proof-of-concept under responsible disclosure, not a breach in the wild. Per that Aim Labs writeup, the team built it in January 2025. They reported it to Microsoft. Microsoft patched it server-side in May 2025. That same writeup notes no proof of real-world exploitation. There is also an arXiv preprint. It was later published at the AAAI Fall Symposium Series 2025. It documents the exploit as zero-click prompt injection. Shown in a live LLM system. Shown, to be clear. Not caught in the wild.
EchoLeak hit M365 Copilot broadly, not Fabric specifically. I haven't seen a documented in-the-wild PII exfiltration tied to Fabric data agents. The point is the pattern, not a Fabric body count. Say an agent's goals get hijacked. Then only one kind of data stays safe. Data the agent cannot read in raw form. Say the card is labeled, but still there in raw form. An allowed agent can still read it, once someone hijacks that agent. A token hands that agent nothing worth stealing.
Why the MCP boundary is the right place to enforce this
More Fabric and Copilot work is moving through tool calls. The Model Context Protocol is how a lot of it gets wired. MCP defines Hosts, Clients, and Servers. They expose Resources, Prompts, and Tools over JSON-RPC 2.0. Credit where it's due. The spec is refreshingly honest about its own limits. It states this outright. MCP itself cannot enforce security at the protocol level. Implementors SHOULD build consent and authorization flows. They should add proper access controls and data protections too. The MCP security best practices doc goes further still. It covers indirect prompt injection and confused-deputy risks.
So the protocol will not save you on its own. Governance has to live at the boundary. Enforce it at ingest, before data reaches OneLake. Or enforce it at the MCP boundary, where the agent asks for data. That is exactly where per-call authorization belongs. The system checks each tool call against live entitlements. You can revoke it mid-session too. That beats trusting an old token. One minted twenty minutes ago, never checked since.
Fabric PII governance best practices for AI agents
This is an engineering threat model, not a compliance checklist. Treat it as a starting frame, not full advice. Here is what I would actually do.
- Keep labels and DLP. They're good at classification, discovery, encryption, and blocking whole assets. Don't rip them out. They just aren't a data-removal control. Stop treating them like one.
- Tokenize the high-risk IDs at ingest. Names, national IDs, card numbers, emails, account numbers. Do it before the data lands in OneLake. Then the agent plane never holds the raw value.
- Hold your own keys. Customer-controlled KMS or HSM. Without that, crypto-shred and real erasure are just theater.
- Authorize per tool call, not per session. Bind each MCP call to current entitlements. Support mid-session revocation too.
- Log every call into a tamper-evident chain. You want to prove what the agent read, not just assert it. This is also where your Article 12 logging story lives.
- Pick your enforcement point on purpose. Choose ingest, the MCP boundary, or both. Don't assume the model or the protocol will do it for you. Neither one will.
Govern what agents can reach, not only what they will do
Most agent security spends its budget on behavior. Guardrails, output filters, refusal training. Useful stuff, but it's all guessing at intent. The deeper lesson from Microsoft Fabric sensitivity labels vs tokenization is this. Shrink what the agent can reach. That is the cheaper move, and the more durable one.
That is the plane DataShield works in. Its Ontology applies join-preserving tokenization at ingest. A hijacked agent then turns up tokens instead of raw PII. Authorization happens per tool call, with mid-session revocation. Each call gets sealed into a tamper-evident audit chain. You can check it yourself at /verify. One honest limit, stated plainly. DataShield's SOC 2 attestation is planned, not attested yet. It isn't a masking policy bolted onto Copilot. It isn't a prompt proxy parked in front of the model either. It governs the data plane, at ingest or at the MCP boundary.
Fabric leans hard Microsoft-native. Entra identity, Purview governance, tight Copilot and agent integration. The authorization layer is already strong there. The gap isn't who can read. The gap is that the raw ID still sits there. You close it by removing the value, not by relabeling it. Want the numbers behind that tradeoff for your own data? The quote worksheet is where to start.
Watch: related explainers
How Purview DLP gates Copilot prompts. Plus the injection risk it leaves in place.
Frequently asked questions
Do Purview sensitivity labels remove PII from OneLake?
No. A sensitivity label is metadata, stored in clear text alongside the item. Per Microsoft, it stays with the content no matter where it's saved. It can trigger encryption or access control. The raw ID still sits in OneLake in clear form.
Can a Fabric data agent read raw PII if the data has a sensitivity label on it?
Yes, if the querying user is allowed. And no DLP exclusion or access-restriction policy blocks that asset. Fabric data agents run under the requesting user's permissions. So the agent gets the raw value back. A control that blocks the whole asset blocks the answer too.
What is the difference between a sensitivity label and tokenization?
A label is metadata attached to a value. That value stays in place. Tokenization at ingest replaces the value with a join-preserving token. The real data then sits in a customer-controlled vault. In short: labels govern who may read the ID. Tokenization governs something else. Whether the ID is even there, for the agent to reach.
Does Purview DLP stop prompt-injection attacks like EchoLeak?
Not directly. DLP for M365 Copilot works as a block-and-exclude gate. It decides which items to process. EchoLeak (CVE-2025-32711, CVSS 9.3 Critical) was a zero-click indirect prompt injection. It was shown as a proof-of-concept and later patched by Microsoft. Its discoverers at Aim Labs noted no proof of real-world exploitation. The durable defense turns the value into a token. Then there is nothing raw left to steal, even from a hijacked agent.
Where should I govern PII for Fabric agents, at ingest or at the MCP boundary?
Ideally both. Tokenize high-risk IDs at ingest. That way they never land in OneLake in clear form. Enforce per-call authorization at the MCP boundary too. The MCP spec states the protocol cannot enforce security by itself. So implementors have to add access controls and data protections there.