Head-to-head · updated 13 September 2026
DataShield vs Google Dataplex: context for your agents, or control over them?
In April 2026 Google renamed Dataplex Universal Catalog to Knowledge Catalog, and the new name is honest about the plan. The product page calls it "always-on context and governance for your agents". If your data sits in BigQuery, it harvests the metadata with no connectors to run, draws column-level lineage you would struggle to rebuild by hand, writes descriptions with Gemini, and serves the lot to Claude or Gemini CLI from a remote MCP endpoint. Nothing to install. Nothing to procure.
We sell the next question. Which agent asked, was it still allowed at that second, and can you prove the record later? Datasets are tokenized at ingest; agents query tokenized data over MCP; detokenization is a privileged, audited operation. Every governed tool call in Auth is re-checked against the agent's current authority, and each decision is sealed into a hash chain you can verify without trusting us. On GCP, most buyers run both. Here is the honest split, including the rows Google wins.
The short version
Pick DataShield when
- Someone will one day ask you to prove an agent's access record wasn't edited. Cloud Audit Logs are retained and exportable. Ours are chained and signed, and the verifier names the failure. Run it yourself.
- The catalog has to run where Google does not. Self-hosted in your own cloud or data center, or a dedicated single-tenant server we operate, with your keys. See the architecture.
- You need to pull an agent's authority mid-session and have the very next tool call fail closed, not wait out a token. How Auth does it.
- You want PII and PHI labels as first-class catalog objects, not a separate scanning product. We classify columns against 129 field classes, and the catalog stores no column values by default. What Ontology holds.
Pick Google Dataplex when
- Your estate is BigQuery. Metadata from tables, views, models, Spanner, Bigtable, Cloud SQL, Looker and Pub/Sub arrives on its own. We cover PostgreSQL today, not your whole estate.
- You want column-level lineage for BigQuery and Vertex AI pipelines. Google reads its own query logs. No third party gets that depth from outside.
- Gemini in the catalog is worth money to you: generated descriptions, semantic search, inferred relationships, and profile scans over unstructured data.
- Procurement is the bottleneck. It is already on the GCP bill, the first 100 DCU-hours each month are free, and there is no new vendor to onboard.
Bottom line: Google gives your agents context. We decide which agent may act on it, and keep proof that the decision was made. On GCP, run both. Off GCP, or in front of an auditor, start here.
Feature by feature: a platform catalog against a governed data plane
Competitor cells describe what Google's public docs, release notes and pricing page say as of the date above. If we've mischaracterised something, email support@myorg.ai and we'll correct it, credited.
| What matters | DataShield | Google Dataplex | Edge |
|---|---|---|---|
| Audit evidence | SHA-256 hash chain with Ed25519-signed checkpoints that are themselves chained. Verification returns clean, attested damage, or tampered, and names the break. Try the verifier. | Cloud Audit Logs: Admin Activity plus Data Access logs you switch on. The audit-logging doc describes no integrity proof, no signatures and no immutability for the log itself. | ◆ |
| Agent authorization | Every governed tool call in Auth passes a scope ceiling, a consented-tool allowlist, a declared authority tier, and a revocation re-check before dispatch. Delegation is RFC 8693 token exchange with an enforced ceiling. | The MCP server uses OAuth 2.0 with IAM roles such as roles/mcp.toolUser. That tells the catalog which human or service account is behind the call. It does not give the agent its own authority, or re-check it. | ◆ |
| Break-glass | Scoped, time-boxed emergency access for agents. It auto-revokes and it cannot be quietly removed from the record. | Not offered. You would assemble it from IAM conditions and hope the rollback is clean. | ◆ |
| GDPR erasure | Crypto-shred of per-subject key material plus ISO 27560 consent receipts. Actor identities are HMAC-committed, so the chain still verifies after a subject is erased. | Not a stated function of the catalog. It records where the data is. Deleting it, and proving you did, is yours. | ◆ |
| Tokenization and data handling | Datasets are tokenized at ingest; agents query tokenized data over MCP; detokenization is a privileged, audited operation. Tokens are deterministic and join-preserving. Quasi-identifier generalization is there when you switch it on: dates to year, decade or age band, ZIPs to 3 or 4 digits, partial phones, SSNs and emails, with a measured cardinality-reduction score per column. | The catalog is metadata. De-identification is a different Google product (Sensitive Data Protection), billed and configured separately. Column-level security runs on BigQuery policy tags. | ◆ |
| Sensitive-data classification in the catalog | 129 field classes covering PII, PHI, financial data and secrets, including all 18 HIPAA Safe Harbor identifiers. Regex plus checksum validation (Luhn, NPI, Verhoeff, ABA, IBAN, GTIN) plus column-name lexicons and anti-pattern suppressors. No model, so verdicts are reproducible from a config digest. Column values are not copied into the catalog by default. | Data profiling and quality scans in the Premium tier, with Gemini-generated semantics. Sensitive-data labelling leans on Sensitive Data Protection and BigQuery policy tags rather than the catalog itself. | ◆ |
| Catalog coverage and lineage | We scan, profile and classify a live PostgreSQL source in place, with no rows leaving it. Typed lineage traversal with per-hop access gating, a glossary materialized from entity types, and stewardship worklists. Lineage is derived from the pipelines that own it, not a stored column-level graph. PostgreSQL today, not your whole estate. | Auto-harvested metadata across BigQuery, Vertex AI, Spanner, Bigtable, Cloud SQL, Looker and Pub/Sub, plus table- and column-level BigQuery lineage, business glossary, aspects, data products and Iceberg REST cataloging. This is their round, and it isn't close. | ◇ |
| Context for agents | Agents query the catalog and act on it through MCP. Ontology's catalog, stewardship and lineage surfaces are MCP-native, and Auth issues MCP tool tokens with scope ceilings for its own surface. | A genuinely good story. Remote MCP server at dataplex.googleapis.com/mcp, a data-products endpoint, lineage MCP in Preview, and a lookupContext bundle built for agent workflows. Gemini CLI, ChatGPT and Claude are all listed clients. | ◇ |
| MCP and agents overall | More than 200 MCP tools across Ontology, Auth, Corpus and Lighthouse, with per-call metering attributed to the agent and typed refusals when a quota bites. | MCP for discovery and data products, with lineage in Preview. No agent identity of its own, no per-call authority re-check, no metering attributed to an agent. | ◆ |
| Deployment | Self-hosted in your own cloud or data center, or on a dedicated single-tenant server we operate. Docker images ship for Auth, Ontology, Corpus and Lighthouse, and Guardian verifies a signed deploy manifest. | A managed Google Cloud service. No self-hosted option, no on-premises option. Third-party and on-prem sources arrive through connectors, not as first-class citizens. | ◆ |
| Maturity signals | Auth, Guardian and Lighthouse are live in production (Guardian and Lighthouse since April 2026). SOC 2 not yet certified, and we say so. | Alphabet. GCP compliance coverage, a release note most weeks, and Google's own claim that over 95% of its top analytics customers use Dataplex. Three product renames in four years is the one wobble. | ◇ |
| Pricing | Published model, scoped instant quote, no sales wall. | Published rates, metered in Data Compute Unit hours: $0.06 per DCU-hour Standard with 100 free each month, $0.089 Premium for lineage, quality and profiling with no free allowance, and about $2 per GiB-month for custom metadata. Cheap to start. Hard to forecast. | ◈ |
◆ DataShield leads◇ Google Dataplex leads◈ comparable
Google claims are drawn from cloud.google.com product, docs, release-notes and pricing pages, last checked 13 September 2026. We link them below rather than paraphrase from memory.
Three things you get here that you won't get from a platform catalog
Proof the record wasn't edited
A retained log is not a tamper-evident log. Ours is a hash chain with signed checkpoints, and the verifier tells you what went wrong: tampering, insertion, deletion or truncation. That is the property EU AI Act Article 12 and HIPAA §164.312(b) reviewers keep asking about. Try it in your browser, no signup.
Authority that changes mid-flight
An analyst resigns on a Friday afternoon. Their agent is 20 minutes into a 40-minute job. With DataShield the next governed tool call is re-checked against current authority and fails closed. An IAM change is eventually consistent, and the call in flight rarely notices. How Auth does it.
A catalog that can leave the cloud it governs
Knowledge Catalog is a Google Cloud service, and it is not going anywhere else. If your estate spans clouds, or your regulator wants the governance layer outside the provider's blast radius, the catalog has to be yours. Ours runs in your data center or on a single-tenant server we operate. See the architecture.
Where Google Dataplex is genuinely stronger
Let's be blunt about the rows we lose. If your data is in BigQuery, Google reads its own query logs and hands you table- and column-level lineage that no outsider can reconstruct. Metadata harvesting needs no connectors. Gemini writes the descriptions, infers relationships, proposes SQL patterns, and since June 2026 runs profile scans over unstructured data. Data products went GA in May 2026, so the unit an agent consumes is curated rather than raw. The MCP endpoint is live today and any Claude or ChatGPT client can hit it. The first 100 DCU-hours each month are free. Against that, our catalog is GA for PostgreSQL and declared for everything else, and we are not going to pretend otherwise.
Here is the push-back. All of that is context, and context is not control. When an agent calls their MCP server it acts under IAM, as a service account or a user, and the catalog has no idea that the agent's authority changed 90 seconds ago. The record of that call lands in Cloud Audit Logs, which you can retain and export, and which carry no integrity proof of their own. Gartner expects most unauthorized agent transactions through 2028 to be internal policy violations rather than attacks. A better description of your tables will not help you there. A signed chain showing who asked, what was allowed, and when the authority was pulled, will.
Questions worth asking both of us
These are the questions we'd want answered if we were buying. Ask them on every vendor call, including ours.
Can you cryptographically prove an audit record wasn't deleted?
DataShield: yes. Each record commits to the one before it, checkpoints are Ed25519-signed and themselves chained, and verification separates deletion from truncation from tampering. Try it on a sample chain at /verify. Google: catalog activity lands in Cloud Audit Logs, with Data Access logs enabled by you. We found no published tamper-evidence for the log itself. Ask them how you would detect a removed entry six months later.
What happens to a revoked agent mid-session?
DataShield re-checks authority on every governed tool call, so a revocation bites on the next call and the context drops to anonymous. With Dataplex you are relying on IAM. Ask how long a running agent keeps catalog access after you remove the role, and whether the call in flight fails or completes.
How does GDPR erasure interact with the audit trail?
DataShield crypto-shreds the per-subject key material and issues an ISO 27560 consent receipt. Actor identities in the chain are HMAC-committed, so the evidence still verifies after the subject is gone. Knowledge Catalog has no erasure function; it is metadata about where things live. Someone still has to own the deletion and the proof.
Is this a Dataplex alternative, or something you run next to it?
Next to it, if you live in BigQuery. Their harvesting and column-level lineage are better than ours and cost you almost nothing to start. We become the alternative when the estate spans clouds, when the deployment has to be self-hosted, or when the obligation is proving an agent's actions rather than describing tables. Keep their context. Put the governed data plane and the evidence here.
When an agent calls your MCP server, what identity is it acting under?
Ask Google this one. Their remote MCP server uses OAuth 2.0 with IAM scopes and roles such as roles/mcp.toolUser, which means the agent borrows a human or service-account identity. In DataShield Auth the agent holds its own MCP tool token with a scope ceiling, its authority is re-checked per call, and each call is metered and attributed to that agent.
Google is discontinuing the old Data Catalog. Does that change anything?
It changes the timing. Legacy Data Catalog went end of life on 1 June 2026, and Dataplex Universal Catalog was renamed Knowledge Catalog in April 2026. Plenty of teams are opening the migration ticket anyway. That is a good moment to decide which parts of governance you want inside a cloud provider and which parts you want to own.
Does DataShield have SOC 2?
Not yet, and we won't imply otherwise. Google Cloud's compliance coverage is enormous and that is a fair thing to weigh against a small vendor. What we offer instead is checkable: a published threat model, a verifier you can run, and Auth, Guardian and Lighthouse live in production since April 2026. Design-partner terms include source escrow. Details on the security page.
- Google's positioning: "Knowledge Catalog (formerly Dataplex)", with the sub-headline "Always-on context and governance for your agents". — cloud.google.com product page, 13 Sep 2026
- Dataplex Universal Catalog was renamed Knowledge Catalog on 10 April 2026, with API, client library, CLI and IAM names unchanged. Google describes it as a Gemini-powered catalog that "builds a dynamic context graph that grounds AI agents in enterprise truth". — Knowledge Catalog introduction, 13 Sep 2026
- Remote MCP server at dataplex.googleapis.com/mcp uses OAuth 2.0 with IAM, scopes dataplex.readonly and dataplex.read-write, and roles including roles/mcp.toolUser. Listed clients: Gemini CLI, ChatGPT, Claude and custom applications. — Knowledge Catalog remote MCP docs, 13 Sep 2026
- Data products reached GA on 25 May 2026, the data-lineage MCP server entered Preview on 27 May 2026, unstructured profile scans with Gemini on 11 June 2026, dbt Core import on 26 August 2026, and data domains on 7 September 2026. — Knowledge Catalog release notes, 13 Sep 2026
- Published rates: $0.06 per DCU-hour Standard processing with the first 100 DCU-hours per month free, $0.089 per DCU-hour Premium processing for lineage, data quality and profiling with no free allowance, and roughly $2 per GiB-month for custom metadata storage. — Dataplex pricing, 13 Sep 2026
- The legacy Data Catalog product was discontinued on 1 June 2026, with a documented transition to the Dataplex catalog. — Google transition docs, 13 Sep 2026
- Google states that over 95% of its top Google Cloud data analytics customers use Dataplex. — Google Cloud blog, 9 Oct 2024
Other head-to-heads
DataShield vs Unity Catalog
The other platform catalog, and the same question about independence.
Platform catalogDataShield vs Microsoft Purview
Governance bundled with the cloud you already bought.
Cloud-nativeDataShield vs Google Cloud DLP
Google's other half of this story: finding the sensitive data.
AllEvery comparison
One honest scorecard per vendor, including the rows we lose.
Keep Google's context. Add the part it doesn't do: break a live audit chain in your browser, revoke an agent mid-session, then decide. Demo Center access is free with a work email.
Get free Demo Center accessYou've seen the proof
Ready for a number? Scope your deployment and we'll price it against your own economics.
Get your quote →