Head-to-head · updated 13 September 2026

DataShield vs Google Dataplex: context for your agents, or control over them?

In April 2026 Google renamed Dataplex Universal Catalog to Knowledge Catalog, and the new name is honest about the plan. The product page calls it "always-on context and governance for your agents". If your data sits in BigQuery, it harvests the metadata with no connectors to run, draws column-level lineage you would struggle to rebuild by hand, writes descriptions with Gemini, and serves the lot to Claude or Gemini CLI from a remote MCP endpoint. Nothing to install. Nothing to procure.

We sell the next question. Which agent asked, was it still allowed at that second, and can you prove the record later? Datasets are tokenized at ingest; agents query tokenized data over MCP; detokenization is a privileged, audited operation. Every governed tool call in Auth is re-checked against the agent's current authority, and each decision is sealed into a hash chain you can verify without trusting us. On GCP, most buyers run both. Here is the honest split, including the rows Google wins.

DataShield vs Google Dataplex at a glanceEight questions regulated buyers ask us. Scored from each vendor's public documentation. DataShield vs Google Dataplex at a glance Eight questions regulated buyers ask us. Scored from each vendor's public documentation. DataShield Google Dataplex Tamper-evident audit chain you can verify Authority re-checked on every tool call Break-glass access for agents GDPR erasure that keeps the chain valid Runs outside Google Cloud Field-level PII and PHI classification in the catalog Column-level lineage across the analytics estate Metadata coverage of a large cloud estate shipped partial / roadmap not offered Sources at the bottom of this page.

The short version

Pick DataShield when

  • Someone will one day ask you to prove an agent's access record wasn't edited. Cloud Audit Logs are retained and exportable. Ours are chained and signed, and the verifier names the failure. Run it yourself.
  • The catalog has to run where Google does not. Self-hosted in your own cloud or data center, or a dedicated single-tenant server we operate, with your keys. See the architecture.
  • You need to pull an agent's authority mid-session and have the very next tool call fail closed, not wait out a token. How Auth does it.
  • You want PII and PHI labels as first-class catalog objects, not a separate scanning product. We classify columns against 129 field classes, and the catalog stores no column values by default. What Ontology holds.

Pick Google Dataplex when

  • Your estate is BigQuery. Metadata from tables, views, models, Spanner, Bigtable, Cloud SQL, Looker and Pub/Sub arrives on its own. We cover PostgreSQL today, not your whole estate.
  • You want column-level lineage for BigQuery and Vertex AI pipelines. Google reads its own query logs. No third party gets that depth from outside.
  • Gemini in the catalog is worth money to you: generated descriptions, semantic search, inferred relationships, and profile scans over unstructured data.
  • Procurement is the bottleneck. It is already on the GCP bill, the first 100 DCU-hours each month are free, and there is no new vendor to onboard.

Bottom line: Google gives your agents context. We decide which agent may act on it, and keep proof that the decision was made. On GCP, run both. Off GCP, or in front of an auditor, start here.

Feature by feature: a platform catalog against a governed data plane

Competitor cells describe what Google's public docs, release notes and pricing page say as of the date above. If we've mischaracterised something, email support@myorg.ai and we'll correct it, credited.

What mattersDataShieldGoogle DataplexEdge
Audit evidenceSHA-256 hash chain with Ed25519-signed checkpoints that are themselves chained. Verification returns clean, attested damage, or tampered, and names the break. Try the verifier.Cloud Audit Logs: Admin Activity plus Data Access logs you switch on. The audit-logging doc describes no integrity proof, no signatures and no immutability for the log itself.
Agent authorizationEvery governed tool call in Auth passes a scope ceiling, a consented-tool allowlist, a declared authority tier, and a revocation re-check before dispatch. Delegation is RFC 8693 token exchange with an enforced ceiling.The MCP server uses OAuth 2.0 with IAM roles such as roles/mcp.toolUser. That tells the catalog which human or service account is behind the call. It does not give the agent its own authority, or re-check it.
Break-glassScoped, time-boxed emergency access for agents. It auto-revokes and it cannot be quietly removed from the record.Not offered. You would assemble it from IAM conditions and hope the rollback is clean.
GDPR erasureCrypto-shred of per-subject key material plus ISO 27560 consent receipts. Actor identities are HMAC-committed, so the chain still verifies after a subject is erased.Not a stated function of the catalog. It records where the data is. Deleting it, and proving you did, is yours.
Tokenization and data handlingDatasets are tokenized at ingest; agents query tokenized data over MCP; detokenization is a privileged, audited operation. Tokens are deterministic and join-preserving. Quasi-identifier generalization is there when you switch it on: dates to year, decade or age band, ZIPs to 3 or 4 digits, partial phones, SSNs and emails, with a measured cardinality-reduction score per column.The catalog is metadata. De-identification is a different Google product (Sensitive Data Protection), billed and configured separately. Column-level security runs on BigQuery policy tags.
Sensitive-data classification in the catalog129 field classes covering PII, PHI, financial data and secrets, including all 18 HIPAA Safe Harbor identifiers. Regex plus checksum validation (Luhn, NPI, Verhoeff, ABA, IBAN, GTIN) plus column-name lexicons and anti-pattern suppressors. No model, so verdicts are reproducible from a config digest. Column values are not copied into the catalog by default.Data profiling and quality scans in the Premium tier, with Gemini-generated semantics. Sensitive-data labelling leans on Sensitive Data Protection and BigQuery policy tags rather than the catalog itself.
Catalog coverage and lineageWe scan, profile and classify a live PostgreSQL source in place, with no rows leaving it. Typed lineage traversal with per-hop access gating, a glossary materialized from entity types, and stewardship worklists. Lineage is derived from the pipelines that own it, not a stored column-level graph. PostgreSQL today, not your whole estate.Auto-harvested metadata across BigQuery, Vertex AI, Spanner, Bigtable, Cloud SQL, Looker and Pub/Sub, plus table- and column-level BigQuery lineage, business glossary, aspects, data products and Iceberg REST cataloging. This is their round, and it isn't close.
Context for agentsAgents query the catalog and act on it through MCP. Ontology's catalog, stewardship and lineage surfaces are MCP-native, and Auth issues MCP tool tokens with scope ceilings for its own surface.A genuinely good story. Remote MCP server at dataplex.googleapis.com/mcp, a data-products endpoint, lineage MCP in Preview, and a lookupContext bundle built for agent workflows. Gemini CLI, ChatGPT and Claude are all listed clients.
MCP and agents overallMore than 200 MCP tools across Ontology, Auth, Corpus and Lighthouse, with per-call metering attributed to the agent and typed refusals when a quota bites.MCP for discovery and data products, with lineage in Preview. No agent identity of its own, no per-call authority re-check, no metering attributed to an agent.
DeploymentSelf-hosted in your own cloud or data center, or on a dedicated single-tenant server we operate. Docker images ship for Auth, Ontology, Corpus and Lighthouse, and Guardian verifies a signed deploy manifest.A managed Google Cloud service. No self-hosted option, no on-premises option. Third-party and on-prem sources arrive through connectors, not as first-class citizens.
Maturity signalsAuth, Guardian and Lighthouse are live in production (Guardian and Lighthouse since April 2026). SOC 2 not yet certified, and we say so.Alphabet. GCP compliance coverage, a release note most weeks, and Google's own claim that over 95% of its top analytics customers use Dataplex. Three product renames in four years is the one wobble.
PricingPublished model, scoped instant quote, no sales wall.Published rates, metered in Data Compute Unit hours: $0.06 per DCU-hour Standard with 100 free each month, $0.089 Premium for lineage, quality and profiling with no free allowance, and about $2 per GiB-month for custom metadata. Cheap to start. Hard to forecast.

◆ DataShield leads◇ Google Dataplex leads◈ comparable

Google claims are drawn from cloud.google.com product, docs, release-notes and pricing pages, last checked 13 September 2026. We link them below rather than paraphrase from memory.

Three things you get here that you won't get from a platform catalog

Proof the record wasn't edited

A retained log is not a tamper-evident log. Ours is a hash chain with signed checkpoints, and the verifier tells you what went wrong: tampering, insertion, deletion or truncation. That is the property EU AI Act Article 12 and HIPAA §164.312(b) reviewers keep asking about. Try it in your browser, no signup.

Authority that changes mid-flight

An analyst resigns on a Friday afternoon. Their agent is 20 minutes into a 40-minute job. With DataShield the next governed tool call is re-checked against current authority and fails closed. An IAM change is eventually consistent, and the call in flight rarely notices. How Auth does it.

A catalog that can leave the cloud it governs

Knowledge Catalog is a Google Cloud service, and it is not going anywhere else. If your estate spans clouds, or your regulator wants the governance layer outside the provider's blast radius, the catalog has to be yours. Ours runs in your data center or on a single-tenant server we operate. See the architecture.

Where Google Dataplex is genuinely stronger

Let's be blunt about the rows we lose. If your data is in BigQuery, Google reads its own query logs and hands you table- and column-level lineage that no outsider can reconstruct. Metadata harvesting needs no connectors. Gemini writes the descriptions, infers relationships, proposes SQL patterns, and since June 2026 runs profile scans over unstructured data. Data products went GA in May 2026, so the unit an agent consumes is curated rather than raw. The MCP endpoint is live today and any Claude or ChatGPT client can hit it. The first 100 DCU-hours each month are free. Against that, our catalog is GA for PostgreSQL and declared for everything else, and we are not going to pretend otherwise.

Here is the push-back. All of that is context, and context is not control. When an agent calls their MCP server it acts under IAM, as a service account or a user, and the catalog has no idea that the agent's authority changed 90 seconds ago. The record of that call lands in Cloud Audit Logs, which you can retain and export, and which carry no integrity proof of their own. Gartner expects most unauthorized agent transactions through 2028 to be internal policy violations rather than attacks. A better description of your tables will not help you there. A signed chain showing who asked, what was allowed, and when the authority was pulled, will.

Questions worth asking both of us

These are the questions we'd want answered if we were buying. Ask them on every vendor call, including ours.

Can you cryptographically prove an audit record wasn't deleted?

DataShield: yes. Each record commits to the one before it, checkpoints are Ed25519-signed and themselves chained, and verification separates deletion from truncation from tampering. Try it on a sample chain at /verify. Google: catalog activity lands in Cloud Audit Logs, with Data Access logs enabled by you. We found no published tamper-evidence for the log itself. Ask them how you would detect a removed entry six months later.

What happens to a revoked agent mid-session?

DataShield re-checks authority on every governed tool call, so a revocation bites on the next call and the context drops to anonymous. With Dataplex you are relying on IAM. Ask how long a running agent keeps catalog access after you remove the role, and whether the call in flight fails or completes.

How does GDPR erasure interact with the audit trail?

DataShield crypto-shreds the per-subject key material and issues an ISO 27560 consent receipt. Actor identities in the chain are HMAC-committed, so the evidence still verifies after the subject is gone. Knowledge Catalog has no erasure function; it is metadata about where things live. Someone still has to own the deletion and the proof.

Is this a Dataplex alternative, or something you run next to it?

Next to it, if you live in BigQuery. Their harvesting and column-level lineage are better than ours and cost you almost nothing to start. We become the alternative when the estate spans clouds, when the deployment has to be self-hosted, or when the obligation is proving an agent's actions rather than describing tables. Keep their context. Put the governed data plane and the evidence here.

When an agent calls your MCP server, what identity is it acting under?

Ask Google this one. Their remote MCP server uses OAuth 2.0 with IAM scopes and roles such as roles/mcp.toolUser, which means the agent borrows a human or service-account identity. In DataShield Auth the agent holds its own MCP tool token with a scope ceiling, its authority is re-checked per call, and each call is metered and attributed to that agent.

Google is discontinuing the old Data Catalog. Does that change anything?

It changes the timing. Legacy Data Catalog went end of life on 1 June 2026, and Dataplex Universal Catalog was renamed Knowledge Catalog in April 2026. Plenty of teams are opening the migration ticket anyway. That is a good moment to decide which parts of governance you want inside a cloud provider and which parts you want to own.

Does DataShield have SOC 2?

Not yet, and we won't imply otherwise. Google Cloud's compliance coverage is enormous and that is a fair thing to weigh against a small vendor. What we offer instead is checkable: a published threat model, a verifier you can run, and Auth, Guardian and Lighthouse live in production since April 2026. Design-partner terms include source escrow. Details on the security page.

Other head-to-heads

Platform catalog

DataShield vs Unity Catalog

The other platform catalog, and the same question about independence.

Platform catalog

DataShield vs Microsoft Purview

Governance bundled with the cloud you already bought.

Cloud-native

DataShield vs Google Cloud DLP

Google's other half of this story: finding the sensitive data.

All

Every comparison

One honest scorecard per vendor, including the rows we lose.

Keep Google's context. Add the part it doesn't do: break a live audit chain in your browser, revoke an agent mid-session, then decide. Demo Center access is free with a work email.

Get free Demo Center access

You've seen the proof

Ready for a number? Scope your deployment and we'll price it against your own economics.

Get your quote →