Head-to-head · updated 13 September 2026

DataShield vs DataHub: your context graph knows everything. Who tells the agent no?

DataHub started at LinkedIn in 2015 and it shows. The metadata model is the most bendable of the open catalogs, the lineage runs deep, and the connector list passes a hundred sources. The company used to be called Acryl Data. Today it calls itself "The Context Platform for AI Agents" and sells a context graph that makes analytics agents answer better. Miro told them accuracy went from about 50% to around 90%. The core is Apache-2.0 and free to self-host forever, with about 12,700 stars behind it. We are not going to argue with any of that.

We are not a catalog for your whole estate. We scan and classify a live PostgreSQL source in place, and that is the honest edge of it. What we add is the part a catalog leaves to you. Datasets are tokenized at ingest; agents query tokenized data over MCP; detokenization is a privileged, audited operation. Every governed tool call is checked against the agent's authority right then, and the decision is sealed into a hash chain you can verify without trusting us. Here is the split, rows they win included.

DataShield vs DataHub at a glanceEight questions self-hosted regulated buyers ask us. Scored from each vendor's public material. DataShield vs DataHub at a glance Eight questions self-hosted regulated buyers ask us. Scored from each vendor's public material. DataShield DataHub Tamper-evident audit chain you can verify Authority re-checked on every tool call Break-glass access for agents Field-level PII and PHI classes built in Prices published before a sales call Connector breadth across the estate Lineage depth and impact analysis Apache-2.0 core and open community shipped partial / roadmap not offered Sources at the bottom of this page.

The short version

Pick DataShield when

  • Someone will ask you to prove an agent's access log was not edited. An auditor, an examiner, or Article 12 of the EU AI Act. Our chain answers with math, not a policy PDF. Run the verifier.
  • You need to pull an agent's authority mid-session and have the very next tool call fail. Not the next token refresh. The next call.
  • You want field classification, tokenization, agent authority and audit in one self-hosted stack, instead of bolting four products onto a catalog.
  • You would like to see a price before you book a demo.

Pick DataHub when

  • You need to catalog the whole estate. A hundred plus connectors, deep upstream and downstream lineage, impact analysis. That is their home ground and we are nowhere near it.
  • Your problem is an analytics agent giving wrong answers. Context Intelligence and the Context Hub are aimed exactly there, and they have a customer number to wave at you.
  • You want an Apache-2.0 core you can fork, read and run forever, with a large community and a decade of hyperscale use behind it.
  • You live on Google Cloud or Snowflake. The Knowledge Catalog sync and the Open Semantic Interchange work are real alliances, not logos on a slide.

Bottom line: DataHub tells your agent what exists and what it means. We decide what the agent may touch and keep proof of the call. If the job is estate-wide discovery, buy DataHub. If agents are already reading regulated data and someone will ask you to account for it, that gap is ours.

Feature by feature

Competitor cells describe what DataHub's public site, docs and press releases say as of the date above. If we have mischaracterised something, email support@myorg.ai and we will correct it, credited.

What mattersDataShieldDataHubEdge
Catalog and connector breadthWe register a provider over a live connection, then scan and profile its objects and columns in place with no rows leaving the source. GA for PostgreSQL only. Snowflake, BigQuery, Databricks, SQL Server, MySQL, S3, Kafka and Salesforce are declared with no handler yet, and the config audit refuses to let us mark them shipped.Context Ingestion unifies metadata from 100+ sources, plus semantic definitions from dbt and Power BI and unstructured institutional knowledge. This is their strongest row and the reason most people buy them.
LineageTyped lineage over an edges surface with four commands and both-endpoint access gating at every hop. It is derived from the pipelines that own the relationships, not a stored column-level graph, and there is no edge-write command by design.Fine-grained upstream and downstream lineage, impact analysis, and a metadata graph proven at LinkedIn and Netflix scale. Deeper than ours, and we would say so on a call.
Field-level PII and PHI classification129 built-in field classes covering PII, PHI, financial data and secrets, including all 18 HIPAA Safe Harbor identifiers and eight non-US national ID formats. Regex plus checksum validation (Luhn, NPI, Verhoeff, ABA, IBAN, GTIN), column-name lexicons and anti-pattern suppressors. No model, so verdicts are reproducible from a config digest. It cut over from shadow mode days ago, which is new, not battle-tested.Tags, glossary terms and classification you apply or ingest. We found no built-in field-class library with checksum validators in their docs.
Audit evidenceSHA-256 row chain with Ed25519-signed checkpoints that are themselves chained. Verification returns clean, attested damage, or tampered, and names the failure as insertion, deletion, or truncation. Try the verifier demo.Versioned, timestamped metadata change history, and their v1 release says agents "know why it changed." Their access-policy docs describe no tamper-evident or cryptographic audit mechanism. Ask them for one.
Agent authorizationEvery governed tool call passes a scope ceiling, an authority tier and a revocation re-check before dispatch, then gets metered and sealed into the audit chain. Cedar handles admin, config and token decisions, not the dispatch hot path, and we will not claim otherwise.The MCP server enforces the access controls and policies already set in your DataHub instance. Platform and metadata policies over resources, privileges and actors. Service accounts can carry a default view that narrows what the server searches.
Break-glassScoped, time-boxed emergency access for agents, admin and IP-allowlist gated with step-up, fully audited and auto-revoking.Not described in their public material.
GDPR erasureCrypto-shred: deleting a subject's key row destroys every ciphertext for that subject at once, cited to ISO/IEC 27040 and GDPR Art. 17. Actor identities in the chain are HMAC-committed, so the audit still verifies afterwards. ISO 27560 consent receipts are signed at grant and at withdrawal.Soft and hard delete of entities and aspects from the metadata graph. No erasure mechanism for the underlying data, which is not their job, and no equivalent to a consent receipt.
Tokenization and generalizationDeterministic, join-preserving, vault-reversible tokens applied at ingest, plus quasi-identifier generalization (dates to year, decade or age band; ZIPs to 3 or 4 digits; partial phones, SSNs and emails) with a measured cardinality-reduction score per column. Both are features you switch on, not defaults.Not a thing they do. A catalog describes data; it does not transform it.
MCP and agentsMore than 200 MCP tools across Ontology, Auth, Corpus and Lighthouse. Agents query the governed data itself, under Auth-issued MCP tool tokens with scope ceilings, with per-call metering attributed to the agent.A real MCP server, and a good one. Search, entity fetch, schema, lineage traversal, SQL retrieval and drafting, plus mutation tools from v0.5.0 for tags, terms, ownership and proposals. OAuth2 with dynamic client registration. It touches metadata only, never table rows. Block has run it in production with Goose since June 2025.
Prompt egressWe do not proxy your LLM traffic. We do gate every value that leaves a governed dataset for a prompt, and a PHI dataset refuses an endpoint without a BAA or a zero-retention agreement.Out of scope. The context graph feeds the prompt; nothing in it inspects what the prompt then carries.
Deployment and licensingSelf-hosted in your own cloud or data center, or a dedicated single-tenant server we operate. Docker images for Auth, Ontology, Corpus and Lighthouse, with a signed deploy manifest Guardian verifies. Ed25519 audit-signing keys can live in your KMS or HSM. HMAC tokenization keys sit in your environment today, not in a KMS. Commercial licence, not open source.Apache-2.0 core you can self-host forever, or DataHub Cloud managed. That licence is a real advantage and the reason a lot of platform teams never call a vendor at all.
Maturity signalsAuth, Guardian and Lighthouse are live in production (Guardian and Lighthouse since April 2026). SOC 2 not yet certified, and we say so. Small team, source escrow on design-partner terms.Project born at LinkedIn in 2015, company founded 2021, $65M raised including a $35M Series B led by Bessemer in May 2025. Netflix, Pinterest, Optum, Miro and Block as references. Google Cloud and Snowflake alliances.
PricingPublished model, scoped instant quote, no sales wall.Core is free under Apache-2.0, which is hard to beat. DataHub Cloud is demo-and-quote only. Their pricing page returns a 404 and every button says "Get a Demo."

◆ DataShield leads◇ DataHub leads◈ comparable

DataHub claims are drawn from datahub.com, docs.datahub.com and DataHub's own news posts, last checked 13 September 2026. We link them below rather than work from memory.

Three things you get here that you won't get from a metadata catalog

Proof that survives an audit

A log an admin can quietly edit proves nothing. Ours is a hash chain with signed checkpoints, and the verifier names what broke. It answers clean, attested damage, or tampered, and it tells insertion apart from deletion and truncation. That is the property EU AI Act Article 12 and HIPAA §164.312(b) reviewers ask about. Try it in your browser, no signup.

Authority that can change mid-flight

An analyst quits on a Friday. Their agent is 20 minutes into a 40 minute job. With DataShield the next governed tool call re-checks current authority and fails closed. Edit a catalog policy instead and you wait for the agent's token to come round again. How Auth does it.

An erasure you can defend

GDPR says delete. Your auditor says keep the log. Crypto-shred settles it. We destroy the subject's key material, the data goes unreadable, and the chain still verifies. See the diagram.

Where DataHub is genuinely stronger

We would rather you heard this from us. DataHub has ten years of hyperscale use behind it and an Apache-2.0 licence you can read line by line. The connector list runs past a hundred sources. The lineage and impact analysis are deeper than ours and will stay that way. Their 2026 has been strong: a Google Cloud collaboration in April that syncs both ways with Knowledge Catalog, a seat at Snowflake's Open Semantic Interchange, and DataHub Cloud v1 in May with four context modules and a Miro accuracy number attached. Their MCP server is not a demo either. Block has run it with Goose since June 2025, it now has mutation tools and OAuth2, and it is careful about never touching table rows. If your problem is that nobody knows what any of your data means, call them, not us.

Here is the push-back. Context makes an agent smarter. It does not make an agent safe. The MCP server inherits the policies you already set in your DataHub instance, which is the right design for a catalog and a thin answer for a regulated shop. Their docs describe versioned change history, which tells you what a record looks like now and looked like before. They do not describe a way to prove nobody edited the record of what an agent did. And the agent still queries the warehouse on its own credentials, with nothing from DataHub in that path. Gartner expects most unauthorized agent transactions through 2028 to be internal policy violations rather than attacks. That is the case where better context is no help at all, because the agent was well informed and still did the wrong thing.

Questions worth asking both of us

These are the questions we would want answered if we were the ones buying. Ask them on every call, ours included.

Can you cryptographically prove an audit log entry wasn't deleted?

DataShield: yes. Each record commits to the one before it, checkpoints are signed and chained, and the verdict tells deletion apart from truncation and from tampering. Run it against a sample chain at /verify. DataHub: their docs cover versioned metadata history and access policies. We found no tamper-evidence mechanism in them. Ask to see one, and ask who can write to the database under it.

What happens to a revoked agent mid-session?

DataShield re-checks authority on every governed tool call, so revocation lands on the next call. DataHub's MCP server applies the policies set in your instance, and service accounts can carry a view that narrows what they search. We could not find a documented mid-session revocation path. Ask how long a compromised agent keeps working after you pull its access.

How does GDPR erasure interact with the audit trail?

DataShield crypto-shreds the subject's key material and issues an ISO 27560 consent receipt. Actor identities in the chain are HMAC-committed, so the evidence still verifies once the subject is gone. DataHub can soft or hard delete entities and aspects from the graph. That removes metadata, not the data itself, and the erasure story for the source system stays yours to solve.

We already run DataHub Core. Do we rip it out?

No. If it is cataloguing a hundred sources for you, keep it. Run us on the datasets agents actually read, where the job is classification, tokens, authority and proof rather than inventory. The overlap is smaller than it looks: we are PostgreSQL today, not your whole estate, and we have no opinion about your Kafka topics. Plenty of teams will sensibly run both.

DataHub ships an MCP server too. What's different?

Theirs hands an agent metadata: search, schemas, lineage, past SQL, and since v0.5.0 the ability to edit tags and terms or file a proposal. It never touches table rows, which is a deliberate and good choice. Ours is where the governed data itself is queried, so the tool token carries a scope ceiling, the call is authorized before dispatch, and the decision is sealed into the chain. Different jobs. Use theirs to ask what a table means. Use ours to read the table under a policy you can later prove.

Does DataShield have SOC 2?

Not yet, and we will not imply otherwise. Auth is live with a published threat model and a verifier anyone can run. Guardian and Lighthouse have been in production since April 2026. We have no SCIM endpoint either, which matters if you were counting on one. Design-partner terms include source escrow, so a small vendor is not a single point of failure. Details on the security page.

Other head-to-heads

OSS catalog

DataShield vs OpenMetadata

The other open catalog, and the same missing enforcement seam.

OSS catalog

DataShield vs Amundsen

Search and discovery for humans, versus authority for agents.

Platform catalog

DataShield vs Unity Catalog

Governance inside one platform, versus evidence across them.

All

Every comparison

One honest scorecard per vendor.

See both mechanisms run in your browser: break a live audit chain, revoke an agent mid-session, then decide what your catalog still owes you. Demo Center access is free with a work email.

Get free Demo Center access

You've seen the proof

Ready for a number? Scope your deployment and we'll price it against your own economics.

Get your quote →