Head-to-head · updated 13 September 2026

DataShield vs Milvus / Zilliz: who governs the text before it becomes a vector?

Milvus is a very good vector database. It is Apache 2.0, it graduated in the LF AI & Data Foundation, and it will hold a billion vectors without complaining. Zilliz sells the managed version and now calls the whole thing a "Vector Lakebase for Enterprise AI, Powered by Milvus". Milvus 3.0, shipped in July 2026, indexes Parquet, Lance and Iceberg files where they already sit, so you stop paying to keep a second copy. If your problem is retrieval scale, that is a fine answer and we are not going to talk you out of it.

We are not a faster index. DataShield's RAG runs on pgvector with local embeddings, and it is bolted to the rest of a governance stack. Text gets classified before it is embedded. Datasets are tokenized at ingest; agents query tokenized data over MCP; detokenization is a privileged, audited operation. Every governed tool call is re-checked against the agent's current authority, and the decision is sealed into a hash chain you can verify without trusting us. Below is the split, including the rows Milvus wins outright.

DataShield vs Milvus / Zilliz at a glanceEight questions regulated buyers ask us. Scored from each vendor's public material. DataShield vs Milvus / Zilliz at a glance Eight questions regulated buyers ask us. Scored from each vendor's public material. DataShield Milvus / Zilliz Tamper-evident audit chain you can verify Authority re-checked on every tool call PII and PHI classified before embedding Tokenized values in the retrieval path Vector search at billion scale Index tuning and open formats Open-source engine and community Self-hosting and published pricing shipped partial / roadmap not offered Sources at the bottom of this page.

The short version

Pick DataShield when

  • The corpus has PHI or customer records in it, and someone needs the sensitive fields handled before the text is embedded. We classify against 129 field classes, then tokenize at ingest.
  • Someone will eventually ask you to prove an agent's retrieval log was not edited. Our chain answers with math, not a policy PDF. Run the verifier.
  • You need to pull an agent's authority mid-session and have the very next tool call fail. Not the next token refresh. The next call.
  • You want retrieval, catalog, classification, tokenization and evidence in one self-hosted stack, rather than five products you glue together.

Pick Milvus / Zilliz when

  • Scale is the whole problem. A billion vectors, tight latency budgets, and real choice of index type. That is what Milvus was built for and we are not going to match it.
  • You want the data to stay in the lake. Milvus 3.0 external collections index Parquet, Lance, Iceberg and Vortex files in place, read-only and zero-copy. Nobody else we track does that as cleanly.
  • You want an Apache 2.0 engine with a large community and foundation governance, so the thing outlives any one vendor's fundraising.
  • You want managed operations with an SLA, and Zilliz Cloud's compliance list already satisfies your procurement checklist.

Bottom line: Milvus stores vectors well. It has no opinion about what was in the text before it became one. If your corpus is public documentation, that is fine. If it is patient notes, the obligation lands somewhere, and a vector index is not where it lands.

Feature by feature

Competitor cells describe what the Milvus docs, the Milvus blog and Zilliz's public material say as of the date above. If we have mischaracterised something, email support@myorg.ai and we will correct it, credited.

What mattersDataShieldMilvus / ZillizEdge
Vector search at scalepgvector with 768-dimension vectors, one chunk table per corpus, hybrid retrieval that runs a keyword leg and a vector leg together. Good for governed corpora. Not a billion-vector engine, and we publish no retrieval accuracy number because we have not benchmarked one.The core product, and very strong. Multiple index types, horizontal scaling, and a 3.0 release with faster sparse retrieval and server-side ranking.
Lake-native retrievalWe ingest into a governed dataset. Analytics run on DuckDB over Parquet, but that reads our masked golden views, not your lake.Milvus 3.0 external collections index data that stays in object storage, in open formats, zero-copy. Plus snapshots and a Spark connector. Genuinely good engineering.
Classification before embedding129 field classes covering PII, PHI, financial data and secrets. Deterministic, with checksum validators and reproducible verdicts stamped with a config digest. All 18 HIPAA Safe Harbor identifiers are discrete classes.None. Milvus indexes the vectors and metadata you hand it. What was in the source text is your problem.
Tokenization and data handlingDeterministic, join-preserving, vault-reversible tokens applied at ingest, plus quasi-identifier generalization (dates to year, decade or age band; ZIPs to 3 or 4 digits; partial phones, SSNs and emails) with a measured cardinality-reduction score per column. Detokenization is privileged and logged.Not part of the product. Encryption at rest and in transit on Zilliz Cloud, which is a different thing.
EmbeddingsLocal nomic-embed-text, 768 dimensions, running on your own infrastructure. No embedding call leaves the building.Bring your own. Self-hosted Milvus can sit behind a local model too, so this is a tie on the engine. On Zilliz Cloud, ask where the embedding call goes.
Agent authorizationEvery governed tool call passes a scope ceiling, a consented-tool allowlist, an authority tier and a revocation re-check before dispatch. It fails closed.RBAC with users, roles, privileges and privilege groups, plus a built-in admin role. That is database-level access control, evaluated when the connection is made, not a per-call decision against an agent's current authority.
MCP and agentsMore than 200 MCP tools across Ontology, Auth, Corpus and Lighthouse. MCP tool tokens carry scope ceilings, and per-call metering is attributed to the agent.A real first-party MCP server. It exposes search and query, and also create collection, insert data and delete entities, behind one optional MILVUS_TOKEN environment variable. One static credential, read and write.
Audit evidenceSHA-256 hash chain with Ed25519-signed checkpoints that are themselves chained. Verification returns clean, attested damage, or tampered, and names the failure mode. See how.Zilliz Cloud lists audit logs among its enterprise features. We found no published tamper-evidence mechanism, and open-source Milvus does not ship one.
Break-glassScoped, time-boxed emergency access for agents. It auto-revokes and cannot be quietly deleted from the log.Not described in their public material.
GDPR erasureCrypto-shred of per-subject key material plus ISO 27560 consent receipts. The audit chain still verifies afterwards.You can delete entities. Whether the subject's data also left every index, snapshot and external collection is a question worth asking, especially now that snapshots record point-in-time views.
DeploymentSelf-hosted in your own cloud or data center, or a dedicated single-tenant server we operate. Docker images for Auth, Ontology, Corpus and Lighthouse, with a signed deploy manifest Guardian verifies. Ed25519 audit-signing keys can live in your KMS or HSM. HMAC tokenization keys sit in your environment today, not in a KMS.Apache 2.0 self-hosting with no control-plane dependency, or Zilliz Cloud serverless, dedicated and BYOC. Both good options.
Maturity signalsAuth, Guardian and Lighthouse are live in production (Guardian and Lighthouse since April 2026). SOC 2 not yet certified, and we say so.Milvus open-sourced in 2019 and a graduated LF AI & Data project. Zilliz founded 2017, around $113M raised, and Zilliz Cloud self-reports SOC 2 Type II and ISO 27001. Much longer track record than ours.
PricingPublished model, scoped instant quote, no sales wall.Milvus is free under Apache 2.0, which is hard to beat. Zilliz Cloud publishes consumption pricing and standardized storage at $0.04 per GB per month across AWS, Azure and GCP from 1 January 2026. Enterprise is quote-only.

◆ DataShield leads◇ Milvus / Zilliz leads◈ comparable

Milvus and Zilliz claims are drawn from milvus.io docs and blog posts, the zilliztech/mcp-server-milvus repository, and zilliz.com, last checked 13 September 2026. We link them below rather than work from memory.

Three things you get here that you won't get from a vector database

The sensitive fields handled before the embedding

An embedding is not a safe container. It is derived from the text, and the text had a medical record number in it. We classify the source columns first, then tokenize at ingest, so the string that reaches the model is already a token. How the data plane works.

Authority that can change mid-flight

An analyst leaves on a Friday. Their agent is 20 minutes into a 40-minute retrieval job. With DataShield the next governed tool call is re-checked against current authority and fails closed. A database role grant does not do that, because nobody re-reads it until the next connection. How Auth does it.

Proof that survives an audit

A log that can be silently edited proves nothing. Ours is a hash chain with signed checkpoints, and the verifier tells you what broke, not just that something did. That is the property EU AI Act Article 12 and HIPAA §164.312(b) reviewers care about. Try it in your browser, no signup.

Where Milvus / Zilliz is genuinely stronger

Retrieval is their craft and it shows. Milvus 3.0 landed in July 2026. It brought lake-native indexing, zero-copy external collections over Parquet, Lance, Iceberg and Vortex, point-in-time snapshots, a Spark connector, and server-side aggregation. Then they spent August publishing the engineering detail behind each piece. That is what a team does when it expects to be read by other engineers. The index-type choice is real: you can tune recall against cost at a scale we have never been asked to reach. The licence is Apache 2.0 and the project is foundation-governed, so your dependency is not one company's cap table. And the storage price cut to a flat rate across three clouds is the kind of unglamorous work that makes a procurement team's year.

Here is the push-back, and it is narrow on purpose. Milvus governs the index. It does not govern the corpus. Their own MCP server hands an agent search, query, create collection, insert data and delete entities. All of it sits behind one optional token. That is a fine developer default. It is a bad production posture for a regulated corpus. RBAC answers who connected. It does not answer whether this call, from this agent, right now, is still allowed, and it leaves no artefact an examiner can verify independently. Gartner expects most unauthorized agent transactions through 2028 to be internal policy violations rather than attacks, which is exactly the case a connection-time role grant handles worst. Run Milvus for the index. Put the obligation somewhere it can actually be enforced and proved.

Questions worth asking both of us

These are the questions we would want answered if we were the ones buying. Ask them on every call, ours included.

Can you cryptographically prove an audit log entry wasn't deleted?

DataShield: yes. Each record commits to the one before it, checkpoints are signed and chained, and verification tells deletion apart from truncation and from tampering. Try it at /verify. Milvus / Zilliz: Zilliz Cloud lists audit logs as an enterprise feature. We found no published tamper-evidence mechanism in their docs, and the open-source engine has none. Ask them to show you one.

What happens to a revoked agent mid-session?

DataShield re-checks authority on every governed tool call, so revocation lands on the next call. Milvus has users, roles, privileges and privilege groups. Revoking a role is a database operation, and we could not find documentation of how an in-flight agent session is affected. Ask how long a compromised agent keeps retrieving after you pull its access.

How does GDPR erasure interact with the audit trail?

DataShield crypto-shreds per-subject key material and issues an ISO 27560 consent receipt. Actor identities in the chain are HMAC-committed, so the evidence still verifies once the subject is gone. Milvus gives you delete entities. Ask what happens to the vectors derived from that subject's text, to older snapshots, and to any external collection built over files you do not control.

Is DataShield a Milvus alternative? Do we rip out our vector database?

Usually not. If you already run Milvus at scale and the corpus is not sensitive, keep it. We are worth a look when the corpus is sensitive. Then the job is classification, tokenization, per-call authorization and evidence. A vector index does none of those. Plenty of teams will sensibly run both: Milvus for the index, us for the governed datasets and the proof.

Milvus 3.0 indexes data in the lake without moving it. Can you do that?

No, and we are not going to pretend the workaround is equivalent. We ingest into a governed dataset, which is the point: the classification, tokenization and access gating happen on the way in. External collections are read-only indexes over files whose contents nobody inspected. That is a great fit for a product catalogue and a poor one for an HR export. Different trade, honestly made.

Does DataShield have SOC 2?

Not yet, and we will not imply otherwise. Zilliz self-reports SOC 2 Type II and ISO 27001, and on a procurement checklist that beats us today. What we offer instead is a public threat model and a verifier you can run. Design-partner terms include source escrow, so a small vendor is not a single point of failure. Details on the security page.

Other head-to-heads

Vector DB

DataShield vs Pinecone

Managed retrieval at speed, versus governance on the way in.

Vector DB

DataShield vs Qdrant

A fast open-source index, and the obligations it leaves you.

Vector DB

DataShield vs Weaviate

Hybrid search and modules, versus tokenized data and proof.

All

Every comparison

One honest scorecard per vendor.

See both mechanisms run in your browser: break a live audit chain, revoke an agent mid-session, then decide what your vector database still owes you. Demo Center access is free with a work email.

Get free Demo Center access

You've seen the proof

Ready for a number? Scope your deployment and we'll price it against your own economics.

Get your quote →