Does Unity Catalog give you tamper-evident, independent evidence of what your agents did?
Databricks shipped real agent governance. Its "Governing AI agents at scale with Unity Catalog" post is worth your time. It ran May 20, 2026. The Unity AI Gateway writes the full payload of each model call. It logs the prompt, the response, token counts, and latency. Unity Catalog audit logs capture the caller, the agent, and the time. MLflow tracing auto-instruments common agent frameworks and model SDKs. Traces land as Unity Catalog tables too. That is a strong insight story. It is documented as production-ready.
But here is the catch. Each bit is data the platform emits about itself. It sits in tables the same platform governs. Databricks does market these inference tables. It calls them an audit-ready record for compliance review. That is the obvious pushback to raise. So let's answer it directly. Audit-ready means queryable and complete by design. Tamper-evident means something stronger. It means you can prove something with cryptography. Nobody changed, added, or removed a record later. Those are two distinct claims. Nowhere in those posts does Databricks call it hash-chained or signed. So the logs answer what happened. But they say nothing about whether a record could change later. Those are two distinct questions. Only one is a compliance question.
One caveat on sourcing. The mutable, non-cryptographic read is my own analysis. It covers table-based storage under one control plane. It is not a Databricks quote. Databricks never claims no one can change it. But it never rules that possibility out either.
Why native telemetry falls short as audit evidence
Visibility tells you what your system did. It helps your engineers debug it. Audit evidence is a stronger claim. It lets a third party confirm what happened. That holds even when someone has a motive to lie. The difference sounds academic, at first. Then a regulator or a customer's security team asks the hard question. How do you know a given log is complete and unedited?
With table-based logs, the honest answer is that the platform says so. The platform that ran the agent also wrote the log. It stores the log and holds the keys too. That is the system under audit acting as its own only witness. In most other fields, we would call that a control weakness. You do not let a trader reconcile his own book. Separation of duties exists for a reason. Self-checking is not proof, no matter how detailed it is.
This is not a knock on Databricks alone. It holds for any native data. That data lives in storage the platform can change. Such logs are rich and useful for engineering work. But they are not the artifact you hand an auditor. Not with a straight claim that no one could have touched it.
Agent goal hijack is not theoretical, and that raises the stakes on logging
Why obsess over whether a log could be altered? Because the attacks that make logs matter are already working. In December 2025, the OWASP GenAI Security Project released its Top 10 for Agentic Applications. The number one category is ASI01 Agent Goal Hijack. It folds prompt-injection-style manipulation into a broader problem. Picture a hacker rewriting an agent's goal. That manipulation can hide as instructions in documents, emails, or RAG results. Simon Willison's lethal trifecta is the crisp version of why this bites. An agent turns dangerous once three things line up. Access to private data. Exposure to untrusted content. A channel to send data back out.
How well does the attack work? In January 2025, NIST worked with the UK AI Security Institute. They used an enhanced version of the AgentDojo framework. Red-teamers tailored attacks to how LLM agents actually behave. Task-hijacking success climbed from an 11% baseline to 81%. That is roughly a sevenfold jump over the base rate. Check this figure against the NIST source before you quote it. NIST's broader adversarial ML taxonomy, AI 100-2e2025 is from March 2025. It catalogs the same family of techniques. That includes indirect prompt injection and memory poisoning. It also includes supply-chain attacks on agent tools.
The attack has also left the lab at least once. EchoLeak (CVE-2025-32711) was a zero-click indirect prompt injection in Microsoft 365 Copilot. Aim Labs (Aim Security) disclosed it in June 2025. One crafted email made Copilot read internal files. It then tried to send them out, with no user interaction. The linked arXiv paper is a secondary write-up. It is not the primary Aim Labs advisory or the CVE record. So treat its finer details with care. It was responsibly disclosed, with no known exploitation in the wild. Microsoft patched it server-side. Sources report the exact patch date inconsistently. So hold onto the June 2025 disclosure as the fact that matters.
It was a proof of concept, not an active breach. Even so, it is a clean demonstration of ASI01. This is exactly when your audit trail matters most.
How a tamper-evident audit chain actually works
The mechanism is deliberately simple. That simplicity is much of why it holds up. Give each log record a SHA-256 digest. That digest includes the previous record's digest too. The sequence then becomes append-only, in a way you can verify. Change one record, drop one, insert one, or swap two. Each later digest stops matching. Anyone who recomputes the chain can see the break.
You then sign the head of the chain from time to time. Say, with an Ed25519 checkpoint. That signature lets a third party check it holds. You need no trust in the log's host. They need no access to your storage. You need no assumption of good intent, either. They recompute the chain. Then they check it against the signed checkpoint. The result is a plain yes or no. The AuditableLLM framework shows this well. MDPI's Electronics journal published it. It's a hash-chain-backed, compliance-aware audit trail for LLM systems. It supports the general mechanism, not any one vendor. It shows the approach works. Does the venue matter to your assurance team? Then confirm its review status directly.
This is the exact property missing from Unity Catalog. Its inference tables don't give it to you. Nor do its audit logs or MLflow traces. That is not a failing on Databricks' part. No one built ordinary lakehouse tables to prove themselves.
Limit what a hijacked agent can reach
There is a second half to this that pure logging cannot cover. It changes where you spend effort. Most agent-safety work tries to constrain how the model behaves. That is useful. But a hijacked agent does just what the hacker wants. So behavioral guardrails end up fighting on the hacker's terms. A better move exists. Limit what the agent can reach, from the start.
That is the data-plane angle. It is where DataShield lives. Tokenize sensitive fields at ingest. A hijacked agent reading a record finds tokens instead. Not the raw data. The PII sits behind those tokens. Authorize per tool call instead of per session, with mid-session revocation. That way, the system cuts a compromised agent's reach the instant something looks wrong. Then seal each of those calls into a tamper-evident audit chain. For the full shape of it, see the architecture. The ontology model lays out how reachability is scoped. /security covers the field-level tokenization. None of this replaces behavioral guards. But guards fail sometimes, and NIST's numbers say they will. When they do, the blast radius stays limited to tokens. The evidence stays intact too.
Agent governance best practices worth adopting now
These are the practices the sources point to. Rough order of leverage, below.
- Strong identity for agents. The Model Context Protocol authorization spec, revision 2025-11-25 is the current bar. HTTP-based MCP servers act as OAuth 2.1 resource servers. Clients must implement PKCE. They should use the S256 method when they can. Clients must also send the RFC 8707
resourceparameter. That ties tokens to one server. Servers must reject tokens not issued for them. They must never pass tokens through to upstream APIs. That is how you avoid the confused-deputy problem. Watch the details: authorization is optional in MCP. STDIO-transport servers take credentials from the environment instead. So cite the specific 2025-11-25 revision, not "the MCP spec" in general. - Least privilege and human-in-the-loop. Scope tools tightly. Require human approval on high-impact calls. It is an unglamorous control that works.
- Input and output isolation with content provenance. This is your blunt instrument against indirect prompt injection. It comes straight out of NIST AI 100-2e2025 and OWASP ASI01 and ASI06.
- Tamper-evident logging. The hash chain plus signed checkpoints from the section above. This is the one most teams skip. Most later regret it.
Want the taxonomy of how these pieces differ from gateways and guardrails? I wrote that up separately in agent governance vs gateways vs guardrails.
Unity Catalog vs independent governance: bring both
None of this argues for ripping out Unity Catalog. It argues against one thing. One system should not both act and audit the action. Keep Unity Catalog for what it is genuinely great at. That's policy, model governance, and deep insight. It covers each agent framework you run. Then place an outside, cryptographically tamper-evident layer beneath it. Have an outside party hold it, not the platform under audit.
A fair question: why isn't that layer just another flavor of self-checking? The answer is mechanical. It runs as its own control plane, with its own signing keys. So one party proves the record is unaltered. That party is not the platform that ran the agent. That independence is the whole product. DataShield's evidence layer chains each entry with SHA-256. It anchors each entry with Ed25519 checkpoints. It returns named verdicts, not a bare valid or invalid. It reports tampering, insertion, deletion, and truncation as distinct, named outcomes. You can run that check yourself against a sample chain at /verify. No sales call required.
To be clear about the claim: this is a technical property. Namely, independent hash-chained verification you can check for yourself. It is not a compliance certification. DataShield does not hold a SOC 2 attestation yet. It is better to say so than to let you assume otherwise.
Say your regime cares about immutable, complete logging. EU AI Act Article 12 is the concrete example. It requires high-risk AI systems to allow automatic recording of events. Those are the logs, kept over the life of the system. We go deeper on that in a linked breakdown. Here is the real gap. A platform stores a table. An outside party proves the record is unaltered. That is the gap between data and proof. Want to scope what that would look like against your own agent traffic? The quote worksheet is the fastest way in.
Watch: related explainers
Databricks governance for AI agents. Plus the injection risk that makes independent proof matter.
Frequently asked questions
Is Unity Catalog tamper-evident?
Not on its own. Unity Catalog stores AI Gateway inference tables, audit logs, and MLflow traces. They sit as plain lakehouse tables. One platform governs all of them. Databricks calls the tables audit-ready, but audit-ready is not tamper-evident. Its primary blogs never claim the tables are hash-chained or signed. They are strong on insight. They are also changeable and non-cryptographic. So on their own, they cannot prove one thing. They cannot prove a record stayed unchanged.
What is agent goal hijack?
It is ASI01, the top category in the OWASP Top 10 for Agentic Applications. That list came out December 9, 2025. A hacker alters an agent's goal or call path using bad content. That can mean hidden instructions in documents, emails, or RAG results. In January 2025, NIST and the UK AI Security Institute ran a test. They tried tailored versions of this attack. Task-hijacking success climbed from an 11% baseline to 81%.
Does DataShield replace Unity Catalog?
No. It complements Unity Catalog. You keep Unity Catalog for policy, model governance, and insight. You add an independent, tamper-evident evidence layer beneath it. Think hash-chained SHA-256 records with Ed25519 checkpoints. A party outside the platform being audited holds it. It runs as its own control plane with its own signing keys. That makes it independent, not self-checking. The goal is separation of duties, not replacement.
What makes an audit log independent?
Independence means a party other than the log's own source can check it. They confirm it wasn't changed, without trusting that source. In practice, each record is chained with SHA-256 to the prior record. The chain head is signed from time to time. A third party can then recompute the chain and check the signature. Say the platform under audit is the only witness. Or say it holds the signing keys. Either way, the log is not independent.
Do I need OAuth 2.1 for MCP servers?
The MCP authorization spec, revision 2025-11-25, sets the rule. HTTP-based MCP servers act as OAuth 2.1 resource servers. Clients must use PKCE. They use the S256 method when they can. Tokens must also tie to one server, via the RFC 8707 resource parameter. Authorization is optional overall, though. STDIO-transport servers take credentials from the environment instead. So it depends on your transport. For HTTP servers handling sensitive tools, it is the current bar.