How Do You Prove a Human Checked an AI-Assisted Clinical Decision?
When an AI helps make a clinical decision, what record proves what the AI did, whether a person checked it, and that no one altered the record afterward? This note describes that gap in plain terms, the instrument that closes it, and — just as carefully — what that instrument does not claim to do. It is written for clinicians, ethicists, lawyers, data scientists, and kaitiaki, not for engineers. It names no requirement that your own governance principles have not already named first.
1. The gap
AI is already in New Zealand clinical care — drafting consultation notes for emergency and general-practice clinicians, summarising, suggesting. In almost every safe deployment, the stated safeguard is the same and it is a good one: a clinician reviews and confirms. The human holds the pen.
But there is a quiet assumption underneath that safeguard. An AI can produce an answer that is confident, fluent, and plausible — and not grounded in the patient in front of it. (The technical word is confabulation: a well-formed output that isn't anchored to fact.) Under time pressure, in a busy clinic, the review that is supposed to catch this can become a glance. None of that is unusual or blameworthy; it is how pressure works on people.
The danger is not that the AI was wrong. The danger is that, right now, in most deployments, there is no record of what the AI produced, whether the review actually happened, or proof that none of it was changed afterward. The decision is documented; the AI's part in it, and the integrity of that account, usually is not. The blindness is the problem — not any single error.
2. What exists today, and what is missing
Most systems already log. The finished note is saved; timestamps exist; the record can be produced on request. That is genuine and useful. But a log written by the same system whose conduct is in question answers a narrower thing than people assume.
"We logged it" is not the same as "we can prove it." A log that can still be added to, edited, or tidied — after an incident, during a complaint, in the calm of hindsight — cannot establish that what it shows is the complete account as it stood at the moment of care. That is the missing property. It is not more logging. It is a different kind of record.
3. The instrument
The instrument is a small, deliberate addition to whatever AI is already running. It does not change the AI, the clinician's judgment, or the workflow. It produces, alongside the decision, a record with three properties. We call each entry a receipt — in the everyday sense: proof that a thing happened, the way a receipt proves a transaction.
- The record — what the AI did, and in what order. The AI produced a draft; a person reviewed it; a decision was confirmed. Each step is its own receipt.
- Tamper-evident — once the sequence is finished it is sealed: any later addition, alteration, or reopening leaves a detectable break in the record. The account is fixed at the moment of care.
- Independently verifiable — anyone can check the record is genuine and unaltered on their own, offline, without having to trust — or even contact — the people who produced it.
The third property is the one that matters most for accountability. A record the operator can verify only by asking the operator is a promise. A record a patient's advocate, an auditor, or the Health and Disability Commissioner can verify themselves is evidence.
4. A worked example — the AI scribe
Consider an AI scribe drafting a consultation note. Today: the scribe drafts, the clinician edits and signs, the note is saved. If a question arises months later — was the AI's draft actually reviewed, or waved through? — the honest answer is usually that the record cannot say, and could in principle have been edited since.
With a sealed record, the same consultation also produces a short, fixed account. In plain language, it reads like this:
The record need not contain the clinical content itself to do its work — it can reference the note without copying it, so the sensitive material stays where it belongs (see §7). What the record establishes is the shape and integrity of the decision: that the AI's part happened, that a review step happened, and that the account is the original, not a later tidy-up.
5. What this proves — and what it does not
This is the most important section, and the one we ask you to hold us to. An instrument that overstates itself is worse than none, because it invites a trust it has not earned.
| It proves | It does not — and does not claim to |
|---|---|
| That the AI produced output, and what step it occupied in the decision. | That the AI's output was clinically correct. The record is silent on clinical truth. |
| That a review step occurred, and in what order. | That the review was thorough. It records that the step happened, not the quality of the clinician's attention. |
| That the account is complete and unaltered since the moment of care. | That any harm was prevented. This is an evidence instrument, not a safety control on the AI itself. |
| That an independent party can confirm all of the above without trusting the operator. | That the AI is unbiased or equitable. Those are properties of the model and its data, addressed elsewhere — not by this record. |
In short: the instrument makes the decision accountable and reviewable after the fact. It does not make it correct, safe, or fair — those remain the work of clinicians, model developers, and your own governance. We think being exact about this boundary is what makes the instrument trustworthy.
6. How this relates to the governance framework's eight domains
The Waitematā AI Governance Group's published framework evaluates any proposed AI across eight domains. This instrument speaks directly to two of them and supports one more. On a fourth — Māori perspectives — it is careful to disclaim more than it offers. It is honestly silent on the remaining four, which concern the AI itself, not the record of its use.
| Governance domain | How the record relates |
|---|---|
| Ethical principles, including transparency | Directly. The AI's role in a decision becomes legible and checkable after the fact, by people outside the system that produced it. |
| Legal and contractual requirements | Directly. The framework describes this domain as "the need for clear accountability and responsibilities," the service's role as "guardians (kaitiakitanga)" of health information, and "responsibility for ongoing monitoring and audit, accountability if an AI tool should fail." A sealed, independently verifiable record is precisely that artefact: it fixes who did what and when, it is the monitoring-and-audit trail itself, and it survives a vendor being sold — because it can hold none of the clinical data itself, referencing rather than copying it, and verifies against published keys. |
| Māori perspectives | Does not provide sovereignty — and does not claim to. This instrument does not offer data or language sovereignty; that is not what it is. The one honest, narrow thing it does for this domain: the record proves what happened by referencing clinical material rather than copying it, so using the record does not itself move data out of the system that already holds it (see §7). Control over where health data lives, and who may access it, stays exactly where it already sits. |
| Technical guidance | Supports. The record is a cryptographic, security-grade artefact — signed and verifiable offline — aligned with the cyber-security considerations this domain raises. |
| Consumer perspectives · Equity and fairness · Clinical perspectives · Data issues | Not addressed by this instrument. These concern the AI and its deployment, not the record of its use. We don't claim otherwise. |
7. What the record holds — and what it does not
A record of accountability must not become a new place where sensitive information is copied or stored. So the instrument is built to hold as little as possible. A receipt can establish that a decision occurred, and prove its integrity, by referencing the clinical material — for example, by a one-way fingerprint of it — without containing the material itself. The patient's information stays in the system that already holds it, under the authority that already governs it. The record proves the event; it does not relocate the data.
To be plain about the boundary: this is not a data-sovereignty tool, and it does not claim to be. It does not decide where health data lives, or who may access it — those remain the responsibility of the systems and authorities that already hold them. Data sovereignty is a real and unresolved question; this instrument does not answer it and does not pretend to. What it offers is narrower and honest: a record of what happened that does not itself move the data anywhere.
8. Who built this, and why
This was built in Hokianga, by a New Zealand company, by someone who carries a kaitiaki responsibility for what gets recorded and how. The honest origin is personal: watching a clinician reach for an AI during a child's check-up, and realising that whatever the AI contributed, nothing in the room would ever record that it had. Not a scandal — a blind spot. The instrument is an attempt to fit the missing earth pin, quietly, to a system that is otherwise working.
We are not claiming to have found a problem the field has missed; clinicians, the HDC, and your own group have all named the accountability question already. We are offering one concrete, verifiable way to answer part of it — and we would rather be corrected early than be polite and wrong. If something here is overstated or mistaken, we want to hear it.
9. See it for yourself
The instrument is real and running, not a slide. The records it produces are signed and can be checked by anyone, with no login, against published keys:
This note is offered, not pitched. If it is useful to your work, or if you can see where it falls short, the door is open: hello@agenticrail.nz.