AI in New Zealand Health Care: The Missing Evidence Layer
New Zealand has deployed artificial intelligence into clinical documentation at national scale, and AI-guided treatment is now in clinical trial. The official safety position rests on two assurances: that "the doctor reviews and confirms" the AI's output, and that the tools "meet all privacy requirements." This brief documents — from primary New Zealand sources — that neither assurance is currently recorded: there is no sealed, tamper-evident, pre-execution record of what an AI produced, nor of whether a human verified it. The brief makes no claim of harm and alleges no breach of law. It documents a structural gap, and identifies the instrument that closes it.
1. Scope
This is a sector brief, not a regulation, and it confers no compliance. It establishes three things from cited New Zealand sources: (a) the scale and form of AI deployment in NZ health care as of mid-2026; (b) the precise point at which the deployment's stated safeguards become unverifiable for want of a record; and (c) the structure of the evidence layer that would make those safeguards provable. Throughout, a careful distinction is kept between a record (which may exist and still be revisable or absent) and evidence (a sealed, contemporaneous, independently verifiable record). The gap is in the second.
2. The Deployment — what is in place
AI scribes, nationally. As of 2026, AI "ambient" scribes — tools that record a consultation and automatically draft clinical notes, referral letters and summaries — are in use by approximately 1,250 emergency-department clinicians across all public EDs, around 250 more than the initial October 2025 target, with a further 1,000 licences being procured for mental-health teams [11][1][2]. Ambient scribes are endorsed for use in Health New Zealand by its National Artificial Intelligence and Algorithm Expert Advisory Group. A Hawke's Bay pilot reduced documentation time from around 17 minutes to just over four minutes per patient [2][3].
AI-guided treatment, in trial. A New Zealand-led clinical trial across roughly 50 intensive-care units in NZ and Australia, will recruit more than 24,000 patients to test whether AI can guide the treatment of critically ill patients on life support ($5M Health Research Council grant, awarded 11 June 2026) [4]. Recruitment has not opened and the trial is not yet registered. This is the highest-consequence end of the spectrum: decisions that cannot be taken back.
The official safeguard. The stated workflow is that the scribe produces a draft and "the doctor reviews and confirms" it; the responsible Minister has stated that "AI will never replace clinical skill or judgement" and that the tools "meet all privacy requirements" [1]. The entire safety case rests on the human-in-the-loop review and on privacy compliance.
3. The Gap — where the safeguards become unrecorded
The safety case depends on (a) a human reviewing the AI's output, and (b) that output being handled safely. Neither is captured in a sealed, pre-execution, tamper-evident record. There is no contemporaneous evidence of what the AI produced, whether a clinician reviewed it, how closely, or on whose authority a resulting decision was made. "The doctor reviews and confirms" is, at present, an assurance with no instrument behind it.
Four documented findings show the gap is not theoretical:
3.1 — Unrecorded consumer-LLM use in clinical notes. In March 2026, Health New Zealand mental-health and addiction staff were found to be using free, general-purpose chatbots — ChatGPT, Claude and Gemini — to draft clinical notes, in some cases transcribing the output into the record. A Rotorua Lakes district memo dated 26 March 2026 warned of disciplinary action, citing "data security, privacy and accountability"; HNZ's director of digital innovation and AI confirmed the tools "presented risks to data security, privacy and accountability." The reporting records no mechanism that captured what those tools produced or whether it was verified before entering a patient's record [5]. This is the evidence gap in its rawest form: an AI materially shaping a clinical note, leaving no sealed trace.
3.2 — Consent and oversight are patchy, and unrecorded. An Otago survey of 197 New Zealand primary-care providers, conducted February–March 2024, found 40% (n=70) had used an AI scribe. Among those reporting, 59% (n=35) sought patient consent and 66% (n=41) had read the tool's terms [13]. The Medical Council guidance requires informed consent for scribe use, and that whether consent was obtained be documented (cl. 9–10) [8]. Where practice diverges from that requirement, the absence of a per-encounter sealed record means the divergence cannot be detected, audited, or disproved after the fact.
3.3 — A complaint is anticipated. A clinical lead at Whakarongorau has stated that a complaint to the Health and Disability Commissioner over AI-scribe use without informed consent is "only a matter of time" [6]. A complaint is the fault event — the moment at which the absence of a contemporaneous, sealed record stops being abstract and becomes the difference between a defensible account and an unprovable one.
3.4 — A security flaw has already occurred. A security flaw in a Health NZ AI tool was reported in March 2026 [7]. Whatever its scope, it establishes that the systems holding and processing clinical AI output are themselves subject to compromise — which is precisely the condition under which an externally signed, tamper-evident receipt, rather than a system-internal log, is the only record that still stands.
4. Why "Review and Confirm" Is Not Yet Evidence
The human-in-the-loop is the load-bearing safeguard, and under time pressure it is the most fragile. Pilot data cited in support of the rollout notes that scribes let doctors see, on average, one additional patient per shift [1] — the same time saving that compresses the "review" of an AI draft toward a confirmation click. Whether a given confirmation was a considered clinical judgement or a reflex under load is exactly the fact that determines accountability if something goes wrong — and it is exactly the fact that nothing currently records.
The Medical Council's own guidance makes the review a professional obligation, not a courtesy: it states that AI "may produce inaccurate or fabricated information," and that a doctor "should check the accuracy of any AI output and confirm it is appropriate for the individual patient before using it for patient care or including it in patient records" [8]. The duty to verify is explicit. What is absent is any contemporaneous, tamper-evident record of whether the verification actually happened — leaving the central safeguard asserted but unprovable.
A sealed pre-execution record resolves this without trusting anyone's memory: the time spent on a draft, the edits made or not made, and the explicit authority under which a decision proceeded, fixed at the moment it happened and verifiable afterward. It does not assume the review was real. It records whether it was.
5. Relationship to New Zealand's Existing Framework
This brief operates beneath — not in place of — the instruments already governing the field. It restates none of them and claims conformance to none.
| NZ instrument | What it requires | Where the evidence layer sits |
|---|---|---|
| Medical Council of NZ — Guidance on using AI in patient care (10 Mar 2026) [8] | The doctor "remain[s] responsible for all your clinical decisions and actions"; AI "may produce inaccurate or fabricated information," so the doctor "should check the accuracy of any AI output and confirm it" before use (cl. 4). AI use that influences decisions must be documented in the patient's notes (cl. 5). Informed consent for scribe use must be obtained, and whether consent was obtained must be documented (cl. 9–10). Only endorsed AI may be used, or the doctor must assure its safety (cl. 11). | Every one of these obligations — the accuracy check, the consent, the documentation — is currently discharged into the revisable patient record, or not recorded at all. A sealed receipt makes the Council's own requirements provable rather than merely asserted: it fixes, at the moment of the decision, that the check happened, that consent was taken, and what the AI produced. |
| NZ AI Pre-Implementation Evaluation Framework (endorsed by the Health NZ Board, 28 Jul 2026) [14] | The methodology the National Artificial Intelligence and Algorithm Expert Advisory Group uses to review AI tools before they are implemented. | It evaluates a tool before deployment. It does not record what any deployed tool then did, on which patient, or whether the required human step occurred — which is the interval this brief describes. |
| Health Information Privacy Code 2020 (incl. IPP3A, Rule 3A in force 1 May 2026) [9] | Governs how patient information may be collected, used and disclosed. | A pre-execution receipt records, at decision time, what data an AI step touched and under what authority — independent of the AI system being governed. |
| Health & Disability Commissioner [6] | Adjudicates complaints about the quality and safety of care, including consent. | The sealed record is the artefact that makes a consent-and-oversight account provable when a complaint arrives. |
| GPNZ AI-in-primary-care working group [10]; Health NZ generative-AI advice | Developing sector guidance on safe AI use. | The completeness requirements (§6) offer a neutral technical specification of the "traceable, tamper-evident record" such guidance presumes but does not yet specify. |
6. The Instrument — a sealed pre-execution receipt
The missing layer is specified, neutrally and in full, in the companion Completeness Specification: an enforcement record is evidence-grade only if it is created before the action (R1), independently of the system being recorded (R4), cryptographically signed and verifiable offline (R3), and — the dividing line — sealed so the account is fixed in time and cannot later be added to, altered, or reopened without detection (R8). A logging-grade record proves nothing was secretly rewritten; only a sealed record proves the account is complete and was fixed at the time — including when the operator is the party later under examination.
At the moment an AI step runs — a scribe drafting a note, a model returning a suggestion — an external gate writes a sealed receipt recording what was produced, what evidence (if any) the clinician reviewed, the time and authority of the human confirmation, and a cryptographic chain to the prior step. The clinician cannot alter it; the AI cannot author it; anyone can verify it offline against a published key, with no call back to the vendor. The verification is automatic — a machine check returning a verdict, requiring no effort from the busy human it protects.
slp8_receipt_v2, AgenticRail's production receipt schema, is offered as one conformant reference implementation. It is named here as the author's own; the specification is implementation-independent, and any vendor's record can be assessed against the same eight requirements.
7. A Deliberate Boundary
This brief asserts no harm and no breach of law or duty by any named body, clinician or vendor. The clinicians described are operating under genuine workload pressure with tools their system endorsed. The brief documents one structural fact: that the safeguards the deployment relies on are not, at present, captured in evidence-grade records — and that the blindness this creates is the danger, independent of whether harm has yet occurred. The argument is for an instrument, not against a person.
8. References
The v1.0 and v1.1 fingerprints above are permanent and unchanged. v1.2 corrects how §3.2 reports the Otago survey. v1.1 described it as a survey of 197 providers using AI scribes and stated that 41% were not seeking explicit patient consent. 197 is the number who completed the survey; 40% (n=70) had used a scribe, and the consent figure is 59% (n=35) among those reporting. The 41% was derived by subtraction and applied to the wrong base, so it is withdrawn rather than restated. The survey ran February–March 2024. v1.2 also removes an uncited sentence claiming roughly half of New Zealand GPs used a scribe by early 2026 — no primary source states this, confirmed by three independent checks — corrects the author list at [13] to Ballantyne, Style, Stubbe, Murton and Dowell, narrows [6] to the quotation it actually carries and dates it to 25 June 2026, and removes an approval date at [8] that appears in no primary Medical Council document. The gap this brief describes, and every conclusion drawn from it, are unchanged. ⚠️ Unlike v1.1, this version changes a canonical token as well as the version and date, because the superseded statistic sat inside the preimage. No earlier version is edited; v1.2 is a separate, independently-fingerprinted record.
The v1.0, v1.1 and v1.2 fingerprints above are permanent and unchanged. v1.3 corrects a single date introduced in v1.2 earlier the same day. Reference [6] was dated 25 June 2026, which is the publication date of the syndicated Pharmacy Today version of that story; the New Zealand Doctor article the reference actually cites was published 24 June 2026, verified at the source. No figure, finding or conclusion changes, and only the version token differs in the canonical string. No earlier version is edited; v1.3 is a separate, independently-fingerprinted record.
The v1.0 to v1.3 fingerprints above are permanent and unchanged. v1.4 makes three changes. It removes a named list of four endorsed ambient scribes: the endorsement mechanism is real and is named here instead, but the specific vendor list could not be traced to any Health New Zealand primary source, and one of the four could not be sourced at all. The verified Hawke’s Bay pilot figure is retained. It corrects the tense of the REVOLUTION trial description — the Health Research Council and MRINZ both say the trial will recruit; recruitment has not opened and no registration exists. And it adds the New Zealand AI Pre-Implementation Evaluation Framework, endorsed by the Health NZ Board on 28 July 2026 as the methodology its expert advisory group uses to review AI tools, together with the in-force date of Rule 3A. The framework evaluates tools before deployment and so sits alongside, rather than closing, the runtime gap this brief describes. ⚠️ A canonical token changes as well as the version, because the withdrawn vendor list was inside the preimage. No earlier version is edited; v1.4 is a separate, independently-fingerprinted record.