When the Marker Is a Machine: What Evidence Supports an AI-Assisted Assessment Decision?
New Zealand marks national writing assessments with automated scoring, quality assured by human check-marking concentrated at the achievement boundaries. The safeguard is real, it is publicly described, and it is better documented than most comparable programmes anywhere. This note does not question whether it happens. It asks a narrower and harder question: when one student challenges one result, what evidence shows the check ran on that script? An agreement rate is a property of a system measured across a population. A challenge is about an instance. Those are different objects, and only one of them is currently recorded.
1. Scope and definitions
This is a sector note. It is not a regulation, it confers no compliance, and it asserts no failure by any New Zealand agency, school, teacher, or marker.
Three terms are used throughout in a fixed sense, because the whole argument depends on holding them apart.
A control is a step required by policy or design — for example, a human reviewing an automatically generated score before a result is released.
A record is any account that a control occurred. A record may be real and still be self-attested, revisable after the fact, incomplete, or produced on request rather than at the time.
Evidence is a record that was created before or at the moment of the action, sealed so it cannot afterwards be added to or altered without detection, and verifiable by a party who does not have to trust the system that produced it.
A control can be genuinely operating and produce no evidence in this sense. That is not a contradiction and it is not an accusation. It is the ordinary state of most automated systems, and it is the only thing this note is about.
2. What is already running
The following are public facts, cited in §8. They are stated here without inference.
2.1 — Automated marking is in national production, not pilot. NZQA's Automated Text Scoring (ATS) ran a large-scale trial of approximately 35,000 writing responses in September 2024, reporting roughly 80% agreement with human markers [1][2]. From May 2025, automated scoring has been applied to all digitally submitted writing assessments, with more than 55,000 assessments marked, and results returned about three and a half weeks faster than the previous cycle [3].
2.2 — Human marking is retained as the named quality-assurance control. Human marking is retained at the achievement boundary, covering over a third of cases, and the human score overrides the automated score where the two differ [3]. This is a deliberate and defensible design: it concentrates scarce human attention at the point where a score change alters the outcome for the student.
2.3 — The programme is expanding under contract. NZQA released RFP 33045390, "NZQA AI Marking Model Design and Implementation," on 20 November 2025; it closed on 5 December 2025 and was awarded to Amazon Web Services on 17 February 2026 [4]. The stated scope is to design and implement an AI-based system for marking, to enhance the accuracy, efficiency and fairness of assessment processes.
2.4 — Automated assistance is planned for moderation, not only marking. NZQA's September 2025 publication on its responsible use of artificial intelligence names moderation support for internally assessed work among its planned tools [5]. Moderation is the check applied to marking. §5 returns to what follows from that.
2.5 — The direction is publicly committed at ministerial level. In August 2025 the responsible Minister stated publicly that AI marking is "as good, if not better than human marking" and described New Zealand as extraordinarily advanced relative to the rest of the world [6]. This note takes that commitment at face value and is addressed to what would substantiate it under challenge.
3. The safeguard is real. The question is whether it is provable.
The quality-assurance mechanism for automated scoring is human check-marking. If a party outside the operating authority asks, of one specific result, "show me that the check ran on this script, and that it ran before the result was released," the answer available today is the authority's own account of its own process. There is no independent verification body for this control, no external attestation of an individual result, and no mechanism by which an appeals panel, an Ombudsman, or a student's representative could confirm the step occurred without re-trusting the same organisation whose process is in question. This is not unusual and it is not alleged to be a breach of anything. It is the standard architecture, and it is the specific thing that a challenge tests.
The distinction being drawn is narrow and worth stating precisely. Nothing here suggests check-marking does not occur. The claim is that "it occurs" and "its occurrence on a given script can be confirmed by someone outside the agency" are two different properties, that a system can hold the first without the second, and that only the second survives a dispute in which the agency's own account is what is being contested.
4. Why an agreement rate cannot answer a single case
An agreement rate — roughly 80%, in the trial figures above — is a statistical property of a system, measured across a population of scripts. It is a good and appropriate measure of whether an automated scorer is fit to deploy. It is the right instrument for the question it answers.
It carries no information about any particular script. A system with a high agreement rate still produces individual results that differ from what a human marker would have given, and the aggregate figure cannot identify which ones. This is a property of aggregate measures generally, not a defect in this one.
The consequence is specific. When a student challenges a result, the question in front of the reviewer is not "is this system accurate overall." It is:
| The question a challenge actually asks | What an aggregate agreement rate can say | What a sealed per-script record can say |
|---|---|---|
| Was this script scored by the automated system, by a human, or by both? | Nothing about this script. | Which steps ran, in order, with times. |
| Did the human check-marking step occur on this script? | Nothing about this script. | Whether it occurred, and the identity and role of the reviewer. |
| Did it occur before the result was released, or after the challenge was raised? | Nothing about this script. | The ordering, fixed at the time and not reconstructable afterwards. |
| If the two scores differed, was the human score the one applied? | Nothing about this script. | The decision recorded at the moment it was taken. |
The boundary-concentration point cuts both ways, and honestly. Because human check-marking is concentrated at the achievement boundary, a script well inside a grade band may correctly have received no human review. That is the design working as intended, and it is defensible. But it means "was this checked by a human" has two legitimate answers, and the difference between them is currently invisible from outside. A record that shows which regime a given script fell under converts a potential accusation into a documented and defensible design decision. This is the case for the record, not against the design.
5. When the check itself is assisted
Moderation is the control applied to marking — the layer that provides assurance that marking was done properly. Where an automated tool assists moderation [5], the assurance layer acquires the same property as the layer beneath it: it now performs a step whose occurrence, on any particular case, is recorded only by the party performing it.
This does not make such a tool inappropriate. Moderation at national scale is genuinely burdensome, and the sector has said so in its own submissions on the replacement of NCEA, describing consistently raised concerns about over-assessment and the heavy demands of moderation, and asking specifically for secure platforms for assessment and moderation [7].
What follows is narrower: the evidence question moves up a level rather than being answered. If marking is checked by moderation, and moderation is assisted by a system whose own steps are self-recorded, then the chain of assurance terminates at an account rather than at a record. It stops being recursive only when some step in the chain produces evidence that does not depend on the account of the party being checked.
6. What the record would need to be
The requirements are specified neutrally, and independently of any vendor, in the companion Completeness Specification. Four of the eight are load-bearing here.
| Requirement | Why it matters for a contested grade |
|---|---|
| Created before or at the moment of the step | A record written after a challenge is raised cannot establish the order of events, which is the fact most challenges turn on. |
| Independent of the system being recorded | If the marking system writes its own account of whether it was checked, the account and the subject are the same party. |
| Signed and verifiable offline | An appeals panel or Ombudsman can confirm the record against a published key without a request to, or cooperation from, the authority. |
| Sealed | Once the sequence closes, the account is fixed. Later additions are detectable rather than silent, which is what makes the record worth anything in a dispute. |
At the moment a marker opens a script flagged for check-marking, an external gate writes a signed receipt recording the script identifier, the reviewer's identity and role, the automated score presented, the time, and the position of this step in that script's assessment sequence. When the score is confirmed or overridden, a second receipt records which. The sequence is sealed when the result is released. The authority cannot alter the record afterwards without the alteration being detectable; a challenge two years later is answered by verification rather than by recollection; and the marker is protected as much as the student, because "the check was never done" becomes a checkable claim rather than an accusation that cannot be disproved. No additional work is asked of the marker — the receipt is a by-product of the step, not a second task.
slp8_receipt_v2, AgenticRail's production receipt schema, is offered as one conformant reference implementation. It is named here as the author's own. The specification is implementation-independent, and any vendor's record — including one built under an existing contract — can be assessed against the same requirements.
7. What this does not do
Each of the following states something true about the record, followed by its limit.
The record establishes what happened and in what order. It does not establish that the result was correct. A receipt showing that a human reviewed a script says the review took place. It says nothing about whether the reviewer was right, attentive, or qualified for that subject. Marking quality is a separate problem with separate instruments.
Tamper detection has a boundary. A sealed, chained record makes later alteration detectable. Detectable is not impossible: a party holding the signing keys could produce a consistent but false chain. What closes that gap is custody — a copy held by someone who is not the operator — not stronger cryptography.
The record narrows a dispute; it does not resolve it. If a marking rule or model is itself flawed, faithful execution of that rule produces impeccable receipts of a flawed rule being applied. The value is that the argument moves from "what happened" to "was the rule right," which is the argument worth having.
This note asserts no failure by any New Zealand agency. NZQA has published more about its automated marking, its agreement rates, and its human-oversight arrangements than most comparable authorities anywhere. That transparency is what makes this note possible to write from public sources, and it is the reason the gap discussed here is a design question rather than a complaint. No individual is named.
8. How to check every claim on this page
Every factual claim in §2 is cited below to a primary or named public source. The claims about the instrument can be checked directly, without contacting anyone:
- Run the live demo. It executes a real multi-step sequence against the production gate and returns a sequence identifier.
- Paste that identifier into the verification tool. It returns every receipt, each with its raw signature and the exact bytes that were signed.
- Fetch the published public keys and run standard Ed25519 verification over any receipt in your own code, offline, with no call back to the operator.
If a claim on this page does not match the live system, the live system is the authority and the page is wrong.