Document type Sector Note — Evidence Brief
Subject Evidence for assessment decisions in which an automated system marked or assisted the marking
Published by TUARA KURI LIMITED — trading as AgenticRail, Hokianga, Aotearoa New Zealand
Date 2026-07-26
Version 1.0
Status Published — open for citation
Related AI in New Zealand Education Assessment · Automated Decisions and the Provable Safeguard · Completeness specification

When the Marker Is a Machine: What Evidence Supports an AI-Assisted Assessment Decision?

New Zealand marks national writing assessments with automated scoring, quality assured by human check-marking concentrated at the achievement boundaries. The safeguard is real, it is publicly described, and it is better documented than most comparable programmes anywhere. This note does not question whether it happens. It asks a narrower and harder question: when one student challenges one result, what evidence shows the check ran on that script? An agreement rate is a property of a system measured across a population. A challenge is about an instance. Those are different objects, and only one of them is currently recorded.

1. Scope and definitions

This is a sector note. It is not a regulation, it confers no compliance, and it asserts no failure by any New Zealand agency, school, teacher, or marker.

Three terms are used throughout in a fixed sense, because the whole argument depends on holding them apart.

A control is a step required by policy or design — for example, a human reviewing an automatically generated score before a result is released.

A record is any account that a control occurred. A record may be real and still be self-attested, revisable after the fact, incomplete, or produced on request rather than at the time.

Evidence is a record that was created before or at the moment of the action, sealed so it cannot afterwards be added to or altered without detection, and verifiable by a party who does not have to trust the system that produced it.

A control can be genuinely operating and produce no evidence in this sense. That is not a contradiction and it is not an accusation. It is the ordinary state of most automated systems, and it is the only thing this note is about.

2. What is already running

The following are public facts, cited in §8. They are stated here without inference.

2.1 — Automated marking is in national production, not pilot. NZQA's Automated Text Scoring (ATS) ran a large-scale trial of approximately 35,000 writing responses in September 2024, reporting roughly 80% agreement with human markers [1][2]. From May 2025, automated scoring has been applied to all digitally submitted writing assessments, with more than 55,000 assessments marked, and results returned about three and a half weeks faster than the previous cycle [3].

2.2 — Human marking is retained as the named quality-assurance control. Human marking is retained at the achievement boundary, covering over a third of cases, and the human score overrides the automated score where the two differ [3]. This is a deliberate and defensible design: it concentrates scarce human attention at the point where a score change alters the outcome for the student.

2.3 — The programme is expanding under contract. NZQA released RFP 33045390, "NZQA AI Marking Model Design and Implementation," on 20 November 2025; it closed on 5 December 2025 and was awarded to Amazon Web Services on 17 February 2026 [4]. The stated scope is to design and implement an AI-based system for marking, to enhance the accuracy, efficiency and fairness of assessment processes.

2.4 — Automated assistance is planned for moderation, not only marking. NZQA's September 2025 publication on its responsible use of artificial intelligence names moderation support for internally assessed work among its planned tools [5]. Moderation is the check applied to marking. §5 returns to what follows from that.

2.5 — The direction is publicly committed at ministerial level. In August 2025 the responsible Minister stated publicly that AI marking is "as good, if not better than human marking" and described New Zealand as extraordinarily advanced relative to the rest of the world [6]. This note takes that commitment at face value and is addressed to what would substantiate it under challenge.

3. The safeguard is real. The question is whether it is provable.

The structural observation

The quality-assurance mechanism for automated scoring is human check-marking. If a party outside the operating authority asks, of one specific result, "show me that the check ran on this script, and that it ran before the result was released," the answer available today is the authority's own account of its own process. There is no independent verification body for this control, no external attestation of an individual result, and no mechanism by which an appeals panel, an Ombudsman, or a student's representative could confirm the step occurred without re-trusting the same organisation whose process is in question. This is not unusual and it is not alleged to be a breach of anything. It is the standard architecture, and it is the specific thing that a challenge tests.

The distinction being drawn is narrow and worth stating precisely. Nothing here suggests check-marking does not occur. The claim is that "it occurs" and "its occurrence on a given script can be confirmed by someone outside the agency" are two different properties, that a system can hold the first without the second, and that only the second survives a dispute in which the agency's own account is what is being contested.

4. Why an agreement rate cannot answer a single case

An agreement rate — roughly 80%, in the trial figures above — is a statistical property of a system, measured across a population of scripts. It is a good and appropriate measure of whether an automated scorer is fit to deploy. It is the right instrument for the question it answers.

It carries no information about any particular script. A system with a high agreement rate still produces individual results that differ from what a human marker would have given, and the aggregate figure cannot identify which ones. This is a property of aggregate measures generally, not a defect in this one.

The consequence is specific. When a student challenges a result, the question in front of the reviewer is not "is this system accurate overall." It is:

The question a challenge actually asksWhat an aggregate agreement rate can sayWhat a sealed per-script record can say
Was this script scored by the automated system, by a human, or by both? Nothing about this script. Which steps ran, in order, with times.
Did the human check-marking step occur on this script? Nothing about this script. Whether it occurred, and the identity and role of the reviewer.
Did it occur before the result was released, or after the challenge was raised? Nothing about this script. The ordering, fixed at the time and not reconstructable afterwards.
If the two scores differed, was the human score the one applied? Nothing about this script. The decision recorded at the moment it was taken.

The boundary-concentration point cuts both ways, and honestly. Because human check-marking is concentrated at the achievement boundary, a script well inside a grade band may correctly have received no human review. That is the design working as intended, and it is defensible. But it means "was this checked by a human" has two legitimate answers, and the difference between them is currently invisible from outside. A record that shows which regime a given script fell under converts a potential accusation into a documented and defensible design decision. This is the case for the record, not against the design.

5. When the check itself is assisted

Moderation is the control applied to marking — the layer that provides assurance that marking was done properly. Where an automated tool assists moderation [5], the assurance layer acquires the same property as the layer beneath it: it now performs a step whose occurrence, on any particular case, is recorded only by the party performing it.

This does not make such a tool inappropriate. Moderation at national scale is genuinely burdensome, and the sector has said so in its own submissions on the replacement of NCEA, describing consistently raised concerns about over-assessment and the heavy demands of moderation, and asking specifically for secure platforms for assessment and moderation [7].

What follows is narrower: the evidence question moves up a level rather than being answered. If marking is checked by moderation, and moderation is assisted by a system whose own steps are self-recorded, then the chain of assurance terminates at an account rather than at a record. It stops being recursive only when some step in the chain produces evidence that does not depend on the account of the party being checked.

6. What the record would need to be

The requirements are specified neutrally, and independently of any vendor, in the companion Completeness Specification. Four of the eight are load-bearing here.

RequirementWhy it matters for a contested grade
Created before or at the moment of the stepA record written after a challenge is raised cannot establish the order of events, which is the fact most challenges turn on.
Independent of the system being recordedIf the marking system writes its own account of whether it was checked, the account and the subject are the same party.
Signed and verifiable offlineAn appeals panel or Ombudsman can confirm the record against a published key without a request to, or cooperation from, the authority.
SealedOnce the sequence closes, the account is fixed. Later additions are detectable rather than silent, which is what makes the record worth anything in a dispute.
Applied to check-marking

At the moment a marker opens a script flagged for check-marking, an external gate writes a signed receipt recording the script identifier, the reviewer's identity and role, the automated score presented, the time, and the position of this step in that script's assessment sequence. When the score is confirmed or overridden, a second receipt records which. The sequence is sealed when the result is released. The authority cannot alter the record afterwards without the alteration being detectable; a challenge two years later is answered by verification rather than by recollection; and the marker is protected as much as the student, because "the check was never done" becomes a checkable claim rather than an accusation that cannot be disproved. No additional work is asked of the marker — the receipt is a by-product of the step, not a second task.

slp8_receipt_v2, AgenticRail's production receipt schema, is offered as one conformant reference implementation. It is named here as the author's own. The specification is implementation-independent, and any vendor's record — including one built under an existing contract — can be assessed against the same requirements.

7. What this does not do

Each of the following states something true about the record, followed by its limit.

The record establishes what happened and in what order. It does not establish that the result was correct. A receipt showing that a human reviewed a script says the review took place. It says nothing about whether the reviewer was right, attentive, or qualified for that subject. Marking quality is a separate problem with separate instruments.

Tamper detection has a boundary. A sealed, chained record makes later alteration detectable. Detectable is not impossible: a party holding the signing keys could produce a consistent but false chain. What closes that gap is custody — a copy held by someone who is not the operator — not stronger cryptography.

The record narrows a dispute; it does not resolve it. If a marking rule or model is itself flawed, faithful execution of that rule produces impeccable receipts of a flawed rule being applied. The value is that the argument moves from "what happened" to "was the rule right," which is the argument worth having.

This note asserts no failure by any New Zealand agency. NZQA has published more about its automated marking, its agreement rates, and its human-oversight arrangements than most comparable authorities anywhere. That transparency is what makes this note possible to write from public sources, and it is the reason the gap discussed here is a design question rather than a complaint. No individual is named.

8. How to check every claim on this page

Every factual claim in §2 is cited below to a primary or named public source. The claims about the instrument can be checked directly, without contacting anyone:

If a claim on this page does not match the live system, the live system is the authority and the page is wrong.

9. References

[1] RNZ, "Artificial intelligence exam-marking on the way for Year 10 writing tests" — rnz.co.nz
[2] SchoolNews NZ, "NZQA: AI-marking now a reality" — schoolnews.co.nz
[3] NZQA, "Embracing AI in student assessments" — nzqa.govt.nz
[4] GETS (Government Electronic Tenders Service), RFx ID 33045390, "NZQA AI Marking Model Design and Implementation," awarded 17 February 2026 — gets.govt.nz
[5] NZQA, "NZQA's Responsible Use of Artificial Intelligence," September 2025 — nzqa.govt.nz
[6] RNZ, coverage of ministerial statements on AI marking, 5 August 2025 — rnz.co.nz
[7] PPTA Te Wehengarua, Submission on the proposal to replace NCEAppta.org.nz (linked from the union's Replacing NCEA campaign page)
[8] Ministry of Education / Te Poutāhū, GenAI in NCEA assessment: FAQs, March 2025 — education.govt.nz
[9] Companion brief — AI in New Zealand Education Assessment: The Missing Evidence Layer, covering student-work authenticity and the internal moderation cycle.
[10] Companion note — Automated Decisions and the Provable Safeguard, covering the asserted / enforced / provable distinction in the public sector generally.
Document Fingerprint — SHA-256 — v1.0
6b8172edeacab28bc163e5a96676a2c3648828218356969fb075bbd282e93bcc
This hash is SHA-256 of the canonical string defined below. It is reproducible independently of this page using any SHA-256 implementation.

Canonical string (pipe-delimited, UTF-8, no trailing newline):
When the Marker Is a Machine|1.0|2026-07-26|TUARA KURI LIMITED|NZQA Automated Text Scoring 35000 writing responses trial September 2024 approx 80 percent agreement with human markers|from May 2025 applied to all digitally submitted writing assessments over 55000 marked results 3.5 weeks faster|human marking retained at achievement boundary over a third of cases human score overrides where they differ|GETS RFx 33045390 NZQA AI Marking Model Design and Implementation awarded Amazon Web Services 17 February 2026|NZQA September 2025 responsible use of AI names moderation support for internally assessed work as planned tool|control record evidence are three distinct things a control can operate and produce no evidence|an agreement rate is a property of a system across a population a challenge is about an instance|aggregate accuracy cannot identify which individual results differ|boundary concentration means was this checked by a human has two legitimate answers and the difference is invisible from outside|where an automated tool assists moderation the evidence question moves up one level rather than being answered|chain of assurance terminates only at a record not dependent on the account of the party being checked|record establishes what happened and in what order not that the result was correct|tamper detection is detectable not impossible custody closes the gap not cryptography|completeness R1-R8|report.agenticrail.nz

Published: 2026-07-26  |  Version: 1.0  |  Entity: TUARA KURI LIMITED

Sourcing note: every factual claim in §2 is cited to a primary or named public source in §9. This note asserts no failure, breach, or bad faith by any New Zealand agency, school, teacher, or marker, and names no individuals.