Document type Sector Note — Evidence Brief
Subject Establishing authenticity of student work where detection has been withdrawn
Published by TUARA KURI LIMITED — trading as AgenticRail, Hokianga, Aotearoa New Zealand
Date 2026-07-26
Version 1.0
Status Published — open for citation
Related When the Marker Is a Machine · AI in New Zealand Education Assessment · Completeness specification

If AI Detectors Don't Work, What Does? Authenticity by Process Record

There are three available responses to generative AI in assessment: prohibit it, detect it, or record the process that produced the work. New Zealand has now tried the first two at national scale and documented what each costs. Detection has been withdrawn by three universities and ruled unsuitable as a sole means by national guidance. Reverting to supervised handwriting works, and shifts the cost onto specific students. This note sets out why the first two options share a single structural weakness — both interrogate the finished artifact — and what the third option looks like when the object of examination is the sequence instead.

1. Scope

This is a sector note. It confers no compliance and asserts no failure by any institution, teacher, or student. It is written from published sources, cited in §8, and its argument is general even though its evidence is mostly from Aotearoa New Zealand — the New Zealand material is unusually useful precisely because the decisions were made publicly and the reasoning was stated on the record.

One distinction is load-bearing throughout. Artifact-based methods examine the finished work and infer how it was made. Process-based methods record how the work was made, as it is being made. Detection and prohibition are both artifact-based, which is why they fail in related ways.

2. Option one — prohibit, or revert to supervised production

This works. It is also the option that has been most visibly taken.

In New Zealand's schools system, reports were discontinued as an assessment method for NCEA Level 1 entirely from 2025, with authenticity concerns "including the rapid advancement in AI tools" given as the reason [1]. In the tertiary system, Victoria University of Wellington's law school moved two third-year courses to handwritten examinations for the June 2025 exam period, after its dean said advances in AI had made it difficult to be confident work was a student's own and that a technical solution for policing it was not yet available [2][3]. The university's provost described handwritten examinations as "the gold standard of ensuring integrity" [3].

That description is accurate, and the control is a reasonable response to an unsolved problem. What it does not do is remove the cost — it relocates it.

Prohibition answers the authenticity question by removing the conditions under which it arises. That is legitimate, and it is not a general answer.

3. Option two — detect. Tried, measured, and withdrawn.

The documented position

The Ministry of Education's guidance on generative AI in NCEA assessment states that AI detection software "should not be relied upon to ensure authenticity," that such tools are "susceptible to returning 'false positive' results," and that they are "therefore unsuitable to use as the sole means of ensuring authenticity" [5]. In the tertiary sector the position moved from caution to withdrawal: Massey University confirmed it no longer uses AI detection tools, switching off the detection feature following student reports of false accusations, after earlier moves by the University of Auckland and Victoria University of Wellington [6]. A Massey spokesperson described the tools as ineffective and noted that academic practice had been inconsistent — some staff treated a result as a guideline, others treated a flagged percentage as grounds for an accusation [6]. The University of Auckland's stated position is that it is extremely difficult to prove work is AI-generated absent obvious signs, and that students can use readily available tools to check whether their work will be flagged and adjust it until it is not [6].

3.1 — The failure is not a maturity problem. A detector infers machine authorship largely from perplexity: how statistically surprising the text is. Text that a language model would find unsurprising is scored as likely machine-written. That inference is structurally unable to distinguish machine-generated text from human text that happens to be plain, formulaic, or written within a narrow range of expression — which is not a defect that a better model removes, because the signal it relies on is genuinely shared by both.

3.2 — The bias is measured, and it falls in a specific direction. A peer-reviewed study of widely used GPT detectors found they consistently misclassified writing by non-native English speakers: more than half of non-native-authored TOEFL essays were incorrectly classified as AI-generated, while the same detectors achieved near-perfect accuracy on essays by US eighth-grade students [7]. The authors hypothesise the mechanism above — lower perplexity arising from restricted linguistic variability — and its senior author's stated recommendation was to be extremely careful about using such detectors, and to consider avoiding them [7][8].

The consequence is that the students most likely to be wrongly accused are the students least equipped to contest the accusation, and the accusation is generated by a tool whose output is a probability with no supporting account of what actually happened. This is the specific harm behind the phrase "false positive results" in the national guidance.

3.3 — And it is evadable by exactly the people it is aimed at. As the University of Auckland's position notes, tools exist that let a student test their work against detectors and revise until it passes [6]. A control that systematically misfires on the innocent and can be routinely evaded by the deliberate is not a control that becomes adequate with tuning.

4. The shared root: both options interrogate the artifact

Prohibition and detection look like opposites — one prevents, one polices — but they ask the same underlying question: what does this finished piece of work tell us about how it was made?

Asked of an artifact, that question has no reliable answer, and the reason is not technological. A finished document is the same document whether it was drafted over three weeks or generated in nine seconds. The information that distinguishes them is not in the artifact. It was in the process, and the process was not recorded.

DetectionProcess record
Object examinedThe finished workThe sequence that produced and checked it
Question askedDoes this look machine-generated?Which steps occurred, in what order, and when?
Answer typeA probabilityA record that is present or absent
When it is producedAfter the factAt the time of each step
Failure modeAccuses the innocent; misses the deliberateShows a step is unrecorded — which is a question, not a verdict
BurdenStudent must disprove a scoreStudent can demonstrate their process

That last row is the one that matters most for the equity problem in §3.2. Under detection, a student flagged by a biased tool is asked to prove a negative about their own mind. Under a process record, a student who did the work has something to point at. The direction of the burden reverses, and it reverses in favour of the person with the least power in the situation.

5. Option three — record the process

The existing verification practices are already process-based; what they lack is durability. National guidance describes the teacher's discharge of the responsibility to verify as knowing the student and their work, verbal questioning, inspecting a document's version history, and having the student sign a declaration or authenticity statement [5]. Every one of those is an examination of process rather than artifact. The gap is that none of them produces a record that survives being disputed: a declaration form is a self-attestation on paper, and version history is held by the same system the work was produced in.

What turns those practices into evidence is specified, neutrally and independently of any vendor, in the companion Completeness Specification. Applied here, four properties do the work:

PropertyWhat it changes
Written at the time of the stepA record created when the check happened cannot be assembled afterwards to fit a conclusion.
Independent of the system being recordedVersion history inside the authoring tool is written by the tool. An external record is not.
Signed and verifiable offlineA student, an appeals panel, or an external reviewer can check the record against a published key without asking the institution's permission.
Sealed at completionOnce submitted, the account of how the work was produced is fixed. Later additions are detectable rather than silent.
What this looks like in practice

A declared sequence for a piece of internally assessed work might be: task issued, plan submitted, draft submitted, feedback given, final submitted, authenticity check completed, result recorded. Each step writes a signed receipt at the moment it occurs, chained to the one before it, and the sequence is sealed when the result is recorded. Nothing in the receipt is the student's work — the receipt records that a step happened, when, and in what order, not what was written. If a question is later raised, the sequence is either intact and complete, or it shows plainly where a step is missing. A missing step is not an accusation; it is the start of an ordinary conversation, held with a record instead of two recollections.

slp8_receipt_v2, AgenticRail's production receipt schema, is offered as one conformant reference implementation. It is named here as the author's own. The specification is implementation-independent, and any vendor's record can be assessed against the same requirements.

6. What this does not do

Each of the following states something true about a process record, followed by its limit.

A process record establishes that a declared sequence occurred in order and was not altered afterwards. It does not establish that the thinking was the student's own. A person determined to submit work they did not produce can perform every step of a recorded process while sourcing the content elsewhere. This note claims no solution to that, and any product claiming one should be treated with suspicion. What the record removes is the far larger class of dispute in which nobody can establish what happened at all.

The record answers a question about process, which is what most integrity disputes actually turn on. The genuine cases of undetectable substitution are a narrower problem than the day-to-day one, which is that a check was required, may well have occurred, and cannot be shown to have occurred.

Tamper detection has a boundary. A sealed, chained record makes later alteration detectable. Detectable is not impossible: a party holding the signing keys could construct a consistent but false chain. What closes that gap is custody — a copy held by a party who is not the institution — not stronger cryptography.

Recording a process adds an obligation, and it should be honest about that. A sequence that nobody follows produces receipts of nothing. This approach suits assessment that already has declared, staged steps; it does not suit a single unstructured submission with no process to record, and it should not be retrofitted onto one by inventing ceremony.

This note takes no position on whether students should use generative AI. That is a curriculum and pedagogy question belonging to teachers and their institutions. The argument here holds whichever way it is answered: a permitted use and a prohibited use both need a record of what actually happened.

7. How to check every claim on this page

Every factual claim in §2 and §3 is cited to a named public source in §8. The claims about the instrument can be checked directly:

If a claim on this page does not match the live system, the live system is the authority and the page is wrong.

8. References

[1] RNZ, "Schools abandon take-home assignments after artificial intelligence used to cheat" — rnz.co.nz
[2] RNZ, "Victoria law students not allowed laptops in exams to prevent AI cheating" — rnz.co.nz
[3] RNZ, "Return to pen and paper for some university exams tough for digitally savvy students" — rnz.co.nz
[4] NZQA, Effective Assessment Practice Guide, February 2020 (proactively released as part of OIA response OC00458) — nzqa.govt.nz
[5] Ministry of Education / Te Poutāhū Curriculum Centre, GenAI in NCEA assessment: FAQs, March 2025 — education.govt.nz
[6] RNZ, "Universities give up using software to detect AI in students' work" — rnz.co.nz
[7] W. Liang, M. Yuksekgonul, Y. Mao, E. Wu, J. Zou, "GPT detectors are biased against non-native English writers," Patterns 4(7), 2023 — arxiv.org/abs/2304.02819
[8] Stanford University / ScienceDaily summary of [7], "GPT detectors can be biased against non-native English writers" — sciencedaily.com
[9] Companion note — When the Marker Is a Machine, on evidence for decisions in which an automated system did the marking.
[10] Companion brief — AI in New Zealand Education Assessment: The Missing Evidence Layer, on the statutory basis and the internal moderation cycle.
Document Fingerprint — SHA-256 — v1.0
c449bd0cb65c183e6c0c8a9c688e6f0067a71b4eb122137b3d907a7066e9f320
This hash is SHA-256 of the canonical string defined below. It is reproducible independently of this page using any SHA-256 implementation.

Canonical string (pipe-delimited, UTF-8, no trailing newline):
If AI Detectors Don't Work What Does|1.0|2026-07-26|TUARA KURI LIMITED|three responses to generative AI in assessment prohibit detect or record the process|prohibition and detection are both artifact-based they interrogate the finished work|NCEA Level 1 reports discontinued as an assessment method from 2025 citing authenticity and rapid advancement in AI tools|Victoria University of Wellington law school moved two third-year courses to handwritten exams June 2025 exam period|provost described handwritten exams as the gold standard of ensuring integrity|handwriting relocates the cost onto students who compose digitally and requires exemptions for students with disabilities needing keyboards|Ministry of Education March 2025 guidance AI detection software susceptible to false positive results unsuitable as the sole means of ensuring authenticity|Massey University no longer uses AI detection tools switched off following student reports of false accusations after University of Auckland and Victoria University of Wellington|University of Auckland position students can use tools to check whether work will be flagged and revise until it is not|Liang Yuksekgonul Mao Wu Zou GPT detectors are biased against non-native English writers Patterns 2023 arXiv 2304.02819|more than half of non-native-authored TOEFL essays misclassified as AI-generated near-perfect accuracy on US eighth-grade essays|detectors infer machine authorship from low perplexity which is shared by plain human writing|under detection the student must disprove a score under a process record the student can demonstrate their process|a process record does not establish that the thinking was the student's own|tamper detection is detectable not impossible custody closes the gap not cryptography|completeness R1-R8|report.agenticrail.nz

Published: 2026-07-26  |  Version: 1.0  |  Entity: TUARA KURI LIMITED

Sourcing note: every factual claim in §2 and §3 is cited to a named public source in §8. This note asserts no failure, breach, or bad faith by any institution, teacher, or student, and names no individuals other than the authors of the cited academic study.