How to Interpret AI Detection Scores in Student Work
An AI detection likelihood score appears on your screen — 78%. Now what? For many UK secondary teachers, that number sits in a difficult space: too high to ignore, but nowhere near a clear enough signal to act on. The anxiety this creates is real, but it is largely the result of a misunderstanding about what these scores actually measure. Once you understand the mechanics, the number stops being a verdict and starts being a useful data point.
This guide explains how to interpret AI detection likelihood scores accurately, how to choose the right detection model for the task at hand, and how to build a fair, evidence-based response that protects both academic integrity and student welfare.
What an AI Detection Likelihood Score Actually Means
AI detection tools — including the one built into GradeOrbit — do not tell you whether a student used ChatGPT, Claude, or any other AI tool. They tell you how closely a piece of text resembles patterns statistically associated with AI-generated writing. That is a meaningful distinction, and it changes everything about how you should respond to the result.
A score of 0% means the text closely resembles patterns typical of human writing. A score of 100% means the text strongly matches the linguistic fingerprint of known AI models. But the scale is not binary: a score of 60% does not mean the student is 60% guilty of anything. It means the text sits in a zone of genuine statistical ambiguity, where both human and AI authorship are plausible explanations.
The middle band — roughly 30% to 70% — deserves the most caution. Scores at the extremes (below 20% or above 85%) carry stronger signal, but even then, your professional knowledge of the student remains the most important interpretive tool you have.
Low, Medium, and High Confidence — What Each Level Means for You
Alongside the likelihood score, GradeOrbit returns a confidence label: Low, Medium, or High. This label tells you how certain the model is about its own assessment — and it is just as important as the score itself.
A Low confidence rating typically appears when the submitted text is short. Detection models work by analysing linguistic patterns across sentences and paragraphs. When there is not enough text to establish a reliable pattern — a single paragraph, for instance — the model flags its own uncertainty. A score of 80% with Low confidence is far less meaningful than a score of 80% with High confidence. Treat low-confidence results as weak signals that need more supporting evidence before any action is warranted.
A Medium confidence rating suggests the model has enough text to work with, but some ambiguity remains — perhaps the writing switches register partway through, or the stylistic signals are mixed. This is a prompt to look more carefully at the text itself rather than rely solely on the number.
A High confidence rating means the model has identified consistent, strong patterns across the full submission. This is still not proof of AI use, but it significantly raises the weight of the score as a signal worth investigating further.
Where AI Detection Lives in GradeOrbit Now
One important correction, because this changed in 2026: GradeOrbit no longer offers a standalone AI detection tool on individual teacher accounts. Detection now exists in two places, and both are deliberately narrower than a general-purpose checker.
- School accounts. Detection runs from a pupil's profile, a class roster, or a completed marking run. Because a school stores each pupil's work under a Data Processing Agreement, detection can do something a one-off checker cannot: compare new work against that same pupil's own earlier writing, and report how far it has moved. It costs one credit per paper.
- Student accounts. A learner who has already marked at least one piece can opt in, for one extra credit, to check whether a new piece reads like their own earlier work.
There is no longer a "faster versus smarter" model choice, and no three-credit deep analysis — that split was retired when the underlying models were unified. Everything below about reading a score still applies; only the product surface changed.
Baseline Comparison Beats a Generic Score
This is the most useful shift in the last two years, and it changes how you should read a number. A generic detector asks "does this text look like AI writing in general?" A baseline-aware check asks "does this look like this pupil's writing?" The second question is far more answerable, because it has a control group: the pupil's own marked work.
It also fails differently, and you need to know how. A pupil who has genuinely improved — because your teaching worked, because they redrafted, because they finally understood paragraphing — will register as changed. Change is not deception. A baseline score tells you something moved; it never tells you why.
Reading Borderline and Mid-Range Scores
The middle of the range is where teachers get stuck, so treat it as a category rather than a failure of the tool. A mid-range score means the text sits in a zone where both explanations are statistically plausible — and that is genuine information, not an error. The correct response to a borderline score is almost always to do nothing formal: note it, keep the work, and see whether a pattern emerges across pieces.
A single borderline score on a single piece is the weakest possible evidence. Three borderline scores across a term, on work that also does not match what you have seen the pupil produce in class, is a reason for a conversation. The pattern carries the weight, never the individual number.
Specific Scenarios Teachers Ask About
Coursework and EPQ drafts
Long-form independent work is where scores are least reliable and the stakes are highest. Drafts are redrafted repeatedly, often with teacher feedback in between, so the writing legitimately becomes cleaner and more formal with each pass — the exact direction that raises a score. Check the first draft against later ones and treat the trajectory as the evidence. A pupil whose first draft already reads like polished prose is more interesting than one whose fourth draft does.
Mock marking
Mocks are usually handwritten and done under supervision, which makes them the strongest baseline material you have. If you are going to run detection at all, use mock scripts to establish what a pupil's unassisted writing actually looks like, then compare coursework against that — not the other way round.
Group and collaborative work
Detection is close to meaningless on genuinely collaborative work, and you should say so in your policy. Multiple authors produce exactly the stylistic inconsistency that detectors read as a signal, and a baseline comparison has no single pupil to compare against. Do not run detection on group submissions and expect an interpretable answer.
Moderation and standardisation
Detection has no place in the moderation cycle itself. Moderation is about whether marks are applied consistently against the criteria; authorship is a separate question, handled through your malpractice policy. Mixing the two produces meetings where a mark is disputed on the basis of a likelihood score, which is unfair to the pupil and unhelpful to the department. Settle authorship concerns before work enters moderation, not during it.
Understanding the Analysis You Get
GradeOrbit offers two AI detection models, and choosing between them well makes a practical difference to both the quality of your results and your credit usage.
A detection run returns the score, a confidence label, and the specific linguistic signals behind it — sentence-length consistency, vocabulary register, argumentation structure. On a school account it also reports how far the work sits from that pupil’s own baseline, which is the part worth reading first.
For high-stakes work — A-Level coursework, NEA submissions, GCSE controlled assessment — the explanation matters more than the number, because a formal conversation or a referral has to rest on something you can articulate. "The score was 78%" is not an evidence base. "This piece is markedly more formal than four earlier pieces, and the pupil could not talk me through their argument" is.
There is no configuration to get right — one detection run, one credit, the same engine every time. What varies is how much weight you give the result, and that is a judgement about stakes rather than a setting: the higher the consequence of being wrong in either direction, the more corroboration you should require before acting.
How to Respond to a High Score Without Jumping to Conclusions
A likelihood score above 70% with High confidence is a genuine signal that merits closer attention. But the correct response is not to confront the student. It is to look more carefully before doing anything at all.
Start by comparing the flagged piece against work you know is authentic — ideally something produced under timed conditions in class, where you were present. If a student's in-class writing is a Year 8 level and the flagged homework reads like a polished university essay, that contextual gap matters far more than the score alone. Conversely, if the flagged piece is entirely consistent with everything else you have seen from that student, the score should prompt curiosity rather than concern.
Next, read the text itself carefully. AI-generated writing — whether from ChatGPT, Claude, or similar tools — tends to show particular patterns: an evenness of sentence length, an absence of genuine personal voice, generic transitions, and a kind of structural competence that hits every required point without any authentic personality beneath it. These qualitative signals are things an experienced teacher can recognise, and they should either add to or reduce your concern about what the score is telling you.
If both the score and your qualitative reading raise concern, the most productive next step is a short, non-accusatory conversation: ask the student to walk you through their argument, explain where they found a particular piece of evidence, or write a short paragraph on the same topic in front of you. A student who wrote their own work will be able to engage fluently. A student who submitted AI-generated content will often struggle to explain ideas they did not actually form.
Why False Positives Happen — and How to Spot Them
False positives — cases where genuinely human-written work scores highly — are well-documented and worth taking seriously. Several student profiles are at elevated risk.
Highly proficient writers, particularly those who have been coached intensively or who naturally produce structured, fluent prose, can trigger elevated scores. Academic writing at its best shares many qualities with AI output: clear topic sentences, logical sequencing, minimal redundancy. A Year 13 student producing A-Level quality work may simply be writing very well.
Students who speak English as an additional language are another category that requires particular care. Students who draft in their first language and then translate — whether manually or using a translation tool — can produce formal, over-structured English that detection models flag. Treating a high score on an EAL student's work as evidence of AI use without careful further investigation could be both educationally harmful and unfair.
Students who have revised their work extensively through multiple drafts, responding to teacher feedback and improving the structure and fluency of their writing, may also produce text that looks statistically cleaner than a first draft would. Good pedagogy produces better writing — and better writing can look more like AI output to a probabilistic model. This is not a failure of teaching; it is a limitation of the tool that you need to account for.
How GradeOrbit Supports Evidence-Based Detection
GradeOrbit's AI detection is designed around the classroom context rather than as a general-purpose checker. On a school account you run it from a pupil's profile, a whole class, or a completed marking run — reusing work already uploaded for marking, with no re-scanning. Results include the score, the confidence label, the detected linguistic signals, and the comparison against that pupil's own earlier writing.
How the work is handled depends on the account, and the two models are deliberate opposites. On a school account, work is retained under a signed Data Processing Agreement — that persistence is what makes baseline comparison possible at all — within a retention window the school controls, and wiped on offboarding. On an individual or team account, uploaded work is never stored, and teachers redact names and personal details before anything is processed. Work is never used to train any AI model on any tier.
For a broader grounding in how detection works and how to build a fair school-wide policy, the guide on how teachers check for AI covers the institutional dimension in more depth.
Use Detection as One Input, Not a Verdict
AI detection likelihood scores are one piece of evidence among many. Used alongside your professional knowledge of the student, a careful reading of the text itself, and a direct conversation where the evidence warrants it, they can play a meaningful role in maintaining academic integrity — without putting students at unfair risk.
If a score is ever the reason a pupil is accused of something, the process has already gone wrong. Used properly it is a prompt to look more closely at work you already had questions about. See how detection works on a GradeOrbit school account, and read our companion guide on false positives and how to respond fairly.