How Teachers Interpret Mid-Range AI Likelihood Scores
The 90% score is easy. The 5% score is easy. The 47% score is the one that keeps teachers staring at the screen at half past nine in the evening, trying to work out what to do next. A mid-range AI likelihood score is not a verdict — it is a prompt to think more carefully, and that thinking is where the professional judgement actually happens.
This post is for the teacher who has just opened a piece of Year 10 coursework, fed it through an AI detection tool, and seen a number that sits squarely in the grey zone. The aim here is not to give you a magic threshold. It is to help you turn that uncomfortable number into a fair, defensible next step.
What a Mid-Range Score Actually Means
AI detection is probabilistic. It is not a fingerprint, a lie detector, or a confession. When GradeOrbit returns a likelihood score of, say, 45%, what the model is telling you is that the linguistic signals in this piece of writing sit between what it typically sees in fully human-written work and what it typically sees in fully AI-generated work. There are several reasons that might happen, and most of them are not "the student cheated".
A student who has used a spellchecker heavily, written in a deliberately formal register because they were told to, drafted an essay from detailed notes, or simply writes in a clean, structured way may produce text that lands in the middle of the range. Equally, a student who used AI to draft and then rewrote substantial parts in their own voice may also land there. The score alone cannot tell you which is which — and that is the point.
The job of a mid-range score is to flag the piece for a closer human read, not to make the decision for you.
Reading the Score Alongside the Rest of the Evidence
Before you do anything else, put the score down and read the work. Read it as if no number were attached. Then ask the questions you would ask of any piece of work that surprised you:
- Does the vocabulary match what this student has produced in class for you?
- Does the argument structure match their usual planning style?
- Are there topic-specific quirks — the textbook example they always reach for, the misconception they have shown twice this term — present or absent?
- Did they submit a planning document, redraft, or earlier version? Does the final piece show the journey?
- Have you seen them write at this length and complexity in a controlled condition recently?
Each of these is a piece of evidence. The likelihood score is one more piece — not the heaviest one. A mid-range score with three other signals pointing to "this does not feel like their voice" is a different situation from a mid-range score on a piece that otherwise reads like the student you teach.
The Fair Conversation — Curiosity, Not Accusation
If the combined picture leaves you uncertain, the next step is a conversation, not a sanction. The single biggest mistake teachers report is opening with the score. "This came back as 60% AI" puts the student straight into defence mode, and you lose the chance to learn anything useful.
A better opening is curiosity. Ask them to walk you through how they approached the piece. Where did they start? What did they find hardest? Which paragraph took the longest? A student who genuinely wrote the work can usually answer those questions in detail. A student who didn't will often hesitate, contradict themselves, or default to vague answers.
This is not entrapment — it is exactly the conversation a head of department would have at moderation. Frame it as routine, because that is what it should be. Our companion post how to talk to students about AI detection results walks through more of the language and tone for these conversations.
When to Re-Run on the 3-Credit Model
GradeOrbit offers two AI detection models: a 1-credit model for everyday triage and a 3-credit model for closer scrutiny. The two-tier setup exists exactly for situations like this. If you have a mid-range score and the rest of the evidence is mixed, re-running the same piece on the 3-credit model is a reasonable next step before you decide whether to escalate.
The 3-credit model takes a longer, more careful look. It is not infallible — nothing is — but it gives you a second probabilistic read that you can place alongside the first. If both models land in the grey zone and the work otherwise reads like the student, that is meaningful. If the second model swings high or low, that is meaningful too. Either way, you have more information to take into the conversation or the moderation note.
Logging the Decision for Moderation and Policy Alignment
Whatever you decide, write it down at the time. A short, dated note in the student's record — score from the first run, score from the second run if you did one, the evidence you weighed, the conversation you had, the conclusion — is worth its weight in gold three weeks later when a parent asks, or six months later when a moderator queries a grade.
This is also where your school's academic integrity policy should anchor your decision. A consistent department-wide threshold for escalation, written down and agreed in advance, takes a huge amount of pressure off individual teachers in the moment. If your school has not yet written one, our post on how to write a school AI academic integrity policy is a useful starting point.
A mid-range score is not a problem to be solved with a single click. It is the start of a professional process. The score gives you a reason to look closer; your judgement, the evidence around the piece, and a fair conversation with the student do the rest.
Try GradeOrbit's AI Detection Free
GradeOrbit's AI detection gives you a likelihood score from 0 to 100% with a 1-credit triage model and a 3-credit closer-read model, so you can act on the easy scores quickly and give the grey-zone ones the attention they deserve. Every new account starts with free credits — no card, no commitment. Get started with GradeOrbit and see how the two-tier model fits into your marking week.