How Teachers Handle a High AI Likelihood Score Fairly
Picture this: you upload a Year 11 essay to GradeOrbit, the AI likelihood score comes back at 87%, and the rest of your free period is suddenly not free at all. What you do in the next 24 hours matters more than the number on the screen. A high score is the start of a process, not the end of one — and how teachers handle that process determines whether a student leaves the conversation with their integrity intact, whether your professional judgment holds up under scrutiny, and whether your school's academic integrity policy means anything in practice.
This post walks through the fair, evidence-based steps a UK secondary teacher should take when a detection tool flags student work. It is written for the moment after the score lands, not before — because the question worth answering is not "what signals reveal AI use?" but "what do I do now that one has fired?"
What a Likelihood Score Means (and What It Doesn't)
Every responsible AI detection tool, GradeOrbit included, returns a probabilistic score between 0% and 100%. The number expresses how closely the linguistic patterns in the text — sentence rhythm, vocabulary distribution, predictability of word choices — resemble patterns commonly produced by large language models. It is not a confession. It is not a verdict. It is one piece of statistical evidence that has to be weighed against everything else you know about the student and the work.
An 87% score does not mean eighty-seven out of a hundred markers would agree the work is AI-generated. It does not account for whether the student is an unusually polished writer, whether they spent three weekends redrafting with their parent's help, or whether they have a habit of mimicking the register of whichever textbook they last read. Detection scores are calibrated against a population of writing, not against this student. Treating the score as proof is the single fastest way to lose a parental complaint and damage a working relationship with a teenager.
The 24 Hours Before You Speak to the Student
Resist the pull to confront the student at the end of the next lesson. The hours between the score landing and the conversation are for gathering evidence, not for rehearsing accusations. Pull up the student's previous work in the same subject — three or four pieces is usually enough — and read them side by side with the flagged piece. Ask yourself, honestly, whether the voice, vocabulary range, and structural choices feel continuous. A student who has produced grade-5 work all year and suddenly writes at grade-8 polish in one piece is more interesting than the score alone.
Check the practical context too. Was this written in class, under supervision, or as homework? Did you set any drafts or planning stages you can compare against? Is there a Google Docs revision history you can request through your school's normal channels? None of this is about building a case for prosecution — it is about giving yourself enough context that the conversation can be open rather than adversarial.
Document what you have found and what you have not. A short paragraph in your markbook — "92% likelihood; voice inconsistent with previous three pieces; no in-class draft available" — is the kind of note that protects you if the conversation later escalates to the head of department or senior leadership.
Running a Fair Conversation, Not an Interrogation
When you do speak to the student, the goal is to understand, not to extract a confession. Open with the work itself, not the score. Ask them to walk you through how they planned it, what sources they used, which paragraph they found hardest, what they would change if they could redraft it. A student who genuinely wrote the piece can usually talk about it fluently — their thinking is in their head and comes out in answers to follow-up questions. A student who generated it is more likely to repeat surface-level points and struggle when pressed on a specific choice.
If the conversation moves to the detection score, frame it as information you have, not as proof you are wielding. "A piece of automated analysis flagged this as similar to AI-generated writing. I wanted to talk to you about it because the polish and structure look different from your usual work" lands very differently from "the AI detector says you cheated." The first invites a response. The second invites a defensive shutdown — and shuts down any chance you had of learning what actually happened.
Listen for the boring middle. Most cases are not "student wrote it entirely themselves" or "student copied it entirely from a chatbot." They are "student used a chatbot to plan, then wrote it in their own words" or "student wrote a first draft, then asked the chatbot to improve it." Your school's academic integrity policy probably has different responses to those middle cases than it does to a copy-paste — but only if you actually find out which one you're dealing with.
Aligning Your Response with School Policy
Your school has, or should have, an academic integrity policy that covers AI use. Read it before the conversation, not after. Some schools treat any uncredited AI use as misconduct; others permit AI for planning and brainstorming as long as the final writing is the student's own; others have different rules for coursework versus homework versus classwork. The detection score does not change which policy applies — the type of work and your school's stance on it do.
Document the conversation outcome in line with what the policy requires. If the student admits to substantial AI use on a piece of NEA coursework, that probably needs to go to your head of department and may need to go to the exam board. If they admit to using a chatbot to summarise a textbook before writing their own essay on homework, that may be a teaching moment rather than a sanction. The point is that the response is calibrated to the policy and the type of work, not to the size of the likelihood score.
For coursework heading to a moderator, log the detection result in the moderation pack regardless. A score on its own is not evidence of misconduct, but it is the kind of context a moderator may want to see when reviewing borderline pieces. The moderation cycle is where detection scores earn their keep — as one input among many, not as a verdict.
When the Score Was Wrong — Repairing Trust
Sometimes the conversation ends with you genuinely convinced the work was the student's own. The score was wrong. This happens. EAL students, students who have been reading academic prose all summer, students who write in a deliberately formal register because they think that is what good writing looks like — all of these produce text that detection models flag as AI-like.
If you reach that conclusion, say so explicitly to the student. "I called you in because a tool flagged your essay. Having talked to you, I'm confident the work is yours. I'm sorry the process felt accusatory — that wasn't the intention." The repair conversation is short, but skipping it leaves a student carrying the sense that their teacher does not trust them, which is corrosive in ways that take a full term to undo.
If the student or their family complains, the contemporaneous notes you took during the evidence-gathering phase are what protect you. "I followed the fair process: I gathered comparison evidence, I held an open conversation, I reached a conclusion based on the conversation, and I communicated that conclusion" is a defensible position. "The score said 87% so I marked it down" is not.
How GradeOrbit Supports a Fair Process
GradeOrbit's AI detection tool returns a likelihood score between 0% and 100%, and it offers two model choices: a 1-credit model for quick triage on lower-stakes work, and a 3-credit model for higher-confidence analysis on coursework or pieces heading for moderation. The score is the start of the fair process this post describes — not a replacement for it.
We do not save student work. Uploads are processed in memory and discarded after the score is returned, so there is no database of flagged students sitting on a server somewhere. Teachers redact personal information before uploading (the canvas tool burns black boxes into the image), and students are anonymous to the model — referred to only as "Student 1", "Student 2", and so on. The audit trail lives in your markbook and your moderation pack, where it belongs.
For teachers who want to read more about how likelihood scores behave in real cases, our guide on the student conversation goes deeper on the language to use, and our piece on false positives covers the patterns that most often trip up the model.
Try GradeOrbit for Fair AI Detection
If you are looking for an AI detection tool that gives you a clear likelihood score and trusts you to run a fair process around it, GradeOrbit is built for that. New accounts get a small allocation of free credits to try the detection and marking workflows on real student work.
Visit our homepage to learn more, sign up, and run your first detection. The tool exists to support your professional judgment, not to replace it — and the fair process described above is how you make sure it stays that way.