Skip to main content
Back to Blog

AI Detection False Positives: A Teacher's Guide

Breanna Mitchell·Content Writer
7 min read

A student submits a piece of coursework that returns a high AI detection score. Your instinct might be to treat that score as proof — but it is not. AI writing detection false positives are real, they happen more often than many teachers realise, and acting on a score without proper context can damage a student's reputation and your school's trust in you as a fair professional.

This guide is for UK secondary school teachers who want to use AI detection tools responsibly. We will explain how likelihood scores actually work, which student profiles are statistically more likely to trigger false positives, how to use a deeper analysis before drawing conclusions, and how to build a defensible, evidence-based conversation with the student.

How AI Detection Likelihood Scores Work

AI detection tools do not read a student's mind. They analyse patterns in the submitted text — sentence rhythm, vocabulary distribution, structural regularity, and statistical predictability — and compare them against models trained on both human-written and AI-generated text. The output is a probability score, not a verdict.

GradeOrbit's AI detection tool returns a likelihood score between 0% and 100%. A score of 80% does not mean the work is 80% AI-generated. It means that, based on the statistical patterns in the text, the model estimates an 80% probability that the writing shows characteristics associated with AI authorship. Probability is not certainty, and certainty is exactly what you need before taking any formal action.

GradeOrbit no longer splits detection into a cheap scan and a deep analysis — that two-tier model was retired when the underlying engines were unified. There is one run, one credit, and on a school account it is baseline-aware: the comparison is against that pupil’s own earlier marked work, not against writing in general. That single change removes a large share of the false positives a generic detector produces, because "this does not look like AI in general" is a much harder question than "this does not look like Amara".

What Causes False Positives?

Several categories of student are statistically more likely to receive a high detection score even when their work is entirely their own. Understanding these profiles is the first step towards applying professional judgement alongside the tool's output.

English as an Additional Language (EAL) students are at particular risk. EAL learners are often taught to write in formal, structured patterns early in their language acquisition. Their sentences tend to be shorter, their vocabulary more predictable, and their phrasing more formulaic — all characteristics that AI detection models flag as statistically likely to be machine-generated. A student who writes clear, rule-adherent prose because they have worked hard to master English grammar may score higher than a native speaker who writes in a more idiosyncratic, human style.

Students in highly formulaic genres also face this risk. Science report writing, structured history essays following the PEEL model, and religious studies analytical paragraphs all tend toward predictable patterns. When students have been trained effectively to follow a genre convention, their writing may superficially resemble AI output because both are adhering to the same structural rules.

Students who have significantly improved can also trigger suspicion. If a student has worked with a private tutor, attended a revision workshop, or genuinely put in extra effort over a holiday period, their writing may look uncharacteristically polished. A high detection score in this case reflects an improvement in quality, not the use of an AI tool.

How to Interpret a High Score Without Jumping to Conclusions

The most important principle when reviewing an AI detection score is that it must be considered alongside everything else you know about the student. A score does not exist in a vacuum. Your professional knowledge of the student is irreplaceable context that no algorithm can replicate.

Start by comparing the flagged piece against the student's previous written work. If a student has produced three pieces this term with a consistent voice, structure, and level of sophistication, and the fourth suddenly reads very differently — that pattern is meaningful. Conversely, if the flagged work is consistent with their established voice, the score warrants scepticism.

Also consider the circumstances under which the work was produced. Was it completed in a supervised lesson? Was it a timed piece? Did the student submit a draft that you reviewed mid-process? The more controlled the conditions, the less plausible AI involvement becomes, regardless of what the detector returns.

No professional body in UK education — including the Joint Council for Qualifications (JCQ) — treats an AI detection score alone as sufficient evidence of malpractice. It is one piece of evidence among many, and the burden of proof lies firmly with the school to demonstrate, on the balance of probabilities, that AI was used inappropriately.

How to Investigate a High Score Fairly

If a score surprises you — particularly for a pupil whose previous work you have no concerns about — the next step is corroboration, not escalation. Work through these in order, and stop as soon as you have an innocent explanation:

  1. Read the work yourself, properly. Before looking at the number again. Does the argument hold together? Is there a voice in it? Detectors cannot tell you whether a piece is good, and a piece that is fluent but says nothing is a different concern from a piece that was not written by the pupil.
  2. Compare against their own earlier work. Not against a class average. Look for a jump in register, sentence structure, or vocabulary that arrives without an intermediate stage.
  3. Check what changed around them. New tutor, a redraft cycle, a unit they are genuinely interested in, a diagnosis and new support in place — all of these produce real, explainable improvement.
  4. Look at the conditions. Was it produced in class, under supervision, handwritten? Then authorship is largely settled regardless of the score.
  5. Only then, talk to the pupil — as an investigation, not an accusation. See below.

The deep analysis examines the text at a more detailed level, producing a more reliable likelihood score. In many cases, a high first-pass score will drop significantly under deeper scrutiny. If the score remains high after a deep analysis, and if it is also inconsistent with the student's prior writing, you then have a stronger basis for a quiet, exploratory conversation — not an accusation.

Using the deep analysis also demonstrates to senior leadership, parents, and — if it ever came to it — a formal malpractice panel that you applied due diligence before drawing any conclusions. That paper trail matters.

Building a Defensible Conversation With the Student

If you have reviewed the score, compared it with prior work, and you still have genuine concerns, the next step is a private, non-accusatory conversation with the student. The goal of this conversation is to gather information, not to deliver a verdict.

Ask the student to talk you through how they approached the piece. Where did they start? What did they find difficult? Can they explain a specific phrase or structural choice that caught your attention? A student who wrote the work themselves will generally be able to describe their process, even if imperfectly. A student who submitted AI-generated text and did not engage meaningfully with it is much less likely to be able to do so.

For more detailed advice on structuring this kind of conversation, see our post on how to interpret AI detection scores. The key principle throughout is that the conversation is investigative, not punitive. Your role at this stage is to establish the truth, not to assign blame.

Document everything: the original score, the deep analysis result if you ran one, the comparison with prior work, and a brief note on the conversation. If the matter does escalate, that documentation protects both you and the student.

Quotes, Citations, and Part-AI Work

Two edge cases account for a large share of the confusing results teachers bring us.

Heavily quoted work. Embedded quotations are, by definition, not the pupil’s prose. A well-evidenced English Literature essay can be a quarter quoted material, and that published text raises the score without telling you anything about the pupil. Judge the analysis between the quotations. A pupil who quotes well and analyses thinly is a teaching problem, not an integrity one.

Part-AI work. This is the common real case, and it is the one detectors handle worst: a pupil writes the piece and then runs it through a tool to “tidy it up”. The result is genuinely mixed authorship, which produces a mid-range score and no clean answer. Your policy needs a position on it before you meet the case, because “did you use AI” gets a yes that means something different from the one you were asking about. Most schools land on: assistance with proofreading is acceptable and must be declared; generating argument or content is not.

When a Pupil or Parent Challenges the Score

Assume this will happen, and that the challenge will be reasonable. Three things make it survivable.

Never let the score be the finding. If your written record says "flagged at 82%" and nothing else, you have no position to defend. If it says what you read, what you compared, what the pupil said, and what you concluded, the score is just one line in a professional judgement.

Be able to explain the limitation out loud. A parent who is told "the software says so" will escalate. A parent who is told "the software indicated the writing had changed markedly from her earlier work, so I asked her to talk me through it, and she could not explain her own argument" is being shown reasoning.

Have the policy written before the case. Who decides, what evidence is required, what the pupil is entitled to know, and what happens on appeal. A department improvising this during a live dispute always improvises badly.

Used responsibly, AI detection supports academic integrity without prejudging pupils or creating an atmosphere of suspicion. Used as proof, it will eventually be wrong about a real pupil in a way that is very hard to undo. See how baseline-aware detection works on a GradeOrbit school account.

More on this topic

3 May 202613 min read

How to Interpret AI Detection Scores in Student Work

A high AI detection likelihood score does not automatically mean a student cheated. This guide helps UK teachers read scores accurately, choose the right model, and respond with professional judgement.

Read more
22 July 20267 min read

How Do Teachers Check for AI? What Schools Really Use in 2026

Students and parents often ask how teachers actually check whether work was written by AI. The honest answer: a mix of professional knowledge, drafting evidence, conversation, and detection tools used as an indicator — never as proof. Here is how it works in UK schools.

Read more
22 July 20266 min read

Mark GCSE Maths Non-Calculator Papers Faster With AI

Non-calculator maths papers are marked on method: the working, the intermediate steps, and the marks a student earns even when the final answer is wrong. Here is how AI marking speeds that up without losing exam-board rigour.

Read more

Ready to save time on marking?

Join UK teachers using AI to provide better feedback in less time.

Get Started Free