Skip to main content
Back to Blog

GradeOrbit vs Turnitin: How AI Detection Scoring Actually Works

Breanna Mitchell·Content Writer
7 min read

If you teach in a UK secondary school, the question of how to handle AI-generated student work has probably landed on your desk in the past year. Two names tend to come up: Turnitin, which has bolted AI detection onto its long-standing plagiarism platform, and GradeOrbit, which built AI detection into a marking-first product from the start.

This post explains how the underlying scoring works in plain English, why the two tools give different numbers on the same piece of work, and how to use a likelihood score as one piece of evidence rather than a verdict. The aim is not to crown a winner — it is to help you make an informed call when ChatGPT or Claude shows up uninvited in a Year 11 essay.

What an AI Detection Score Actually Is

Every reputable AI detection tool — including GradeOrbit and Turnitin — produces a probabilistic output. The number you see is a model's estimate of how strongly the text resembles patterns associated with AI generation. It is not a forensic test result. It is closer to a weather forecast: useful, directional, sometimes wrong.

That distinction matters because the way the score is presented shapes how teachers and senior leaders react to it. A bare percentage with no context invites a binary interpretation. A score paired with reasoning, detected signals, and confidence labelling invites professional judgement. The difference between those two experiences is what separates a useful classroom tool from a disciplinary risk.

How Turnitin Approaches AI Detection

Turnitin's AI detection sits inside its broader similarity report. It compares submitted writing against statistical patterns learned from large samples of human and AI-generated text, then returns a percentage describing how much of the submission appears AI-generated. Turnitin has been transparent that its model can produce false positives, particularly on writing from English-as-an-additional-language students and on short passages.

The product makes sense for institutions already using Turnitin for plagiarism. It is also tied to a wider workflow — submission portals, originality scores, similarity matches against archived student work. That archival side is part of why some schools are cautious about it: student work is processed and, in some cases, retained by the platform.

How GradeOrbit's AI Detection Works

GradeOrbit takes a different path. The detection tool sits inside a marking-first product, so it is built for the same moment teachers reach for it in real life — a piece of work has just arrived and you want a second opinion before you mark it.

You can submit student writing as pasted text, an image, or a scanned document. The tool then returns four things together:

  • A likelihood score from 0 to 100% (0 = almost certainly human, 100 = almost certainly AI).
  • A confidence label — Low, Medium, or High — describing how sure the model is.
  • A list of detected signals — specific linguistic or structural patterns that pushed the score up or down.
  • A short reasoning paragraph in plain English.

GradeOrbit offers two models: a faster 1-credit check for quick screening, and a more thorough 3-credit analysis for borderline pieces or formal investigations. Your default model is remembered, so you are not reconfiguring it every lesson.

Crucially, GradeOrbit does not store student work. Submitted text and images are sent to the AI model for analysis and discarded after the result is returned. That choice is deliberate: it protects pupil privacy and keeps schools out of complicated data-retention questions.

Why the Same Essay Can Score Differently

If you run the same paragraph through Turnitin and GradeOrbit, the numbers will almost never match exactly. That is not a bug — it is a property of how these systems work.

Different detectors are trained on different samples of human and AI writing. They weight different signals — sentence rhythm, vocabulary choice, paragraph structure, predictability of the next word — in different combinations. A 64% score in one tool and a 28% in another does not mean one is right and one is wrong. It means the two models found different amounts of evidence in the same text.

This is also why no single number should ever be treated as proof of misconduct. Teachers' colleges, exam boards, and the JCQ have all been clear: AI detection results are corroborating evidence, never a verdict on their own.

Using a Likelihood Score Well

The score is most useful when you use it to structure your professional judgement rather than replace it. A practical workflow looks like this:

Anchor on the Student's Baseline

Your own knowledge of the pupil is the strongest signal you have. If the writing feels stylistically alien to anything they have produced in class, that is meaningful regardless of what any tool reports. A high likelihood score that aligns with your gut tells you to dig in. A high score that does not match your knowledge of the student tells you to be careful and look closer before reacting.

Read the Reasoning, Not Just the Number

GradeOrbit's reasoning paragraph and detected-signals list exist precisely so a 78% score is not a black box. If the signals say "low burstiness, uniform sentence length, generic transitions," you can hold that against the actual essay in front of you. Sometimes the signals are convincing. Sometimes they describe a Year 10 who simply writes in clean, even sentences.

Have the Conversation Before the Sanction

If detection scores and your professional read both point the same way, the next step is almost always a short conversation with the student. Ask them to talk through their argument, defend a choice of evidence, or write a short paragraph on the same topic on paper. We covered this in more depth in our guide to AI detection for teachers, and the principle holds across every detection product on the market: the score opens the door, the conversation closes the case.

Where Each Tool Fits

Turnitin makes sense for schools already deep in its ecosystem, particularly where coursework is centrally submitted and plagiarism similarity is a core part of the workflow. GradeOrbit makes sense for teachers who want AI detection as part of the same place they mark — and who want a product that does not retain student work.

Other tools exist too. None of them, used alone, will give you certainty. All of them are useful when treated as a structured second opinion rather than a judge.

The Honest Limit of All AI Detection

Detection technology is in a race with generative models that are getting better at sounding human. A student who runs AI output through a paraphraser, edits it by hand, and adds a personal anecdote will drop their likelihood score on any tool currently available. That is the reality every detection product — including GradeOrbit's — has to be honest about.

What this means in practice is that detection tools work best as part of a wider academic-integrity culture: clear policies, in-class writing samples, regular drafting checkpoints, and conversations with students about what AI use is and is not acceptable. The tool gives you signal. The school's culture turns signal into integrity.

Try GradeOrbit's AI Detection

If you want to see how the likelihood score, confidence label, detected signals, and reasoning paragraph work together on real student writing, GradeOrbit's AI Detection is available inside the dashboard with no setup. You can run the faster 1-credit check or the deeper 3-credit analysis, and student work is never stored. Head to the GradeOrbit homepage to create a free account and try it on a single piece of work before deciding whether it fits your marking routine.

More on this topic

17 June 20267 min read

How Teachers Detect AI in Summer Bridging Work

Bridging work set over the summer is the easiest place for AI use to slip through unnoticed. A practical guide for UK teachers on detecting AI fairly in transition tasks for incoming Year 12 students.

Read more
25 May 20267 min read

What Teachers Do When a Student Denies Using AI

Likelihood score came back high but the student says no. A UK teacher process for fair investigation, evidence beyond the score, and policy-aligned next steps.

Read more

Ready to save time on marking?

Join UK teachers using AI to provide better feedback in less time.

Get Started Free