A score measures resemblance
Detectors do not watch anyone write. They compare the finished text with patterns they associate with machine output, as explained in how AI text detectors work. A high score means the text looks like what the detector learned to call machine-written. Plenty of human writing looks like that: formulaic assignments, careful non-native English, heavily templated reports. A score of 82 is therefore a statement about the text, and only indirectly about the writer.
The same score means different things in different rooms
Suppose a detector flags 2 percent of honest writing and catches 90 percent of machine writing. In a class where half the essays were machine-written, almost every flag is right. In a class where 5 percent were, roughly one flag in three points at an honest student. The detector has not changed; the population has. This is the base-rate effect worked through in false positives and base rates, and it is why a score cannot be read without knowing both the detector's error rates and the setting.
| What you see | What it supports | What it does not support |
|---|---|---|
| High score on a long, unedited text | The text strongly resembles machine output; worth asking about the process | That the writer used a machine, or how much |
| High score on a short or formulaic text | Very little; short and formulaic texts score high often | Any conclusion about the writer |
| Low score | The text does not resemble what the detector expects | That the text is human-written; editing lowers scores |
| Different scores from different tools | The text is in the zone where detection is unreliable | Picking the tool that matches a suspicion |
Treat a score as a prompt to talk, not as evidence. Ask to see drafts and notes, ask the writer to explain their choices, and decide on that evidence. Record that the score was not the basis of the decision.
Worked example: 400 submissions and one flagged essay
A course receives 400 essays. Suppose 20 of them, 5 percent, were machine-written, and the detector in use catches 90 percent of machine text and flags 2 percent of honest writing. The counts are invented but realistic. The detector flags 18 of the 20 machine essays and 8 of the 380 honest ones, 26 flags in all. Eighteen flags are right, so a flagged essay has roughly a 69 percent chance of being machine-written and a 31 percent chance of belonging to an honest student. The same detector, same threshold, in a course where 1 essay in 100 is machine-written, flags about 4 machine essays and 8 honest ones: most of its flags are now wrong.
| Share of essays machine-written | Machine essays | Correct flags | Honest students flagged | Share of flags that are right |
|---|---|---|---|---|
| 1 in 100 | 4 | 4 | 8 | 33% |
| 5 in 100 | 20 | 18 | 8 | 69% |
| 20 in 100 | 80 | 72 | 6 | 92% |
| 50 in 100 | 200 | 180 | 4 | 98% |
The reader of a single score never sees this table, and the score does not print the row it belongs to. That is the whole problem. The person deciding what an 82 means has to supply the false positive rate, which the detector's maker should have measured, and the prevalence, which nobody knows precisely and which is usually far lower than a worried reader assumes.
Common mistakes when acting on a score
- Reading 82 as an 82 percent chance of machine use. It is a resemblance score on the tool's own scale, and the worked example shows how far the two can differ.
- Adding a second tool and treating two flags as confirmation. Detector errors are correlated: formulaic text trips several tools at once.
- Scoring a 200-word excerpt because the full text was inconvenient. Short inputs make every detector unstable.
- Penalising without asking for process evidence. Drafts and history are cheap to request and settle most cases in minutes.
- Ignoring that the writer is working in a second language, for which published studies found several times the false positive rate.
The 2 percent false positive rate in the worked example is a choice someone made when they set the threshold. At a looser setting that catches 95 percent of machine text and flags 5 percent of honest writing, the 5-in-100 course produces 19 correct flags and 19 wrong ones. The score on the screen looks the same either way.
Prevalence is the number nobody has. Surveys of student and professional writing put the share of fully machine-written submissions anywhere from a few percent to a fifth depending on the setting and the year, and the honest reading of that range is that the same flag can be mostly right in one classroom and mostly wrong in the next. A department that wants to use a detector at all should estimate its own prevalence from process evidence on a sample before it decides what a flag is worth.
Recording a flag properly
If a flag is written down at all, write down the tool and version, the threshold, the measured false positive rate on writing like this, the text length, the score, the date and the decision that was actually taken and on what evidence. A file that says only flagged 82 percent will be read a year later as an accusation. Where a score ever contributes to a decision about a person, the run itself should be kept in the form described in evidence-grade test records. Where it leads to a label on published content rather than a decision about a writer, the label carries the detector's error rate into public view, and the guide to testing AI disclosure labels covers what that label must be tested for before it goes live.
For writers who have been flagged
If your writing has been flagged, the most persuasive response is evidence of process: earlier drafts, version history from your editor, notes and sources, and the ability to talk about why you made the choices you did. Point out what the published research says about detector error rates, including the studies showing higher false positive rates for non-native English writers. The guide to proving you wrote it lists the evidence that carries most weight.
What organisations should publish
- Which detector is used, at what threshold, and its measured false positive rate on writing like theirs.
- That a score alone never leads to a penalty, and what evidence is considered instead.
- How a writer can respond, and who reviews the case.
Guidance from national bodies has moved in this direction. The Quality Assurance Agency for UK higher education, for example, has encouraged universities to design assessment around the process of learning rather than rely on detection, and many institutions now say detector output is not sufficient evidence on its own.
Common questions
What does a high AI detector score mean?
That the text resembles what the detector associates with machine writing. It is not the probability that the writer used a machine, and many human texts score high.
Does a low AI detector score prove a text is human-written?
No. Edited, paraphrased or mixed text often scores low. A low score is weak evidence either way.
Why do base rates matter for detector scores?
Because when machine-written text is uncommon, even a small false positive rate means a large share of flags point at honest writers.
Can a teacher fail a student on a detector score alone?
They should not. Published research shows detectors make errors, sometimes unevenly across groups. A score should lead to a conversation about the writing process.
What evidence is better than a detector score?
Drafts, version history, notes and sources, and a discussion in which the writer explains their work. These speak to how the text was produced.
What false positive rate is acceptable before a detector is used?
There is no universal number, but the harm sets the bar. For decisions about people, most careful policies require a measured rate around 1 percent or lower on writing like theirs, with the upper confidence bound below 2 percent, and still treat a flag as a prompt to ask rather than a finding.
Sources
This guide is part of the testing AI systems hub. It is best read alongside why ai detectors disagree about the same text and proving you wrote it, which cover the neighbouring questions.