記事一覧

Who gets falsely flagged by AI detectors, and why

約 4 分

The finding that should have changed more than it did

In 2023 a Stanford team ran seven widely-used GPT detectors over TOEFL essays written by non-native English speakers and over essays by US eighth-graders. The detectors flagged more than half of the TOEFL essays as AI-generated. They flagged the native-speaker essays at close to zero. The essays were all human-written. The only variable that moved was who wrote them. This is not a bug in one vendor's model. It is a direct consequence of the mechanism, and it reproduces across tools.

Why it happens

Detectors look for low perplexity — text that goes where a language model expects — and low burstiness, meaning little variation in sentence rhythm. Someone writing in a second language typically has a narrower active vocabulary and draws on constructions they are confident are correct. That is careful, competent writing. It is also, measured statistically, more predictable and more regular than the writing of someone improvising freely in their first language. The detector is not detecting AI. It is detecting a property that AI writing shares with careful second-language writing, and it cannot tell the two apart because it is not measuring anything that distinguishes them.

Who else this hits

The same mechanism catches several other groups, for the same reason: Technical and scientific writers, whose fields enforce conventional phrasing and standardised structure. A methods section is supposed to read like every other methods section. Anyone writing to a template — business correspondence, clinical notes, structured report formats. Regularity is the point of the format. Grammar-tool users. Grammarly and similar products smooth exactly the irregularities detectors read as human: they shorten winding sentences, standardise transitions, normalise word choice. Running a human draft through a grammar checker measurably raises its AI score. And students who have been taught rigid essay structure, which is most of them. Formulaic writing is what schooling often rewards.

The asymmetry that makes this serious

False negatives — AI text scored as human — are easy to produce. Light editing or a paraphrasing pass will usually do it. Their cost is that someone gets away with something. False positives cost a person their standing. An academic-misconduct finding follows someone through a degree and sometimes beyond. The two error types are treated as symmetric in accuracy figures and are not remotely symmetric in consequence. And they compound: the same groups most likely to be falsely flagged are frequently the groups least equipped to contest it — international students navigating an unfamiliar disciplinary process in a second language, often without the institutional confidence to push back.

What institutions should take from this

Several universities have responded by turning AI detection off in their submission workflows, or by prohibiting its use as standalone evidence. That is a defensible reading of the data rather than a retreat. A workable policy usually has three parts. Detection scores may open an inquiry but may never close one. Any finding must rest on process evidence — drafts, version history, a discussion of the material. And students must be told when detection is in use, what threshold applies, and how to respond. Assessment design does more than policy. Supervised writing, oral defence of submitted work and process portfolios all make the question far less load-bearing, because they generate the evidence directly instead of inferring it.

What writers can do

Preparation beats argument, because the evidence that resolves these disputes has to exist before the dispute starts.

  1. 1Draft in an editor that keeps version history — Google Docs, Word online, or anything with a revision log
  2. 2Keep outlines, notes and sources rather than deleting them on submission
  3. 3Where a piece matters, self-check it before submitting so a high score is not a surprise
  4. 4Disclose AI assistance according to whatever rule applies to you, in writing
  5. 5If you use a grammar tool, know that it raises your score, and keep the pre-tool draft

Why we show confidence alongside the score

Most of the situations above produce a particular signature: a mid-range score, signals that disagree with each other, or text short enough that the estimate is unstable. That is precisely what a low confidence level describes. A detector that reports 61% and stops has told you almost nothing while sounding like it has told you something. One that reports 61% with low confidence and says the writing patterns are not clear enough for a strong result has told you the truth, which is that it does not know. That is not a marketing distinction. It is the difference between a tool that helps someone make a fair decision and one that supplies a number for an unfair one.

続けて読む

参考文献

accuracyfairnessstudents