How Does AI Text Detection Work? What Detectors Measure and Why They Misfire
- Filed by
- Marketing
- Received
- Length
- 4 min

AI text detectors do not recognize machine writing the way you might recognize a friend's handwriting. Most estimate how predictable a passage is, compare its patterns with samples of known human and machine text, and return a probability. That number is an estimate, not proof, and treating it as a verdict is the most common mistake people make with these tools.
The main approaches
Predictability scores
A language model reads the text and measures how surprising each next token is, a quantity known as perplexity. Text produced by a model tends to follow likely continuations, so it scores as highly predictable, while human writing usually contains more unexpected word choices. The logic is sound, but plain, careful human prose can look predictable too. Our explainer on tokens in AI covers the units these scores are built on.
Variation across sentences
Some detectors look at how much predictability and sentence length change through a passage, sometimes called burstiness. People tend to mix short and long sentences and shift rhythm; machine drafts are often more even.
Trained classifiers
Another approach trains a separate model on large sets of labeled examples, human and machine, and lets it learn whatever features separate them. Such classifiers can do well on text similar to their training data and poorly on anything new.
Watermarks
A generator can deliberately bias its word choices in a hidden statistical pattern that a matching checker can later find. This only works if the tool that wrote the text applied a watermark, and heavy editing or rephrasing weakens it.
Process evidence
Version history, drafts and notes are not detection in the technical sense, yet they are often the most convincing record of how a document came about.
| Method | What it examines | Main weakness |
|---|---|---|
| Predictability score | How expected each word is | Flags simple, formal human writing |
| Sentence variation | Changes in rhythm and length | Easily disturbed by light editing |
| Classifier | Patterns learned from labeled samples | Struggles with newer models and mixed text |
| Watermark | A hidden pattern added by the generator | Absent unless the generating tool added it |
| Process evidence | Drafts and edit history | Not always available |
Why detectors misfire
- False positives. Formulaic writing such as legal notices, product specifications, templates and lists can score as machine-like. People writing in a second language, who often favor common words and safe structures, are at particular risk.
- Short passages. A few sentences give too little evidence for a meaningful score.
- Mixed authorship. A human draft polished by a tool, or a machine draft reworked by a person, sits in between, and scores swing accordingly.
- Moving targets. New generators write differently from the ones a detector was trained on.
How to tell if text is AI generated without a tool
Experienced editors notice certain signals, none of which settles the question on its own, because people produce them too:
- General statements with no concrete examples, names or numbers
- References, quotes or studies that cannot be found anywhere, a classic symptom of the hallucinations described in what a large language model is
- Very even structure: every paragraph the same length, every list exactly three items long
- A voice that does not match the writer's earlier work
- Confident factual errors, especially about recent events
Using results responsibly
- Set the rules first. Say in advance which kinds of AI assistance are acceptable and how they should be disclosed.
- Treat a score as a reason to talk, not a finding. Ask the writer about their process and sources.
- Look at the work around the text. Drafts, outlines and edit history usually reveal more than any detector.
- Never act on a score alone. Grades, payments or disciplinary steps need more evidence than one tool's estimate.
Schools wrestle with the same question; our piece on AI in education looks at measured ways to bring these tools into the classroom.
For content and marketing teams
Readers and clients care whether a text is accurate, specific and useful. A workable policy focuses on those qualities: a named person checks facts and claims, drafts are edited into the brand's own voice, and AI assistance is disclosed wherever a client contract or publishing platform asks for it. Practical guidance on drafting with care, from client emails to formal requests, is in how to write a letter with AI.
Quick answers
Can a detector prove who wrote a text?
No. It gives a probability based on patterns. Process evidence and a conversation with the writer come far closer to establishing authorship.
Are detectors more reliable on long documents?
They generally have more to work with on longer passages, but formulaic long documents can still be misjudged.
Should a team buy a detector at all?
It can serve as one input among several. Clear rules, editorial review and an open conversation with writers usually do more to protect quality than any score.
- 70
- articles
- 10
- topics
- 2
- min average read
- 2020–2026
- years covered



