← Back to Home

How to read a result

Our results are probabilities, not verdicts. This page explains what that difference means in practice, because the gap between the two is where people get hurt.

A score is an estimate, not a finding

When we say a piece of text is 85% likely to be AI-generated, we are not saying we found AI in it. We are saying that, across the detectors that examined it, the weight of evidence leans that way. The same number can come from several detectors agreeing mildly or from one detector being very confident while others abstain — which is why we show you each detector's vote rather than only the total.

Read the individual signals before you trust the headline number. If the detectors disagree with each other, that disagreement is the most useful thing on the page, and we deliberately do not smooth it away.

What a high score does not tell you

  • Who wrote it. Detection cannot identify an author, and a high score does not distinguish between text written by AI, text written by a person and edited by AI, or text written by a person who happens to write in a plain, regular style.
  • Whether anyone did anything wrong. Using AI is not misconduct in most contexts. What matters is the rule that applied, and a detector knows nothing about that.
  • Intent. An edited image might be a forgery or it might have been cropped and compressed by the app it was sent through.

False positives are real and they are not evenly distributed

Every detector, including ours, sometimes flags human work as machine-generated. This falls hardest on people who write in short, regular sentences with common vocabulary — which disproportionately means people writing in a language that is not their first, people writing to a template, and people who have been taught to write plainly.

Accuracy also degrades outside English. If you are checking Norwegian, or any language with less training data behind it, treat the result as weaker evidence than the same number would be in English.

If a result would change how you treat a person, it is not enough on its own. Go and find something that corroborates it, or set it aside.

If you are a teacher, an employer, or reviewing someone's work

Do not open a conversation by presenting a score as proof. It is not proof, and if you are wrong you will have accused someone of dishonesty on the basis of a statistical estimate about their writing style.

What works better:

  • Ask about the work. Someone who did it can discuss their choices, their sources, and their earlier drafts.
  • Look for corroboration that is independent of the detector — version history, drafts, citations that do not exist, a sudden change in voice mid-document.
  • Tell the person a tool flagged the work, and let them respond, rather than treating the flag as a conclusion you are testing them against.
  • Have a written rule about AI use, so the question is whether a rule was broken rather than whether a number is high.

Documents, images, video and speech

For uploaded documents, the most concrete evidence is usually the structural findings rather than the score — for example that a PDF was re-saved after its original export, or that it carries edit markers from a specific editor. Those describe something we actually observed in the file. Read them first.

Some checks do not apply to some file types, and when that happens we say so rather than printing a zero. A check that did not run is not a check that passed.

For images and video, ordinary processing leaves traces that resemble manipulation. Screenshots, re-encodes, social platform compression and filters all alter a file. A flag means the file was altered, not that it was altered to deceive.

When to trust a result more

  • Several independent detectors agree, rather than one being decisive.
  • The content is long. Short samples carry much less signal, and a sentence or two is close to unreadable for any detector.
  • The content is in English.
  • Structural or forensic findings back up the score with something specific about the file.
  • It matches something you already had reason to suspect, for an independent reason.

Why we tell you this

Most tools in this category advertise a single accuracy figure and let you assume it applies to your case. We think that is how people end up wrongly accused, so we would rather explain the limits up front and be trusted on the results that matter. We wrote about the industry's accuracy claims, and how we aggregate detector scores differently, in why every AI detector claims 99% accuracy.


← Back to Home