Detection
Why false positives happen
Classifiers and writing patterns are probabilities, not proof. Here is what trips them, what a flag does and does not mean, and what to do about a wrong one.
Updated
Detection is probabilistic. It can flag human work as AI (a false positive) and miss AI work (a false negative). The card exists so you can see which detector produced a flag and how strong it is.
Signals that are near-certain when they fire
A verified Content Credentials signature, a matching watermark check, or metadata declaring "trained algorithmic media" mean the file itself says it was generated. These are wrong only when someone deliberately attached another file's provenance, which is rare.
Signals that are guesses
- Writing patterns. Formal, edited or non-native English often uses even sentences, lists of three, em dashes and tidy structure. Technical writers and lawyers trip these constantly. A model asked to write casually evades them.
- Generator classifiers. Stylised photography, 3D renders, illustrations, heavy filters, upscaled or heavily compressed images, and screenshots of generated images all confuse pixel classifiers.
- Model rechecks. A vision model describes what it sees and is sometimes confidently wrong about hands, text and lighting.
- Community reports reflect what other people believed, not what is true.
What a flag means
"This content scored above your threshold on these detectors." It is not a statement about the person who posted it, it is not verified by a human, and the terms prohibit using it to accuse, discipline or target anyone.
What to do about a wrong flag
- Open the card and choose Not AI. Enough honest reports move the score for everyone.
- Raise your sensitivity preset to Strict if weak flags bother you.
- Turn the extension off for that site with Scan this site.
- Email support with a link to the post if you think a detector is systematically wrong there.
Pro users can also Always allow this account or Always allow this post.