What question are you actually asking?
"Was AI involved?" sounds binary, but real writing histories are messy. One person drafts an article, asks a model to restructure a paragraph, accepts a grammar fix and rewrites the conclusion by hand. Another asks for an outline and writes every sentence personally. A classifier sees the final text and none of those steps.
So define your question before you read the output. Are you looking at fully generated passages, heavy AI revision, or any use of a writing assistant at all? A detector evaluated on one definition may not answer another, and length, genre, language and editing history all belong in the question too. Without them, a single headline accuracy figure says little about the text in front of you.
Why doesn't "99% accurate" mean a flag is 99% right?
Take 10,000 documents. Say 1% are fully AI-generated, the detector catches 90% of those, and it wrongly flags 1% of the human ones. That's 100 AI documents, 90 caught, plus 9,900 human documents, 99 wrongly flagged. Of the 189 flagged, 90 are AI-generated: about 48%.
These numbers are hypothetical and aren't measurements of this site. They show why prevalence matters. Run the same detector, with the same sensitivity and false-positive rate, on a collection that's 10% AI and you get 900 true positives and 90 false positives, so about 91% of flags are correct. The classifier stayed the same and the collection changed.
That's why "99% accurate" can't be read as "a flagged document has a 99% chance of being AI." When you see a performance claim, ask for the confusion matrix, the class balance and the conditions of the evaluation.
Is the Zyphh score a probability?
No. Zyphh’s 0–100 score is the output of a classifier trained on labeled RAID examples, displayed on an easy-to-read scale. It has not been calibrated as a probability of AI authorship. The methodology and benchmark report a source-grouped held-out test, including human false positives and missed AI samples.
A benchmark result does not establish performance on every text. Familiar vocabulary, genre, model generation settings and editing can change a score. The public corpus covers an earlier generation of models, and the evaluation does not separately validate current releases or mixed human–AI drafts.
The report also measures sentence rhythm, repeated wording and vocabulary. These descriptive observations are separate from the classifier score. Highlights do not identify AI-authored sentences or verify a statistical watermark.
What happens when a detector is wrong?
A false positive is a human-written document labeled as generated. Run into one during a personal experiment and it's an annoyance. Run into one in a classroom, a hiring process, a newsroom or a workplace and it can become a serious accusation. The same technical uncertainty carries very different consequences depending on what someone does with it.
Treat a detector result as a reason to look closer, not as a replacement for a fair process. Ask for drafts, notes, sources and the writer's own account. Don't make anyone "prove their humanity" by rewriting until the score moves, because that measures how the classifier reacts to edits and tells you nothing about what happened when the text was first written.
Predictable wording isn't evidence of misconduct. Standard professional phrases, repetitive instructions and tightly constrained factual answers can come from people as easily as from machines. A short excerpt stripped of its context is especially easy to misread.
Why do language and genre matter?
Test on material that looks like what you'll actually assess. If you're judging short technical abstracts from multilingual researchers, a benchmark of long English news articles leaves the important questions open. Genres differ in how much they repeat themselves, how they build sentences and how much specialized terminology they use. Researchers have also reported that some GPT detectors are biased against non-native English writers.
Record the language and genre of both your human and generated samples, and include edited and mixed examples if those occur in your setting. Keep translations separate from original-language writing, since translation changes the text and may involve another model. Don't stretch a reported result to languages or document types nobody evaluated.
This site's detector currently targets English prose and asks for at least 80 words. That's a product rule. It isn't a scientific line where detection suddenly becomes reliable.
How do you run your own evaluation?
Build a small dataset you have permission to use, with known origins and a written labeling policy. Keep the human texts independent of the generators you're studying. For generated texts, record the model, prompt, settings and date. For mixed texts, keep the edit history and decide in advance how you'll label partial assistance.
Split the collection before you touch any threshold: a development set to pick your approach, and a separate held-out set for the final measurement. If you keep choosing examples until the detector looks good, the number you end up with won't tell you how it performs on new material.
Report counts as well as percentages. "One false positive in 20 human documents" shows the sample size far better than an apparently precise rate. Give the number of human, generated and mixed documents along with their lengths, languages and genres, and include uncertain outputs instead of quietly dropping them.
Save the algorithm version, the date and the exact text you evaluated. Rule changes can alter later scores, so a screenshot won't let anyone reproduce your experiment.
What does an editing experiment actually show?
Comparing a draft with its paraphrase shows that the score changed after an edit. It doesn't show that the author changed, that the revision is human-written or that every watermark disappeared. A paraphrase can also damage facts or bring in a different model's statistical signal.
For robustness studies, log each transformation separately: translation, sentence reorder, deletion, human edit or model rewrite. Measure meaning preservation alongside detection. Shuffled words may strip out statistical context, but they destroy the document too, which makes them a poor stand-in for a real prose revision.
Trying edit after edit until a score looks good also bends the statistics, since the final result was picked from many attempts. Keep the whole record, and separate a single planned comparison from an adaptive search.
What should you do with a result?
If the result is inconclusive, read the stated reason. A short sample, uncertain language support or code-heavy text may not suit this method. If it's scored, look at the contributing features and ask whether the genre or writing task explains them. More matches aren't stronger evidence of authorship until a proper evaluation shows they are.
Suspect invisible characters? Use the watermark checker. Want a clearer draft? Try the paraphraser and review each suggestion. For research, export the algorithm version, source text and signal report. Keeping those questions apart stops a score from picking up authority it hasn't earned.