What does AI detection precision measure?
AI detection precision answers a specific question: among the passages a detector flagged, how many had an AI label in the evaluation data? Divide the correctly flagged AI passages by all flagged passages. A false positive belongs in that denominator because the detector flagged it too.
Suppose a small test produces 40 flags. The test records say that 32 came from AI and eight came from human writers. Precision is 32 divided by 40, or 80%. This example says nothing about the passages that received no flag. There could be many missed AI passages outside that group.
Google's classification guide defines precision, recall, accuracy and false-positive rate separately. Write down the denominator before comparing percentages. For the broader question of whether a detector suits a particular kind of writing, start with how accurate AI detectors are.
A worked example from Zyphh's evaluation
Zyphh's published evaluation contains 2,643 accepted passages: 1,320 labeled AI and 1,323 labeled human. At the high-signal threshold of 0.65, the outcomes are:
| Recorded origin | Flagged in the high band | Below the high band | Total |
|---|---|---|---|
| AI | 923 | 397 | 1,320 |
| Human | 30 | 1,293 | 1,323 |
| Total | 953 | 1,690 | 2,643 |
The machine-readable model card records the measurements and evaluation scope. These passages form a source-grouped holdout from RAID's public training data. This is Zyphh's own evaluation, not a submission to RAID's official hidden test.
Below the high band means unflagged for this binary calculation. It includes mixed results and lower scores. It does not mean the tool verified human authorship. The matrix uses the benchmark's origin labels, which were not independently audited by Zyphh.
Calculate precision, recall and the false-positive rate
Precision uses the flagged column: 923 divided by 953 gives 96.85%. Of the 953 high-band results in this test, 923 had an AI label and 30 had a human label.
Recall uses the AI row: 923 divided by 1,320 gives 69.92%. The detector flagged about seven in ten AI passages. It left 397 AI passages below the high threshold.
The human false-positive rate uses the human row: 30 divided by 1,323 gives 2.27%. That is the share of human passages flagged, not the share of flags that were wrong. The latter is 30 divided by 953, or 3.15%.
Overall binary accuracy counts both kinds of correct outcome: 923 plus 1,293, divided by 2,643, gives 83.84%. The model card also reports balanced accuracy of 83.83%, which averages recall and the correct rejection rate.
The same test therefore supports both a high precision figure and a substantial count of missed AI passages. Report the 397 misses and the 30 human false positives alongside the percentages.
Why a score of 85 is not 85% precision
Precision is a measurement over a labeled collection. An individual score comes from applying the classifier to one passage. Zyphh displays its uncalibrated classifier output on a scale from 0 to 100. A score of 85 is not a statement that the passage has an 85% chance of AI authorship, and it is not the percentage of words written by AI.
The reported AUROC of 0.936 is another separate measurement. It summarizes ranking performance across thresholds. Writing "93.6% accurate" next to it would mislabel the result.
The methodology explains the scoring model, threshold selection and eligibility rules. Compare the recorded engine version and threshold before treating two reports as equivalent. Either can change results for identical text.
Keep the original score and band in a report. Do not silently convert a mixed result into a human label because a spreadsheet expects a binary answer.
Can a higher threshold improve precision?
A higher threshold requires a higher classifier score before assigning a flag. That can reduce false positives while missing more AI passages. It does not guarantee that precision rises in every small sample, because the outcome depends on which passages move across the boundary.
Imagine a second hypothetical test. At one threshold, 100 flags contain 90 AI passages and ten human passages: 90% precision. A stricter threshold leaves 50 flags containing 49 AI passages and one human passage: 98% precision. The percentage improved, but the detector now flags 41 fewer AI passages.
Choosing between those settings depends on what happens after a flag. A research triage queue and an authorship accusation have different consequences. Pick the threshold using development data and a stated purpose, then measure it on separate test data.
The AI text detector uses published bands. Moving a passage between those bands by editing it does not establish a different writing history.
How should mixed or inconclusive results be counted?
Keep them visible. Zyphh's evaluation includes 516 mixed-band passages out of 2,643. Its high-band binary metrics count those passages as unflagged. That convention makes the calculation reproducible, but it does not settle their origin.
An evaluation with an abstaining detector needs to report how often it declines to decide. For example, a system could appear precise by flagging only a handful of easy cases and leaving everything else unresolved. Readers need both the error counts and the number of passages that received each result.
Eligibility exclusions also belong in the report. Zyphh's model-card figures describe accepted passages, not every conceivable input. Empty text, unsupported material and passages outside the product's limits should not disappear into a claim about general performance.
When reviewing a real document, the step-by-step checking guide explains how to retain the result and compare it with writing-process evidence.
How to check precision on your own material
Start with material whose origin you can document. Write down what counts as AI for your evaluation before running it. Fully generated prose, human writing with spelling corrections and a heavily rewritten AI draft are different histories. If your use case includes mixed writing, label and report it separately.
Keep related passages together when splitting data. A human source and several generated rewrites of that source should not be scattered between development and test sets. Save the exact detector version, threshold and input text. Choose the final test set before using its results to revise your approach.
For each accepted test passage, record its known label and output band. Count the four matrix cells, calculate the metrics, and retain inconclusive and excluded cases in separate totals. Include lengths, language, genre and generation dates so another reader can see what the evaluation covers.
Small samples deserve plain counts. Nine correct flags out of ten is 90% precision, but one changed outcome moves that figure by ten percentage points. A confident-looking decimal cannot fix limited evidence.
The RAID study tested detector robustness across varied domains, generators and transformations. Its findings support checking those conditions rather than assuming a result from one collection transfers unchanged.
What should a performance claim tell you?
Look for the dataset, threshold, test size and labeling policy beside a performance claim, together with the name of whoever ran the evaluation. Zyphh's 96.85% precision belongs with its 953 high-band results and the limits of its 2024 benchmark source. Current generators, edited outputs and mixed authorship have not been separately validated in that evaluation.
Ask whether competing tools were tested on the same passages under comparable conditions. Zyphh has not run a matched comparison with GPTZero or ZeroGPT, so these figures cannot support a superiority claim.
Also check what the product actually measures. An AI classifier's precision does not describe a Unicode inspector or a provider's watermark verifier. The text watermark guide explains those different methods.
A detector's response to rewriting is a robustness observation. If you compare versions of a passage, keep both and check that the revision preserved the facts. The guide to paraphrasing without changing meaning provides a separate review process for that task.
How to use the number when reviewing a document
Use precision to understand an evaluation, then return to the document and its context. A population-level result cannot reconstruct who wrote a particular sentence. A high-band flag can prompt a review of drafts, sources and revision history, but the flag does not supply those records.
When recording a result, name the engine version and the band the passage received. Attach the score if useful, and link the benchmark's error counts separately. Do not replace that record with "96.85% likely to be AI"; the test did not establish that probability for the passage.
If you need to make a consequential decision, record the evidence separately from the interpretation. Include the original passage, exact result and any writing-process information you considered. Another reviewer should be able to see how you reached the conclusion and where uncertainty remains.