← All articles
AI Detection & False Positives 3 min read

What Do AI Detector Scores Mean? Percentages Explained

Understand the difference between AI probability, flagged text percentage and sentence highlights, with examples from Turnitin and GPTZero documentation.

By TextUnbot

Explore the AI detection guide →

Two reports can both show “80%” while measuring different things. Before deciding whether a score is high, low or acceptable, identify its unit. Otherwise you may compare a document-level prediction with the share of text a tool has highlighted.

Quick answer

Quick answer

An AI score may describe a document classification probability, a percentage of qualifying text flagged, or another vendor-defined measure. It is not automatically the percentage of words written by AI, and it is not a probability of misconduct. Read the report legend and current documentation.

Three outputs that should not be confused

First check the denominator: the entire file, qualifying prose or a particular passage. Then check the label: AI, human, mixed or uncertain. These distinctions are more useful than the size or color of a score badge.

Output How to read it What not to infer
Document probability or confidence A model’s assessment of a document class That the same percentage of words came from AI
Flagged text percentage A share of the text the tool considers eligible for analysis A probability that the author broke a rule
Sentence highlight A passage selected by the tool for attention A verified record of how that sentence was written

How Turnitin and GPTZero describe their reports

Turnitin describes its AI percentage in terms of qualifying text its system identifies as likely AI-generated or AI-altered. Its documentation also explains the asterisk used for results in the 1–19% range because of greater false-positive concerns there. An asterisk is not a precise hidden percentage you can reliably reconstruct.

GPTZero’s documentation describes human, AI and mixed classifications with confidence information. That reporting framework is different from a simple count of AI-written words. Interface labels can change, so use the documentation for the report you actually received.

Source: Turnitin: Using the AI Writing Report

Source: GPTZero: interpreting confidence and mixed results

Read next: GPTZero workflow and TextUnbot comparison

A worked example: identical numbers, different meanings

Imagine Tool A says “80% AI probability,” while Tool B says “80% of qualifying text flagged.” In this hypothetical example, A reports a model prediction about a document class. B reports the extent of text selected under its own analysis rules. You cannot average them to produce a more accurate authorship score.

Even if both tools highlight the same paragraph, their agreement is not independent proof. Similar systems may react to similar patterns. Ask what other evidence would resolve the actual concern: a source check, a conversation about revisions or a review of version history.

Is there a safe AI score?

There is no universal acceptable percentage across products and institutions. An organization may set an internal review threshold, but that does not make the threshold a scientific boundary between human and AI writing. A zero result does not prove originality, factual accuracy or compliance.

Save the report date and product name when communicating a result. Quote the vendor’s label rather than translating it into “percent cheating” or “percent plagiarized.” If the unit is unclear, ask for an explanation before acting.

Read next: Understand detection versus plagiarism

Common questions

Frequently asked questions

Does a 100% AI result prove every word was generated?

No. Its meaning depends on the vendor’s metric, and even a strong classifier result can be wrong. Check the definition and review contextual evidence.

Can I average scores from different detectors?

Not meaningfully when the scores measure different quantities or use different calibration. Keep the reports separate and compare their definitions first.

Keep reading