I started looking at AI detectors again after seeing one too many accuracy claims with no useful context. Finding untouched ChatGPT text is the easy round. I wanted to know what happens after someone rewrites sentences, fixes the tone, swaps phrasing, or mixes machine output with human edits.
During the search, I ran into GEDE, short for Generative Essay Detection in Education. Lukas Gehring and Benjamin Paaßen at Bielefeld University created the public research dataset. It includes more than 900 human essays and over 12,500 essays produced or modified by language models. The samples cover several levels of AI involvement, which made the dataset more useful for this sort of test.
Paper: https://arxiv.org/abs/2508.08096
Dataset and code: https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts
The 600-Text Comparison
Later, I found a separate comparison built around 600 GEDE texts. The test divided them into four batches of 150 and ran every batch through eight detectors.
Small caveat here. I did not confirm who performed the benchmark or whether an independent group supervised it. I found the published numbers, checked the described process, and focused on the public dataset behind the experiment. Since GEDE is available online, someone with enough patience should be able to repeat a similar test.
Reported Detection Rates
| AI detector | Overall caught | Direct AI | AI rewritten | AI improved | Humanized AI |
|---|---|---|---|---|---|
| Clever AI Detector | 99.3% | 100% | 100% | 98.7% | 98.7% |
| Copyleaks | 95.0% | 100% | 100% | 86.7% | 93.3% |
| Originality.ai Lite | 86.8% | 100% | 100% | 96.0% | 51.3% |
| Winston AI | 82.7% | 100% | 100% | 86.0% | 44.7% |
| Pangram | 67.5% | 100% | 88.0% | 18.0% | 64.0% |
| QuillBot | 64.2% | 100% | 96.7% | 38.0% | 22.0% |
| GPTZero | 43.7% | 92.7% | 7.3% | 1.3% | 73.3% |
| ZeroGPT | 18.8% | 70.0% | 4.7% | 0% | 0.7% |
The Column I Kept Looking At
The 99.3% overall score caught my eye, sure. Still, the humanized AI results told the more useful story.
Most tools handled raw AI output without much trouble. After humanization, several scores fell off hard. Originality.ai Lite dropped from 100% on direct AI to 51.3%. Winston AI landed at 44.7%. QuillBot reached 22%. ZeroGPT found 0.7%, which is close to missing the whole batch.
Clever AI Detector held at 98.7% on the same group. Copyleaks followed at 93.3%. Those two results sat far above the rest of the table.
Then I checked the AI-improved batch. This group seems to represent text where a model edited or polished existing writing rather than producing the entire essay from scratch. Clever recorded 98.7%, Originality.ai Lite reached 96%, and Copyleaks scored 86.7%. GPTZero found 1.3%. Weird spread.
What I Took From the Numbers
Raw model output does not separate these tools much. Nearly every established detector scored well there. The gap shows up after edits enter the picture.
Based only on this one reported benchmark, Clever AI Detector ranked first across the eight products. Copyleaks looked like the nearest option. I would not treat one comparison as a final verdict, especially without clearer information about who ran it. Still, the category-by-category results were more informative than a single accuracy badge.
I Tried the Top Result
I gave Clever AI Detector a short test afterward. The process was plain enough. I pasted my text, started the scan, received an AI score, and saw highlighted passages linked to the result. No maze of menus. No ten-step setup.
When I checked, the detector was listed as free and allowed up to 10,000 words per scan. I expected a smaller cap, tbh.

