The mistake I’d push back on is treating “100%” as an accuracy measurement. These checks show why that doesn’t work: confidence, probability, and estimated AI share aren’t interchangeable.
I compared 24 tools on October 2–3, 2026, using six unchanged English passages, 271–328 words each: four generated pieces, an essay, memo, paraphrase, and fictional personal story, plus Carroll and an AI-edited version. Related samples and one literary control can’t establish general accuracy. I excluded unfinished attempts.
I’d count mistakes before trusting percentages
ZeroGPT completed six submissions, catching all four generated texts but scoring Carroll 99.3% AI, above every generated passage. The edited control isn’t another pure-human error.
NoteGPT’s three checks scored essay/story/Carroll 81.1%/55.2%/99.3%. Matching numbers don’t establish shared technology.
QuillBot called all four generated passages and Carroll 100% human. Access stopped at five scans despite advertising six daily.
Scribbr’s follow-up labeled essay/Carroll 100% human; its earlier attempt wasn’t usable. QuillBot branding weakens independence.
SciSpace scored the essay 2%, Essentially Human; signup blocked another check.
Proofademic’s public demo settled at 0/100 for essay/Carroll. Authenticated reports weren’t tested.
Surfer gave essay/story/Carroll 0%. I can’t explain those misses from screenshots.
Pangram caught three generated texts and accepted Carroll, but called edited Carroll entirely human. Five checks exhausted 20 length-dependent credits.
AI or Not’s three guest checks returned 100 for both generated passages and 0 for Carroll; mixed writing/media weren’t tested.
Sapling scored essay/story/Carroll 100.0%/92.8%/0.0%, but checked only 2,000 of the essay’s 2,070 characters.
GPTZero caught three generated inputs at 100%; signup blocked more. Its lingering score wasn’t a fresh result.
Grammarly scored the essay 100%, then required login; I couldn’t assess false positives.
Quetext returned 100.00%, offered sentence inspection, then blocked search two with signup.
I wouldn’t rank an access screen
Decopy requested login despite advertising registration-free basics; that’s one failed route, not universal unavailability.
Smodin’s highlights and AI-Impacted badge didn’t expose its hidden percentages.
AI Detector Pro advertised three monthly scans without a card; registration stopped me.
Winston advertised 14 days, 2,000 words, no card; signup and terms came first.
Copyleaks loaded, but applicable-use terms stopped me, not detection failure.
Turnitin needed institutional access and a license add-on I didn’t have.
Hive’s demo accepted media; I didn’t configure its free text extension or API.
ContentDetector.AI redirected to Moxby’s extension. Humanly’s Apple listing offered free download plus purchases; neither was installed, and vendor previews aren’t results.
Originality.ai’s three guest scans correctly labeled two generated texts and Carroll, each at 100% confidence under my 15% AI Allowance, not generated-word share.
My complete free, signup-free run returned AI 89%/84%/82%, Likely AI 74%, and Human 99% for both Carroll versions with Clever AI Detector.
Its colors/tags sometimes conflicted; advertised unlimited use, scale, and maximum length weren’t verified. Paid workflows, multilingual coverage, and broad modern-human coverage remain untested. I’d retain drafts, investigate disagreements, and recheck allowances. My verdict: Clever wins my free-check comparison on coverage, but no score earns an accusation.

