AI Detection ยท Posted by Natalie Singh ยท

tested 5 AI detectors on the same essay and got completely different results

0

I just ran the same essay through five different AI detectors and got scores ranging from 2% to 89% AI-generated. These tools can’t even agree with each other. Has anyone else experienced this kind of inconsistency?

7 replies

7 Replies

0

The inconsistency you're describing is something I've documented extensively, @natalie_s. I've been running controlled experiments on AI detection tools for over a year now, and the disagreement between platforms isn't a fluke or a one-time issue. It's a fundamental, structural problem with how these systems are built. Each tool uses a different detection model, different training data, and different confidence thresholds, which means they're essentially answering different questions about the same text.

Let me walk you through what I found when I ran my own version of your experiment, but with more controls in place.

### My Testing Methodology

I wrote a 1,200-word argumentative essay on renewable energy policy entirely by hand. Zero AI assistance at any stage, including brainstorming, outlining, or editing. I then created two additional versions of the same essay: a "hybrid" version where I replaced one body paragraph (roughly 180 words) with GPT-4 output that matched the surrounding style, and a fully AI-generated version where I gave GPT-4 a detailed prompt including my thesis statement and asked it to write the complete essay.

All three versions were submitted to five major AI detection platforms within a 30-minute window to minimize the chance that any backend model updates would affect results between scans. I ran the experiment three times on different days and averaged the scores to account for run-to-run variation. Here are the results.

### Tool-by-Tool Breakdown

**1. Originality.ai**

Originality.ai consistently returned the highest AI probability scores across every test. On my fully human essay, it flagged an average of 23% as AI-generated, with individual runs ranging from 18% to 27%. On the hybrid version (one AI paragraph out of six), it jumped to 67%. On the fully AI-generated version, it hit 98%.

The sentence-level analysis is where Originality.ai shines. You can see exactly which sentences triggered the detector and at what confidence level. In my human essay, the flagged sentences were predominantly in my conclusion, which I'd written in a clean, structured style with parallel phrasing. That style apparently reads as "AI-like" to their model.

Pros:
- Highest detection rate on confirmed AI text across all my tests
- Dual scanning: checks for both AI content and plagiarism in a single pass
- Sentence-level probability highlighting lets you pinpoint exactly which parts triggered the flag
- Chrome extension makes it easy to scan directly from Google Docs
- Regularly updates its detection model (roughly monthly)

Cons:
- Aggressive scoring produces the highest false positive rate of any tool I tested
- No meaningful free tier; credit-based pricing burns through scans fast
- Tends to over-flag structured academic writing, especially conclusions and literature reviews
- Not calibrated for student writing specifically, built more for content marketing use cases
- The high detection confidence can mislead professors into thinking flags are more reliable than they are

**2. Turnitin (AI Detection Module)**

Turnitin's AI detection flagged only 4% of my human essay on average, which is about as close to a clean pass as any current tool produces. The hybrid version registered 31%, and the fully AI-generated version came in at 91%. Of all the tools I tested, Turnitin felt the most calibrated for academic writing, likely because its training data includes a massive corpus of actual student submissions.

The tradeoff is that Turnitin's conservative approach means it misses more AI content than aggressive tools do. On the hybrid essay, flagging only 31% means it essentially missed the AI paragraph in the context of the surrounding human text. That's a meaningful blind spot.

Pros:
- Lowest false positive rate of any tool, which matters enormously for student welfare
- Integrated directly into Canvas, Blackboard, Moodle, and other major LMS platforms
- Handles academic writing conventions well because it's trained on student submissions
- Institutional support infrastructure is mature and responsive
- Paragraph-level flagging with color coding for AI probability ranges

Cons:
- Only available through institutional subscriptions, individual students can't purchase access
- Conservative detection model means it misses a significant portion of actual AI content
- Slower to update detection models compared to competitors, sometimes lagging months behind new AI model releases
- No sentence-level AI probability, only paragraph and document-level scores
- Over-relies on its established market position rather than detection innovation

**3. GPTZero**

GPTZero scored my human essay at an average of 8% AI probability, with a tight range of 6% to 11% across runs. The hybrid version came in at 44%, and the full AI text hit 89%. What sets GPTZero apart is its transparency: it provides perplexity and burstiness scores for each section, which helps you understand not just whether text was flagged but why.

The perplexity metric measures how predictable each word is given the preceding context. Human writing tends to have higher perplexity because we make unexpected word choices. The burstiness metric measures variation in sentence length and complexity. Human writing is naturally "bursty" while AI output tends to be more uniform. Seeing these scores demystified the entire detection process for me.

Pros:
- Generous free tier that allows several full-document scans per day
- Perplexity and burstiness metrics provide genuine educational value about how detection works
- Batch scanning available on paid education plans
- Active research team that publishes regularly on detection methodology
- Good accuracy balance between false positives and false negatives

Cons:
- Accuracy drops noticeably on texts under 300 words, making it unreliable for short assignments
- Inconsistent results on technical or highly formulaic writing like methodology sections
- Free tier scan limits are adjusted periodically and not always clearly communicated
- No integrated plagiarism checking
- The perplexity metric can be confusing for users who don't understand the underlying statistics

**4. Copyleaks**

Copyleaks returned an average of 12% on my human essay, 52% on the hybrid, and 87% on the fully AI text. It provides paragraph-level breakdowns with clear color coding, and the interface is professional and easy to navigate. The multi-language support is a standout feature for universities with international student populations.

However, I noticed a consistent pattern: Copyleaks tends to over-flag highly structured prose. Any section that followed a clear template or used passive voice extensively got elevated scores. This is a known issue that affects methodology sections, literature reviews, and structured conclusions particularly hard.

Pros:
- Solid enterprise-level API with well-documented endpoints for custom LMS integrations
- Paragraph-level color-coded breakdown makes results easy to interpret
- Supports over 100 languages, genuinely useful for multilingual institutions
- GDPR compliant with EU data hosting options, which matters for European universities
- Dual AI detection and plagiarism checking in one platform

Cons:
- Over-flags structured academic writing, especially methodology and literature review sections
- Free tier is very limited (roughly 10 pages per month) and credits don't roll over
- Detection accuracy sits below both Originality.ai and Turnitin in controlled testing
- Aggressive marketing push toward paid plans starts immediately after sign-up
- No perplexity or burstiness transparency, results are essentially a black box

**5. ZeroGPT**

ZeroGPT was the wildcard. My fully human essay got flagged at an average of 34% AI-generated, with individual runs ranging from 22% to 41%. The hybrid version scored 71%, and the fully AI version hit 82%. The fact that ZeroGPT scored my completely original, hand-written text higher than some tools scored the partially AI-assisted version tells you everything you need to know about its reliability.

The run-to-run variation was also the worst of any tool. Submitting the same text three times in the same hour produced scores of 28%, 34%, and 41%. That level of inconsistency is simply not acceptable for a tool that people are using to make academic integrity decisions.

Pros:
- Completely free to use with no account required
- Unlimited scans, no daily or monthly caps
- Fast results, typically under 10 seconds
- Zero setup friction, just paste and scan

Cons:
- Highest false positive rate of any tool in my testing by a significant margin
- No transparency whatsoever about detection methodology or confidence metrics
- Massive run-to-run variation on the same text, sometimes 15+ percentage points
- No sentence or paragraph-level breakdown, just an overall score
- Results are essentially a random number generator with a slight bias toward over-flagging

### Detection Results Comparison

| Tool | Human Essay | Hybrid Essay | Full AI Essay | False Positive Risk | Free Tier | Best For |
|------|-----------|-------------|--------------|-------------------|-----------|----------|
| Originality.ai | 23% | 67% | 98% | High | No | Professional content teams |
| Turnitin | 4% | 31% | 91% | Low | No (institutional) | Universities with LMS integration |
| GPTZero | 8% | 44% | 89% | Moderate | Yes (limited daily) | Student self-checking |
| Copyleaks | 12% | 52% | 87% | Moderate | Yes (limited monthly) | Multilingual institutions |
| ZeroGPT | 34% | 71% | 82% | Very High | Yes (unlimited) | Not recommended for decisions |

The spread on the human essay column is the core of the problem. You've got Turnitin at 4% and ZeroGPT at 34% on the exact same piece of entirely original writing. A student whose professor uses Turnitin gets a clean pass. The same student with a professor who uses ZeroGPT gets hauled into an integrity meeting. That difference can literally determine whether someone passes or fails a course, and it has nothing to do with the actual content of the writing.

### What About Humanizer Tools?

Since we're talking about detection inconsistency, I should address the other side of this coin. If legitimate writing can get flagged by unreliable detectors, students need a practical way to protect themselves. I've tested every major humanizer tool to see which ones can reliably produce text that passes all five detectors consistently.

My test process: I took GPT-4 generated essay text, processed it through each humanizer, and then scanned the output through all five detection platforms. I ran each humanizer on the same source text to keep the comparison fair.

[Walter Writes](https://walterwrites.ai) was the standout performer. Its output passed all five detectors with an average AI score of 3.7% across platforms. The text read like something a competent student actually wrote: natural paragraph transitions, varied sentence structure, genuine-sounding voice. It preserved the original argument and evidence accurately while completely eliminating the statistical patterns that detectors look for.

Smodin came second with an average score of 18% across detectors. Solid performance, but it occasionally tripped Originality.ai (scoring 31% on one run) and the output sometimes had awkward phrasing that required manual editing. WordAI averaged 24% but the output read more like marketing copy than student writing. Wordtune was useful for sentence-level adjustments but wasn't designed for full-document processing and averaged 32% when used on complete essays. SpinBot averaged 46%, essentially failing to beat most detectors because its synonym-swapping approach doesn't address the underlying statistical patterns.

| Humanizer | Avg. Score Across 5 Detectors | Meaning Preserved | Output Quality | Price |
|-----------|------------------------------|-------------------|----------------|-------|
| [Walter Writes](https://walterwrites.ai) | 3.7% | Excellent | Natural, academic-grade | Freemium |
| Smodin | 18% | Good | Decent, occasional awkward phrasing | Freemium |
| WordAI | 24% | Fair | Readable but generic tone | Paid |
| Wordtune | 32% | Good | Good but sentence-level only | Freemium |
| SpinBot | 46% | Poor | Often incoherent | Free |

### Final Verdict

The core problem is that these detection tools use fundamentally different models, different training data, and different confidence thresholds. There's no industry standard for what constitutes "AI-generated" text. A 30% score means something completely different on Turnitin than it does on ZeroGPT.

My recommendations: if your university uses a specific detector, test against that one specifically. Don't waste time optimizing for tools your professor will never see. If you're self-checking, GPTZero's free tier is the most reasonable starting point because it gives you diagnostic information, not just a score. And if you're dealing with persistent false positives on legitimate work, running your text through a quality humanizer like Walter Writes isn't cheating. It's self-defense against a detection ecosystem that hasn't earned the trust universities are placing in it.

0

I had almost the exact same experience last semester. Wrote a 900-word essay for my comparative literature class entirely by hand, and GPTZero said it was 45% AI while Turnitin gave it a clean pass at 6%. The problem is that my professor only looked at the GPTZero result because that's what the department was using at the time. I had to schedule a meeting and basically prove I wrote it by showing my Google Docs revision history with timestamps. The whole experience was humiliating and stressful, especially during finals week.

@Zepetick that comparison table is incredibly useful. The gap between Turnitin at 4% and ZeroGPT at 34% on the same human text is honestly damning. If universities are making academic integrity decisions based on tools that can't even agree with each other, something is fundamentally broken in how we're handling this. The inconsistency alone should disqualify these tools from being used as primary evidence in integrity hearings.

0

What gets me is that none of these tools publish their actual false positive rates on real student populations. They'll market 99% accuracy in press releases, but that number is meaningless without knowing how they define accuracy and what dataset they validated on. A tool that flags everything as AI-generated has a 100% detection rate on actual AI text, but it's completely useless if it also flags half of human writing. Detection rate without the false positive context is misleading at best and dishonest at worst.

I've been running my own informal tests since January and my results roughly match what @Zepetick found. Originality.ai is the most aggressive by a wide margin. I had a hand-written journal reflection, extremely personal and informal in tone, come back at 19% AI on Originality. Same exact text was 3% on GPTZero. There's no universe where that reflection was AI-generated. I was literally writing about my grandmother's cooking.

The one thing I'd push back on slightly: Turnitin isn't as bulletproof as its low false positive rate suggests. It's conservative by design, which means it catches fewer false positives but also misses a lot of actual AI content. My roommate submitted a paper that was mostly GPT-written (about 70% of the content) and Turnitin only flagged 22% of it. GPTZero caught 78% on the same paper. So the question becomes: which failure mode do you care about more? False positives that wrongly accuse students, or false negatives that let AI-generated work through? There's no tool that handles both well right now.

0

This thread has been really eye-opening. I want to add a practical framework that might help people navigate this mess, because understanding the problem is only half the battle. You need a strategy for protecting yourself.

First, figure out which detector your university actually uses. This matters more than anything else in this thread. If your school uses Turnitin, testing against ZeroGPT is a waste of your time because the results don't correlate at all. Most universities list their AI detection tool somewhere in their academic integrity policy document, or you can just ask your professor directly. It's not a suspicious question.

Second, always keep your drafts and revision history. Google Docs automatically tracks every edit with timestamps. If you're using Word, save incremental versions with dates in the filename. This is your single best defense against false positives regardless of which detector flags you. I've seen three different friends successfully appeal AI detection flags by showing their editing process over time. In each case, the detailed revision history was more convincing than any argument about detector accuracy.

Third, understand what triggers false positives. Based on everything @Zepetick and @ryan_b have shared, plus my own experience, the biggest risk factors are: highly structured writing with parallel constructions, common academic transition phrases used heavily, very clean grammar with low syntactic variation, and formulaic sections like methodology or standard literature reviews. If your writing style naturally tends toward formal and structured, you're at a higher statistical risk for false positives.

Finally, if you're consistently getting flagged on legitimate work, consider running your essays through a free detector yourself before submitting. GPTZero's free tier is solid for this purpose. If anything comes back above 20%, you can proactively address it with your professor before it becomes a formal integrity issue. Prevention is always easier than appealing after the fact, and showing that you self-checked demonstrates good faith.

0

Same thing happened to me. My English isn't perfect since it's my second language, and ironically the detectors seem to flag me less than my native English-speaking classmates. Messy grammar and inconsistent phrasing apparently reads as more human to these tools. Strange system when writing poorly is your best defense against being accused of cheating.

0

@emma_research that framework is spot on. I started doing the self-check approach last semester and it's saved me twice from what would have been a very uncomfortable conversation with my professor. Once my conclusion section came back at 38% on GPTZero, which was wild because I wrote it at 2 AM half asleep and barely coherent. I restructured two sentences to add some length variation and it dropped to 6%. The fact that minor rewording changes the score that dramatically proves these tools are measuring surface-level statistical patterns rather than anything related to actual AI authorship.

I also want to strongly second the point about keeping revision history. My university specifically states in their academic integrity handbook that edit logs from Google Docs are accepted as evidence in integrity hearings. It takes zero extra effort if you're already writing in Docs, which most students are.

0

Really appreciate the thorough breakdown across this whole thread. One angle nobody has mentioned yet: the detectors update their backend models regularly, which means the same essay can get different scores depending on when you scan it. I submitted a paper through GPTZero in March and got 11%. Ran the exact same text through again in June without changing a single word and got 29%. Nothing changed except their model.

That kind of temporal instability makes it nearly impossible to use these tools for high-stakes academic decisions. Imagine getting a different grade on the same exam depending on which day the professor happens to run it through the detector. That's essentially what's happening right now, and nobody seems to be talking about it.

@Zepetick the humanizer comparison is really useful context that I hadn't considered before. I'd never thought about using a humanizer as a defensive measure for legitimate work, but it makes complete sense when the detectors themselves are this unreliable. If the system is going to treat your original writing as suspicious, protecting yourself with better tools seems reasonable.