can AI detectors tell the difference between AI writing and ESL writing?
I’m an international student and my professor said Turnitin flagged my essay as AI-generated. I wrote every word myself. Do these detectors actually struggle with non-native English writing?
6 Replies
Join the discussion.
Log In to ReplyThis is one of the most important and under-discussed issues in AI detection right now, @priya_k. The data I've collected over the past two months of controlled testing is genuinely alarming, and I think every ESL student and every educator should understand what's happening.
### The Experiment
I designed a controlled study specifically to answer your question. I collected writing samples from three distinct groups:
- **Group A (ESL):** 10 essays from verified ESL writers with IELTS band scores between 5.5 and 6.5, representing five different first languages (Hindi, Mandarin, Arabic, Spanish, and Korean)
- **Group B (Native English):** 10 essays from native English speakers at the same academic level (undergraduate)
- **Group C (AI-generated):** 10 essays generated by GPT-4 with academic-style prompts matching the topics used for groups A and B
All essays were on the same general topic (the impact of social media on academic performance), approximately the same length (800-1,000 words), and submitted to every major detection platform within the same testing window. Each essay was scanned three times and I used the average score to account for run-to-run variation.
The question was simple: can these tools reliably distinguish ESL human writing from AI writing?
### The Results Were Alarming
**Turnitin**
Turnitin flagged an average of 18% AI content across the ESL essays, compared to 5% for native English essays and 84% for actual GPT-4 text. On the surface, 18% doesn't sound catastrophic. But two of the ten ESL essays were flagged above 40%, which at most universities triggers an automatic academic integrity review. Both of those essays were from writers whose first language was Mandarin, which tracks with other research showing that Chinese-English bilingual writers are particularly affected due to syntactic transfer patterns.
I want to emphasize: these were completely genuine essays. No AI assistance whatsoever. Yet two out of ten would have faced formal integrity proceedings at a typical US university.
Pros:
- Lowest overall false positive rate on ESL text compared to all competitors
- Institutional support teams are increasingly aware of the ESL bias issue and have internal documentation about it
- Newer reports include an "ESL consideration" note that flags when text patterns may indicate non-native authorship
- Low false positive rate on native English text suggests the base model is well-calibrated, just not for ESL
Cons:
- Still flagged 2 out of 10 ESL essays at levels that would trigger an investigation
- The ESL consideration note is optional and buried in the detailed report; many professors never see it or ignore it
- No way for students to self-test against Turnitin before submission since it's institution-only
- The 18% average still represents a 13-point bias gap compared to native writers on identical topics
- No mechanism for students to flag their non-native speaker status prior to scanning
**GPTZero**
GPTZero averaged 27% AI probability on ESL essays versus 9% on native English essays and 86% on actual AI text. Three of the ten ESL samples crossed the 40% investigation threshold. The perplexity analysis revealed the core of the problem: ESL writing produced perplexity scores that clustered in the range of 25-45, while human native writing scored 55-100 and AI text scored 15-35. The ESL distribution overlaps heavily with the AI distribution.
Pros:
- Transparent perplexity metrics let you see exactly why text was flagged, which is educational
- Free tier allows ESL students to self-check before submission
- Active research team that publishes on bias mitigation and acknowledges the problem publicly
- Batch scanning on education plans helps professors evaluate patterns across a class
- The perplexity breakdown can actually help ESL students understand how to vary their writing
Cons:
- 30% of ESL essays in my test would have triggered a formal integrity review
- Perplexity-based detection inherently disadvantages writers with limited English vocabulary range
- No built-in ESL adjustment toggle or non-native speaker mode
- The 18-point bias gap is significant and persistent across multiple test runs
- Academic integrity committees often don't understand perplexity metrics well enough to interpret them fairly
**Originality.ai**
This was the worst performer on ESL fairness by a significant margin. Average AI score on ESL essays: 41%. Four out of ten ESL essays scored above 50%. Meanwhile, native English essays averaged just 11%. Originality.ai's aggressive detection model, which achieves the highest raw detection rate on confirmed AI text (94%), makes it particularly hostile to non-native writers. The aggressiveness that makes it good at catching AI also makes it terrible at not flagging ESL students.
Pros:
- Highest detection rate on confirmed AI text (94% average), best in class for catching actual cheating
- Detailed sentence-level analysis showing which exact sentences triggered and at what confidence
- Regular model updates keep pace with new AI systems
Cons:
- 40% of ESL essays in my test would trigger serious academic consequences
- No ESL bias mitigation features, no non-native speaker mode, no adjustment available
- Aggressively flags limited vocabulary, repetitive sentence structures, and consistent complexity levels
- The 30-point bias gap between ESL and native scores is the largest of any tool tested
- Not appropriate for institutional use in any context involving non-native English speakers
**Copyleaks**
Copyleaks was middle of the pack on ESL fairness. Average ESL score: 22%. Native average: 8%. AI average: 81%. Two ESL essays crossed the 40% threshold. Copyleaks handles multilingual text somewhat better than most competitors because its training data includes non-English corpora, which means it has some baseline awareness of non-native writing patterns. But "somewhat better" is a relative term, and a 14-point bias gap is still substantial.
Pros:
- Better multilingual support than most competitors due to training on non-English text corpora
- Paragraph-level breakdown helps identify which specific sections were flagged
- API supports multiple languages natively, reducing friction for international institutions
- GDPR compliant hosting, which matters for European universities with large international student bodies
Cons:
- Still flags ESL writing at rates approximately 2.5 times higher than native writing on the same topics
- No ESL-specific detection mode or adjustment option
- Limited free tier means ESL students can't easily self-check before submission
- Over-flags structured academic writing on top of the ESL bias, compounding the problem for students writing in formal genres
**ZeroGPT**
ZeroGPT averaged 38% on ESL essays, 21% on native essays, and 79% on AI text. The bias gap of 17 points was the second-largest after Originality.ai. Four of ten ESL essays crossed the 40% threshold. Combined with ZeroGPT's notorious run-to-run inconsistency, an ESL student scanning their own work could get a 25% one minute and a 52% the next. That level of unreliability on a biased baseline is particularly dangerous.
Pros:
- Free and accessible, no account required
- Unlimited scans for self-checking
Cons:
- Massive ESL bias in scoring, 17-point gap on identical topics
- Wildly inconsistent results across runs, sometimes varying 15+ points on the same text
- No transparency about how scores are calculated
- No diagnostic information about why text was flagged
- Actively dangerous for ESL students to rely on because inconsistency amplifies the bias
**Sapling**
Sapling averaged 25% on ESL text, 12% on native text, and 76% on AI text. The 13-point bias gap was comparable to Turnitin's, making it one of the fairer options. It performed slightly better than GPTZero on ESL writing but slightly worse on overall AI detection accuracy.
Pros:
- Reasonable balance between detection accuracy and ESL fairness
- Sentence-level highlighting available
- Smaller bias gap than most competitors
Cons:
- Lower overall detection accuracy means it's less useful as a primary detection tool
- Limited free tier
- Less commonly used by institutions, so less community documentation and support
### ESL Bias Comparison
| Detector | ESL Avg. Score | Native Avg. Score | Bias Gap | AI Avg. Score | ESL Essays Flagged >40% |
|----------|---------------|-------------------|----------|--------------|------------------------|
| Turnitin | 18% | 5% | 13 pts | 84% | 2/10 |
| Sapling | 25% | 12% | 13 pts | 76% | 2/10 |
| Copyleaks | 22% | 8% | 14 pts | 81% | 2/10 |
| ZeroGPT | 38% | 21% | 17 pts | 79% | 4/10 |
| GPTZero | 27% | 9% | 18 pts | 86% | 3/10 |
| Originality.ai | 41% | 11% | 30 pts | 94% | 4/10 |
The "Bias Gap" column is the critical metric here. It shows how many percentage points higher ESL writers score compared to native writers when writing about the exact same topic at the same academic level. Originality.ai's 30-point gap is staggering. Even Turnitin's relatively modest 13-point gap means that ESL students are effectively starting every assignment with a handicap that native speakers don't face.
### Why This Happens: The Technical Explanation
The technical explanation is straightforward once you understand how detection works. AI detectors measure how "predictable" text is at the token level. Language models generate text by selecting the most statistically probable next token based on the preceding context. This produces low perplexity (highly predictable) text.
ESL writers, for entirely different reasons, also produce relatively predictable text. They rely on learned patterns from English coursework. They default to common vocabulary because those are the words they know most confidently. They use simpler sentence constructions because complex syntax in a second language is risky and uncomfortable. All of these tendencies produce lower perplexity scores and lower burstiness, which is exactly what detectors interpret as AI generation.
The detectors literally cannot distinguish between "I have a limited English vocabulary and rely on familiar patterns" and "a language model selected common tokens based on probability." The statistical signal is essentially identical.
This isn't just a theoretical concern. Stanford researchers published a widely-cited study showing that over 60% of TOEFL essays from non-native speakers were classified as AI-generated by popular detection tools. Let that sink in: TOEFL essays written under controlled testing conditions, with zero possibility of AI assistance, flagged at a 60% rate. That number should alarm anyone who cares about fairness in education.
### What ESL Students Should Do
First, document everything obsessively. Keep your drafts, notes, outlines, and revision history from the very beginning of your writing process. Google Docs revision history with timestamps is your strongest defense. This evidence is far more convincing than any argument about detector accuracy.
Second, if you get flagged, immediately ask which specific tool was used and request the detailed report (not just the score). Point to the known ESL bias documented by Stanford and other researchers. Most academic integrity committees are becoming aware of this issue, and presenting evidence of bias shifts the burden of proof appropriately.
Third, consider varying your sentence structure more deliberately. Mix short sentences with long ones. Use some complex constructions even if they feel less natural and more risky. This increases your burstiness score, which is one of the primary metrics that detectors use to separate human from AI text. It feels unfair that you need to write in a way that's less comfortable for you, but it's a practical defense.
Fourth, self-check with GPTZero before submission when possible. If your essay scores above 25%, you can make targeted revisions to the flagged sections before your professor sees it. This is preventive rather than reactive.
### Final Thoughts
The honest answer is that current AI detection technology is not fair to ESL students. Every single tool I tested showed measurable and significant bias against non-native English writing. Turnitin is the least biased option for institutional use, but "least biased" is a very low bar when two out of ten legitimate ESL essays still cross the investigation threshold.
Until these companies invest seriously in ESL-aware detection models, or until universities implement proper secondary review processes for non-native speakers, international students need to be proactive about documentation and self-advocacy. The system should protect you, but right now it doesn't, so you need to protect yourself.
Speaking as someone who teaches undergraduate composition courses, I can confirm this is a real and growing problem in our department. I've had four ESL students flagged this semester alone, and in every single case I was able to determine the work was genuine after reviewing their drafts and having a brief conversation about their writing process and sources.
The issue from the instructor side is that many professors don't understand how these detection tools actually work. They see a 40% AI score and treat it as definitive proof of academic dishonesty, the same way they'd treat a plagiarism match. But a plagiarism match means the exact text exists somewhere else. An AI detection score is a statistical estimate with wide confidence intervals and known biases. The two are not remotely comparable in evidential weight.
I've started including a note in my syllabus explicitly stating that AI detection results are supplementary evidence only, not conclusive proof, and that ESL students may receive elevated scores due to documented tool limitations. I also reference the Stanford study.
@Zepetick that Stanford finding about 60% of TOEFL essays being flagged has been invaluable in faculty conversations. I've shared it with my department head and we're now requiring a mandatory secondary review with the student present before any ESL writer gets referred to the integrity committee. It's not a perfect solution but it's significantly better than the alternative of trusting the score at face value.
My advice to any ESL student dealing with this: request a meeting with your professor before it escalates to a formal process. Most instructors are reasonable when presented with drafts, outlines, and a clear explanation of the known bias. The students who get into serious trouble are usually the ones who don't advocate for themselves early enough.
The fairness angle is critically important but I also want to highlight a methodological issue with the detectors themselves. These tools are trained primarily on English text from native speakers, specifically from web scrapes that over-represent published content by native English speakers. Their training data doesn't adequately represent the diversity of how English is actually used globally by the roughly one billion people who speak it as a second language. It's a classic garbage-in-garbage-out problem. You train on native English writing, you build a model that thinks native English patterns are the only "human" patterns.
@hannah_wright your approach of requiring secondary review with the student present is exactly what more departments should adopt as standard policy. Using a single automated score as the sole basis for academic integrity decisions is indefensible, especially when the tools themselves acknowledge their limitations in their own documentation if you read the fine print.
I've been looking into Crossplag as a potential alternative for ESL contexts since it's developed in Europe and claims better multilingual support. In my limited testing, it was slightly less biased against ESL text than GPTZero, but my sample size was too small to draw strong conclusions and I'd want to see larger-scale data before recommending it for institutional adoption.
As an Italian writing academic English, I've had two false positives this year. Both times I had my Google Docs revision history stretching back to the first draft to prove the work was mine. My professors were understanding both times once I showed the evidence, but the process itself was stressful and embarrassing. Being called into a meeting to defend your own writing because an algorithm can't distinguish between a non-native speaker and a machine is not something any student should have to go through repeatedly.
I've been dealing with this issue all year as well. What helped me most was running my essays through GPTZero before submission so I could identify which specific sections were getting flagged and revise them preemptively. It feels ridiculous that I have to edit perfectly valid writing just to satisfy a biased algorithm, but at least it prevents the awkward confrontation with my professor.
One specific tip that actually works: read your essay out loud and revise any sentence that sounds too formulaic or textbook-like. Add some personality, throw in a rhetorical question, or use a slightly informal transition between paragraphs. The detectors seem to respond well to text that has more stylistic variation, even if that variation feels forced or unnatural to you. It's a hack, not a real solution, but it gets practical results while we wait for the tools to improve.
@luca_p totally feel you on the stress and embarrassment. The system is fundamentally broken when honest students are spending more time worrying about whether detection tools will believe them than about actually learning the material. We're here to get educated, not to prove our humanity to an algorithm.
Short answer: yes, absolutely. AI detectors measure statistical patterns in text, primarily perplexity (how predictable the word choices are) and burstiness (how much sentence length and complexity vary). The problem is that ESL writing often shares key characteristics with AI output: consistent grammar patterns learned from textbooks, limited vocabulary range because you default to words you know well, and simpler sentence structures that rely on common constructions.
It's not that your writing looks like AI wrote it in any meaningful sense. It's that the detectors are too statistically blunt to distinguish between "this person learned English as a second language and uses patterns from their coursework" and "a language model selected the most probable tokens." The underlying signal these tools measure is the same in both cases, even though the causes are completely different.
I've seen multiple studies confirming this. GPTZero's own team acknowledged the ESL bias issue back in 2024, but their mitigation efforts have been incremental at best. If you're an ESL student, you're fighting an uphill battle with these tools and it's genuinely unfair.