~/blog / ai-writing-detectors.md
AI writing detectors and false positives: what the tools get wrong
TL;DR: AI writing detectors are less accurate than they claim, and their mistakes land on real people: writers have lost assignments, gone unpaid, and been terminated over false positives. The tools hunt surface patterns that careful human writing shares with AI writing, so the more polished your prose, the more flaggable it gets. I’m Julie Kaiser, a science writer who edits AI-drafted work, and this post covers both halves: why the detectors fail, and what actually gives AI writing away.
Are AI writing detectors accurate?
Not reliably, no. AI writing detectors are less accurate than they claim, and many of the writers being flagged never touched AI.
The trouble starts with how the score gets read. The tool outputs a percentage. The client, editor, or teacher reads a verdict. Nothing in between asks whether the percentage means anything, and for careful human writing it often doesn’t.
The score survives because it’s convenient. Someone who can’t judge the writing themselves needs a cheap answer to a scary question, and a percentage feels like due diligence. That’s how a number nobody can explain ends up outranking a writer somebody hired.
I write about AI making things up about the world in the flagship post on AI hallucinations. Detectors are the trust problem running in the other direction: software making things up about you.
What happens when a detector gets it wrong
Real work, real money, real jobs.
In r/freelanceWriters, one writer went and interviewed colleagues about their experiences with detection tools. What came back wasn’t one bad week. It was a pattern: “Many of the writers I spoke with have faced dire consequences, including loss of assignments, unpaid invoices, and even termination from freelance positions after their work was evaluated by an AI checker.” (r/freelanceWriters)
Read the specifics again. Assignments lost. Invoices unpaid. Positions terminated. In each case the tool’s percentage counted as evidence and the writer’s word didn’t, and the writer was left trying to prove a negative to someone who’d already decided.
I have my own version, and it isn’t even a dramatic story. Some of my old work – for a client who was careful to make sure all their writers wrote consistently with each other – comes out flagged today. Written by hand, years ago, to a house style. The consistency that made it good client work is exactly the predictability a detector flags.
Why does AI detection fail?
Detectors fail because they hunt surface tells – vocabulary, punctuation, sentence patterns – and careful human writing hits those same tells.
Here’s the paradox underneath. AI was trained to imitate polished, formal writing. So when you write carefully, in a formal register, your prose resembles the thing AI is imitating, and the detector flags the resemblance. The tools are checking against a moving target while flagging the imitation-target as the fake.
That’s the whole mechanism. It’s not that these tools are badly built. It’s that the task – separating a good imitation from the thing it imitates, using surface features alone – may not be winnable. At least not this way.
What actually gives AI writing away
Start small. There are the surface tells everyone lists: certain words and punctuation most people don’t normally use – though stripping the em dashes stopped meaning much once everyone did it – and certain phrases and constructions that show up again and again.
One level up sits the best of the surface tells: burst variance. Human prose has it – a long sentence carrying nested clauses, then a short one, then a medium, then a very short. AI prose runs at one length: medium, medium, medium, medium. Even when it shifts rhythm on the surface, the total length stays inside a narrow band. Detectors don’t catch this; careful readers do, and once you notice the pattern you’ll see it everywhere.
These are all surface things, though. Look deeper and the real tell is the structure of the story underneath the words. A human writes by drilling down toward specifics: sentence two tells you something sentence one didn’t, and the third paragraph knows more than the first. Unedited AI writing circles – the repetitive, context-free wall of text, the same idea restated in slightly different words, no new information per sentence. Put the two side by side and they look structurally different, whatever the vocabulary says.
I notice these patterns because I edit AI-drafted work for clients – catching them is part of the writing work on my Work with me page. And an honest caveat: a careful read beats a percentage score, but even a careful read can be wrong. I can’t prove a text’s origin from its surface. None of us can.
How do I prove I wrote something, not AI?
You prove it with process, not with a counter-score: the trail your writing leaves behind while you make it.
- Keep version history on. The revision log in Google Docs or Word shows a document growing over hours, with wrong turns – the one thing a paste-in doesn’t have.
- Keep your working materials. Outlines, research notes, drafts with dates. Boring, and exactly what a fabricated authorship story lacks.
- If you’re flagged, ask two questions: which passages, and what’s the tool’s documented false-positive rate? People treating a score as proof have usually never looked up either.
- Run the same detector over something you wrote years before these tools existed, and bring the result to the conversation.
One thing I wouldn’t count on: rewriting to dodge the score. Lots of people try – rewriting the flagged passages, or running one of the skills and tools that promise to “humanize” AI text. I’ve tried them too, without any good luck. Text that starts out AI-written has an underlying sameness to it, and the surface swap-outs those tools make either don’t work or work at the expense of correctness. Writing worse to satisfy a broken tool gets the priority backwards.
The bigger frame
Both failure modes in this pillar are the same mistake: treating fluent, confident output as evidence. AI invents facts about the world – citations have their own version of this problem, covered in Is ChatGPT accurate? – and detectors invent facts about writers. Different victims, same posture error.
What’s the fix? A percentage is a claim, not a verdict, and a citation is a claim, not a source. To resolve either one, somebody has to actually read the thing – and the detector isn’t a somebody. Careful reading is still the job.
New here? I’m Julie – the homepage is the two-minute version of who I am and what this is about. Came with one specific worry, like an AI that forgets you or lies to you? The blog page is sorted by exactly those questions – start at yours.
The newsletter
If this was useful
The newsletter is where I send what I learn next – a letter every week or so on what I built, what broke, and what I’d tell you to try. No hype, ever.
Double opt-in · unsubscribe anytime · GDPR-compliant

I’m a scientist by training and a science writer by profession: chemistry and biology, 14 years at the lab bench, 8 peer-reviewed papers, and regulated biotech and pharma clients since 2011 – work where being wrong has consequences. For the last three years I’ve used AI on that real work, and here I document what actually happened: what worked, what broke, and what I’d tell you to try next. My best tip: if I can do it, you can do it.
