Check before you share
Media Literacy Guide

Do AI Detectors Work? An Honest Look at Accuracy, False Positives, and Real-World Limits

An honest assessment of AI detector accuracy, explaining why high false positive rates—especially for non-native English writers—make them unreliable as sole evidence.

The Sales Pitch vs. The Stress Test

Ask the blunt question "do AI detectors work accuracy false positives" and you will get a confident yes from a vendor and a long pause from a researcher. The gap between those two answers is the whole story. Detector companies market 99% accuracy the way a hotel markets a sea view: true from one angle, meaningless from the one that matters. Independent testing keeps finding the same thing. A tool that flags a piece of text as machine-written is not reporting a fact. It is reporting a probability from a model that has never seen your writing before. That distinction is where the false positive problem lives.

Do AI detectors work?
Epoch AI , CC BY 4.0 via Wikimedia Commons

What the Vendors Claim vs. What Tests Show

The Numbers You Hear vs. The Numbers That Matter

AI verification tool reliability collapses the moment you put two numbers side by side. GPTZero's own developer claimed 99% accuracy on human text and 85% on AI text back in 2023. The Stanford University study from July of that year tested seven systems, including GPTZero, and found something else entirely: they correctly identified only 5.1% of AI-generated text as human, but they flagged 61.3% of non-native English writing as AI-generated. The vendor's number describes its best case. The researcher's number describes a classroom full of students whose first language is not English.

Why Paraphrasing Makes a Mockery of the Score

Accuracy for AI-generated text sits somewhere between 50% and 80% depending on the tool and the content type. The worst of them are little better than a coin flip. The failure gets worse when you introduce a paraphraser. A University of Maryland study in 2023 found that simply rephrasing AI-generated text dropped detection rates from around 95% down to about 50% across every tool they tested. A system that catches a lazy copy-paste job misses the same text once a human has run it through a thesaurus. The person using the tool never sees any of this. They see a red score and a percentage, and they make a decision on it.

Why the Stanford Numbers Matter for Real People

The Stanford finding is not an academic footnote. It is the difference between a student being accused of cheating and being cleared. Non-native English writers get caught in the net because their sentence structures are more regular, more predictable, more like the patterns a language model produces. The tool is not biased because it is malicious. It is biased because it was trained on data that looks like formal written English, and a competent non-native writer produces something close to that. The over-trust problem appears the moment a teacher uploads an essay, sees an 80% AI score, and stops reading. The screening tool just became the judge, jury, and executioner, and it never told anyone it was wrong twenty times out of a hundred.

Turnitin, the plagiarism checker used by most universities, publishes a false positive rate of about 1% at the sentence level. That sounds reassuring until you consider that a single essay contains dozens of sentences. The tool is not flagging the whole document; it is flagging fragments. A student accused on the basis of one sentence has to fight a battle they did not create. Turnitin itself admits the false positive rate is higher for non-native English writers, though it has never published the exact figure. The number that would settle the argument is the one they will not release.

What Happens When the Tool Fails

Independent testing of image screening tools tells the same story from a different angle. Hive Moderation, one of the more established names, scored around 85% to 90% accuracy on Midjourney v6 images in a 2024 independent test. That sounds strong until you flip it: one image in ten is wrong. For a scanner checking a news feed, a 10% error rate means a real photograph gets labelled fake, and a synthetic image gets labelled real. Neither outcome helps the person trying to decide what to believe. The tools that exist for spotting AI-generated video are even less reliable outside a laboratory. A deepfake that has been compressed, resized, and re-recorded off a screen loses the artefacts the system was trained to find.

The exact numbers shift depending on when the test was run and which model generation was used. What does not shift is the direction of the result. Every independent test finds the same gap between the claimed accuracy and the real-world performance. The tools are useful as a signal. They are dangerous as a verdict.

  • Claimed accuracy (typical vendor claim): 99% on human text, 85% on AI text
  • Real-world detection range (AI text): 50-80% depending on tool and content type
  • Non-native English writers misclassified as AI (Stanford, 2023): 61.3%
  • Native English writers misclassified as AI (Stanford, 2023): 5.1%
  • Detection drop after paraphrasing (Maryland, 2023): Roughly 95% down to about 50%
  • Turnitin sentence-level false positive rate: Approximately 1%, higher for non-native writers

What You Should Do Instead of Trusting a Single Score

Stop Outsourcing Your Judgement

AI screening tools fail not because the technology is useless but because it is being used for the wrong job. No tool is reliable enough for standalone verification. The ideal use is as one signal among several. A detector might suggest something is worth a closer look. It should never be the only thing you look at. The person who uploads a document to a single tool and accepts the result is committing the exact failure mode that researchers identify as verification bypass. They are outsourcing a judgement that requires context to a system that has none.

Learn Lateral Reading Instead

Lateral reading is the more durable skill. When you encounter a claim, open a new tab and check the source before you believe it. Who published it? Do they have a stake in the outcome? What do other sources say? Factually, the Singapore government's fact-checking portal, publishes clarifications on false claims circulating locally, and Black Dot Research, an independent organisation, does the same work. Neither one needs an AI detector to do their jobs. They read, they check, they publish their reasoning. That is a methodology disclosure you can evaluate. A detector gives you a percentage with no reasoning behind it, which means you can never audit its work.

Why the Tool Over-Trust Failure Is So Dangerous

Tool over-trust happens because the interface is seductive. A red score with a percentage feels like a fact. It looks like a measurement, like a thermometer reading or a speedometer. But a thermometer measures temperature directly. A detector measures probability across a statistical model, and that probability is wrong often enough to ruin a life. The student falsely accused of cheating will remember the humiliation long after the tool vendor has updated their model and quietly improved their numbers. The person who forwarded a deepfake video of a politician and got caught by a false positive will simply stop trusting verification altogether.

Uploading an image or text to a single tool and treating the result as definitive, without understanding the tool's error rate, is not an edge case. That is the default behaviour of most users. The people who build the tools know this. The responsible ones publish their limitations. The irresponsible ones bury them in a footnote on a page nobody reads.

What the Detectors Get Right

None of this means the tools are worthless. A detector is excellent at catching the lazy case: a student who pastes a full paragraph from a chatbot into an essay without reading it. It is also good at triage, at sorting content into plausible and suspicious, so a human reviewer knows where to spend their attention. The problem is not the tools. The problem is the leap from a probability to a conclusion. A tool with a 90% true positive rate is still wrong 10% of the time. For a system used in university discipline hearings, a 10% error rate means a meaningful number of innocent students are being punished. Those are not abstract statistics. Those are people with names.

The ScamShield app used in Singapore does not include AI content detection. The Infocomm Media Development Authority's AI Verify toolkit is designed for governance testing, not output detection. No widely validated consumer-grade tool for detecting AI-generated voice clones exists as of 2026. The tools that exist are narrow, specialised, and fallible. The ones that do not exist are more important.

The Evidence Trail That Actually Works

The alternative to a detector score is an evidence trail. C2PA content credentials are a standard that lets a creator cryptographically sign their work with provenance information. A photo taken on a modern smartphone can carry data about when and where it was taken, which camera captured it, and whether it has been edited. Google's SynthID watermarks AI-generated text using a technique that embeds a pattern invisible to the eye but detectable by software. Neither system is perfect. But both have a property that a detector lacks: they are attached to the content at the moment of creation, rather than guessed at after the fact.

The problem with a detector is that it is trying to infer a cause from an effect. A watermark is trying to establish a chain of custody. The first approach is probabilistic; the second is cryptographic. The difference matters because a probability can be wrong, while a signature either verifies or it does not. For a reader in Singapore who wants to know whether a forwarded message is real, the question should be: where did this come from, and can I check the source? A detector cannot answer that. Lateral reading can.

The Cost of the Wrong Tool for the Job

When a detector gives a false positive, the cost is not just the error. It is the trust that is destroyed when the error comes to light. A student cleared of false cheating accusations rarely feels vindicated; they feel the system is broken. A reader who shares a real photograph that a tool flagged as fake looks foolish. The tool was supposed to make the world more honest, and it has made it less so, because it replaced a human judgement call with a bad number.

A fast correction that nobody sees has limited real-world effect. A verdict without visible evidence is unverifiable by the reader. When the detector is wrong, the correction does not spread with the same velocity as the original accusation. The person who was right to be sceptical of the AI content is now sceptical of the verification, and the whole exercise collapses into a mess of competing distrust.

Where This Leaves the Person Actually Trying to Verify

If you are a teacher, a journalist, or a concerned citizen in Singapore trying to figure out whether a piece of content is synthetic, your first move should not be to open a detector. Your first move should be to check the source. Who is telling you this? What do they have to gain? Is this a legitimate news outlet, a government fact-checking portal, or a random account with a following? The International Fact-Checking Network requires its members to publish their methodology and funding. That is the disclosure you should be looking for, not a probability score from a black box.

For images, run a reverse image search. Find out where the image has appeared before. A photo of a burning building that went viral last year will show up in the results. For text, look for the specific claims being made and search for those claims directly. If the claim is real, a credible source will have reported on it. If it is false, a fact-checker will have documented the correction. This takes longer than uploading a file and reading a score. But it works, and it does not depend on the whims of a statistical model that changes with every training run.

What the Research Says About the Future

Watermarking is the most promising direction for AI text detection because it does not rely on statistical guesswork. The University of Maryland researchers who proposed the technique in 2023 also showed that paraphrasing destroys it, which is why it has not yet become the silver bullet that some hoped. SynthID is open-source and integrated into Google's Responsible GenAI Toolkit, but it only works if the generator cooperates. A person using an open-source model that does not embed the watermark can produce text that is invisible to the system.

The C2PA standard is more robust, because it is attached to the content rather than the output. A camera that signs its photos at capture time creates a chain of custody that cannot be broken by a paraphraser. The standard is at version 2.1 as of 2024, and the coalition that maintains it includes most of the major tech companies. But it only helps with new content. It does nothing for the vast archive of images and text already online, which is where most of the misinformation lives.

The Real-World Test: Singapore's Approach

Singapore's approach to AI-generated misinformation is instructive because it did not reach for a detector. The POFMA Office has not endorsed any AI detection tool as of 2026, and its fact-checking partners do not rely on one. The Infocomm Media Development Authority's AI Verify toolkit is about testing whether an AI system is fair and transparent, not about detecting whether a given output is synthetic. The emphasis is on source verification and correction, not on catching the generator.

A Singapore public consultation on AI-generated misinformation closed in January 2024, and the published documents do not endorse any detector. The consensus among fact-checkers in 2024 is that no tool is reliable enough for standalone verification. That is not a Luddite position. That is the conclusion of people who spent the year watching the tools fail at speed and scale. The tools are not useless. They are just not the answer to the question everyone wants them to answer.

The Bottom Line for Anyone Facing a Suspicious Message

The next time you receive a forwarded message on WhatsApp or Telegram, resist the urge to run it through a detector. The message has already been through at least one hop, and by the time it reaches you, it has been stripped of the metadata that might have helped. Instead, treat it the way a journalist treats a tip from an anonymous source: with scepticism, a check of the source, and a search for independent confirmation.

The Source, Understand, Research, Evaluate framework embedded in Singapore's school curriculum and public libraries is a better guide than any detector. It teaches you to question where information comes from, to understand the difference between a claim and a fact, to research the context, and to evaluate the credibility of the source. None of that requires a probability score. It requires the one thing a detector cannot give you: judgement.

Who This Subject Is For

Understanding the limits of AI detectors is for anyone who has ever been asked to verify something, which is everyone. It is for the teacher grading essays, the journalist checking a tip, the parent seeing a scary video of a politician, the professional forwarding a message to a group chat. It is for the person who wants to be right, not just to appear right.

It is not for people who want a single button that tells them the truth with certainty. No such button exists, and anyone who tells you otherwise is selling something. If you are willing to spend a few minutes learning how to verify a source, how to run a reverse image search, and how to read a fact-checking organisation's methodology for bias, you will be better at this than almost anyone sharing unverified content. The tools are not the answer. You are the answer. The detectors are just a hint.

  • OpenAI AI Text Classifier: Discontinued July 2023; at discontinuation, correctly identified 26% of AI-written text
  • Stanford study (July 2023): 7 detectors tested; 61.3% of non-native writing falsely flagged vs. 5.1% of native writing
  • Paraphrasing evasion (Maryland, 2023): Detection rates dropped from ~95% to ~50% after paraphrasing
  • Watermarking status: SynthID open-source (2024); C2PA standard v2.1 (2024)
  • Singapore's position: No official endorsement of AI detectors by POFMA or IMDA as of 2026

The Practical Checklist for Verification Without a Detector

When you are staring at something suspicious, work down this list. First, who published it? If the answer is a stranger on Telegram, that is not a source. Second, what is the claim exactly? Break it into parts. The simpler the claim, the easier it is to check. Third, search for the claim directly, not for the sentence you were given. Type key phrases into a search engine and see what comes back. Fourth, look for a fact-check. Factually, the Singapore government's fact-checking portal, publishes clarifications on false claims circulating locally, and Black Dot Research, an independent organisation, does the same. If neither has covered it, that does not mean it is true; it means it has not been checked yet.

Fifth, check the date. Old news resurfaces, and a story from 2019 may be circulating as if it happened yesterday. Sixth, run a reverse image search on any visual. A photo of a crowded MRT train might be from a different city entirely. Finally, be willing to say you do not know. Not knowing is not a failure. Pretending to know is how misinformation spreads.

Why This Matters More Than Getting the Perfect Tool

The conversation about AI detectors has been framed as a technology arms race: the tools get better, the fakes get better, repeat. That framing is wrong. The issue is not that the detectors are not good enough. It is that they are being used as a substitute for thinking. A detector cannot tell you whether a claim is true, only whether it is likely machine-written. A false claim written by a human is still false. A true claim written by an AI is still true. The provenance of the text tells you nothing about its accuracy, and that is the point the vendors do not want you to notice.

The claimed-vs-real gap exists because the tools are solving a different problem from the one their customers think they are solving. The customers want to know if something is true. The tools tell them if something looks like a machine wrote it. Those are not the same question, and confusing them is how a student gets expelled for an essay they wrote.

The One Skill That Beats Any Detector

Lateral reading is not a hack or a trick. It is the behaviour of someone who does not believe the first thing they see. It is opening new tabs to check the source, the claim, and the context before you decide. It is the difference between a person who forwards a fact-check and a person who checks the fact-check's own sources. The National Library Board's framework, Source, Understand, Research, Evaluate, is the formal name for this skill, and it is embedded in Singapore's schools and public libraries for a reason. It works because it teaches you to think, not to defer.

A detector asks you to defer. It asks you to hand over your judgement to a statistical model and accept the output. Lateral reading asks you to keep the judgement and use the tools. A detector can be a useful input. It should never be the output of your verification process. The person who understands that is the person who will not be fooled, or at least who will be fooled less often, and less badly.

The Final Word on Detector Accuracy

You now have the numbers, the studies, and the failure modes. The question was never really whether the detectors work, because the answer depends on what you mean by work. If work means catching a lazy copy-paste job, then yes, they are fine. If work means reliably distinguishing between human and machine writing, then no, not one of them is up to the task, and the vendors who claim otherwise are lying or deluded. The Stanford study alone is enough to settle the argument: 61.3% of non-native English writing was falsely flagged as AI-generated. That is not an acceptable error rate for a tool used in academic discipline, employment screening, or public verification.

The tools will get better, and then they will get worse, and then the fakes will improve, and the cycle will continue. The constant is the skill of verification. Not the tool, but the skill. And now you have it.