VestigeForensics

Learn

How accurate are AI image detectors? The measured answer

Vendors say 99%. Our benchmark measures 55 to 64% detection at a 5% false-positive cap, and 36 to 45% at the strict setting. Where the gap comes from.

Published July 21, 2026 · Updated August 6, 2026

Ask a vendor and you'll hear a number like "99% accurate." Ask a researcher and you'll get a question back: on which images, at which threshold, against which generators, after how much compression? The honest answer is that an AI image detector doesn't have an accuracy; it has a performance curve, and where you sit on that curve depends on choices someone made and should be willing to disclose. We run a detection benchmark ourselves, so this article uses our own published numbers, including the unflattering ones, to show what those choices do.

Why "99% accurate" is usually a non-statement

A single accuracy percentage silently blends two very different errors:

You can trade one for the other freely by moving the decision threshold, which means a marketing number can be true and useless at the same time. Worse, "accuracy" depends on the mix of the test set. A detector tested on clean, full-resolution images from the same generators it was trained on will post a spectacular score that says almost nothing about the screenshot forwarded to you this morning.

The two errors also matter differently. If you check a thousand ordinary photos, of which a handful are AI, even a 5% false-positive rate produces far more false alarms than true catches. That's why the first question for any detector isn't "how many AI images does it catch?" but "how often does it cry wolf on real photos, and on what kind of real photos was that measured?"

What we measure, and what it shows

Our benchmark run (full_v4) scores 8,064 images across 8 conditions (387,072 scored evaluations), because real images don't arrive clean. They arrive re-compressed by WhatsApp, screenshotted, downscaled, and re-shared. The hardest condition in the grid is a screenshot of a WhatsApp image, which is what actually lands in front of an analyst.

Three of our own numbers, side by side:

Ranking quality (AUC), hardest condition0.933sounds like a marketing number, is not a detection rateDetection at ≤5% false positives55–64%the honest operating number, range across conditionsDetection at ≤1% false positives (evidentiary)36–45%the price of alarms you can defend
One detector, one benchmark run (full_v4). The gap between the first bar and the other two is the honesty gap: ranking quality is real, but decisions run at a threshold, and a threshold that caps false alarms buys fewer catches.

That distance between the impressive-sounding AUC and the operating detection rate is what we call the honesty gap. Every detector has it. Most quote the top number and stay quiet about the bottom two; we print the bottom two on the report itself.

For contrast: a well-known public detector run at its default 0.5 threshold flagged 27% of genuine photos as AI under a messenger re-encode in our benchmark: it reads compression artifacts as evidence of synthesis. Nothing is wrong with that model's science; everything is wrong with using it un-calibrated.

What moves the numbers

Image quality

Compression, resizing, and screenshots destroy exactly the high-frequency statistical traces detectors read. A detector calibrated only on clean images degrades unpredictably on real-world ones, so the calibration has to be done per condition, and you should ask whether it was. We measured how much this costs in the compression article.

Which generator made the image

Detectors are best on generators represented in their training and calibration data. Against a brand-new or unseen generator, detection rates drop, ours included. A detector that won't say this is not being straight with you.

The threshold policy

The same model can be an aggressive tipster (more catches, more false alarms) or a conservative witness (fewer catches, alarms you can defend). Neither is "the accuracy"; they're operating points, and the choice should be yours, stated on the output. Ours issues verdicts at two: standard (≤5% false positives) and strict (≤1%).

Questions to ask any detector, including ours

  1. What is the false-positive rate of a verdict, and on what real images was it measured? ("We don't publish that" is an answer too, just a bad one.)
  2. Was performance measured on compressed, resized, screenshotted images, or only clean ones?
  3. Which generators were tested, and what happens on ones outside that list?
  4. What does a negative result mean? (Correct answer: not much. A calibrated detector deliberately lets borderline images pass rather than inflate false alarms. "Not flagged" is never a certificate of authenticity.)
  5. Is the calibration pinned to a named, versioned run, so the number on yesterday's report still means something after the model updates?

If you take one thing away: an accuracy claim without a false-positive rate, a test-condition description, and a generator list isn't a measurement; it's a slogan. The tell of a trustworthy detector isn't a high number; it's the willingness to show you the low ones.

Faced with a specific tool and a specific claim, there is a checklist for working through exactly this, including a twenty-minute test you can run yourself with photos you already have: is that detector accurate?

Frequently asked questions

What is the most accurate AI image detector?
The question has no stable answer, because every published accuracy number depends on the test set, the threshold, and how degraded the images were. A detector that tops one benchmark can trail on another. The useful comparison is not a single score but disclosure: measured false-positive rates, per-condition results, and a named calibration run you could check.
Are free AI image detectors accurate?
Some use respectable models, but almost none state a false-positive rate or calibrate their threshold for compressed real-world images. In our benchmark, one well-known public detector at its default threshold flagged 27% of genuine photos after a messenger re-encode. Without a stated error rate, a free score is a guess you cannot weigh.
What does an AI detection score like 0.7 actually mean?
On its own, almost nothing. A raw score only becomes a verdict at a threshold, and the right threshold depends on the image's condition and the false-positive rate you can tolerate. A calibrated system converts the score into a decision with a stated error rate; an uncalibrated 0.7 is just a number.
Can any detector really be 99% accurate?
On clean images from generators it was trained on, yes, briefly. Under realistic conditions, with false positives capped so alarms stay meaningful, our own measured detection runs at 55 to 64%, and 36 to 45% at the strict evidentiary setting. Any 99% claim without conditions attached is marketing, including if we ever made one.

Stop guessing. Run the analysis.

Upload the image and read the full forensic report: calibrated verdict, thresholds, limits, hashes and all. Free trial, no card required.