VestigeForensics

Learn

Is that AI image detector accurate? How to judge any vendor's claim

A checklist for evaluating any AI detection tool: the four disclosures a real accuracy claim makes, how to test one yourself in twenty minutes with photos you already have, and the red flags that settle it early.

Published August 6, 2026

You have found a tool. Its homepage says something like 99% accurate, the interface is clean, and you are about to rely on it for something that matters. The useful question is not whether the number is true, because it probably is true of something. The question is what it is true of, and whether that has any bearing on the image in front of you.

This is a method rather than a verdict on any particular product. We build one of these tools, which makes us interested parties, so the checklist is written so you can apply it to us as easily as to anyone else. If it is any good, it should occasionally embarrass us too.

The anatomy of an accuracy claim

"99% accurate"the number on the pricing pageOn whichimages?At whichthreshold?What false-positive rate?Whichgenerators?
A single percentage cannot answer any of the four, and the four are what decide whether the number applies to your image. A vendor that publishes all four is telling you something; a vendor that publishes only the headline is telling you nothing.

A single percentage is the answer to a question nobody stated. Four pieces of context decide whether it transfers to your situation, and a claim that omits them is not a measurement.

On which images?

Every accuracy number describes a specific test set. A detector evaluated on clean, full-resolution images from a handful of well-known generators will post a spectacular figure that says almost nothing about a screenshot forwarded through a messaging app. Ask what the test images were, how many, and above all whether they had been compressed, resized, or screenshotted. If the answer is only clean originals, the number describes a laboratory.

Ask also where the real photos came from. A detector tested against real images from one stock library has been tested against one photographic style. Modern phone photos, with their heavy computational processing, are a much harder case and a much more common one.

At which threshold?

A detector produces a score, and something has to turn that score into a yes or no. That cut-off is a policy choice, not a property of the model: move it one way for more catches and more false alarms, the other way for the reverse. Two vendors quoting different numbers may be running the same model at different cut-offs. A claim that never mentions a threshold has hidden the most consequential choice in the system.

What false-positive rate?

This is the disclosure that matters most and appears least. It is the share of genuine photographs the tool calls AI, and without it a detection rate is uninterpretable, because any detection rate can be reached by accepting enough false alarms.

It is also the error that hurts people. A false positive accuses a real image, and by extension whoever produced or supplied it. If a vendor publishes one number only, and that number is not the false-positive rate, you have learned something about their priorities.

Which generators?

Detectors perform best on generator families represented in their training and calibration data, and degrade against ones they have never seen. New models ship constantly, so any generator list has a shelf life. Ask which generators were tested, when the evaluation was run, and what the vendor claims happens outside that list. "It generalizes" is a hope. A measured number on a held-out generator is a result.

A fifth, once you are asking

Is the calibration pinned to a named, versioned run? Models get retrained. If a report from last month cannot be traced to a specific evaluation, its numbers cannot be reproduced or challenged, which matters enormously if the result ever has to be defended.

Test it yourself in twenty minutes

You do not have to take anyone's word, including ours. The most valuable test set for this purpose is one you already own: photographs you took yourself, whose origin you have no doubt about.

1. The false-alarm test. Upload twenty to thirty of your own recent phone photos, unedited, straight from the camera roll. Every flag is a false positive, because you know the provenance. Include night-mode shots, portrait-mode shots, and heavily processed ones, since those are the hardest real images and the ones most likely to be wrongly flagged.

2. The degradation test. Take one image and make three versions: the original, a screenshot of it, and a copy sent to yourself through a messaging app and downloaded again. Submit all three. A tool that gives confident and different verdicts across them, without ever mentioning that the file condition changed, is not accounting for condition at all.

3. The consistency test. Submit the same file twice, a few minutes apart. The answer should be identical. This sounds trivial and occasionally is not.

4. The known-generation test. Generate a handful of images with a current model, or use recent ones whose origin you are certain of, and see what happens. Prefer a recent model over a famous old one, because that is the realistic case.

Then read the results honestly, which means knowing what this cannot do. Twenty photos cannot measure a 5% error rate: with a sample that small the result is dominated by noise, and you could see zero flags or three from the same tool. What a small test can do is expose disqualifying behavior immediately. A tool that flags six of your thirty holiday photos has an error rate you can rule on without any statistics at all.

Red flags that settle it early

Applying this to us

Fair is fair. Our own detection rates, measured at thresholds calibrated so at most 5% of real photographs are flagged, sit in the range of roughly 55 to 64% depending on how degraded the image is, and lower again at the strict setting. The full accounting, including the conditions and the parts that do not flatter us, is in the measured answer, and the vocabulary used here is defined in the glossary.

Run the twenty-minute test on us. If your own photographs get flagged more than occasionally, we want to know: that is the failure mode we care most about, and the contact page reaches a person.

What this checklist cannot do

Frequently asked questions

How can I test an AI image detector myself?
Upload twenty to thirty of your own recent phone photos, which you know are real, and count how many get flagged. Then submit one image three ways, as an original, a screenshot, and a copy sent through a messaging app, and see whether the verdict changes without the tool acknowledging the change in file condition. Both tests use material you already have and neither requires trusting the vendor.
What should a trustworthy accuracy claim disclose?
Four things: which images it was measured on including how degraded they were, the threshold the verdict is issued at, the false-positive rate at that threshold, and which generators were tested. A fifth is worth asking for, namely whether the calibration is pinned to a named versioned run, so a result can still be reproduced after the model is retrained.
How many images do I need to test a detector on?
More than you can practically gather, if the goal is a precise error rate. A few dozen images cannot distinguish a 2% false-positive rate from an 8% one, because the result is dominated by chance at that sample size. A small test is still worth running, because it reliably exposes gross failure, such as a tool that flags a fifth of your ordinary photographs.
Should I trust a detector that tops a public leaderboard?
Only as far as the leaderboard's test set resembles your images. Rankings usually measure how well a model orders images from most to least suspicious, which is not the same as being correct at a specific threshold, and public sets often use clean high-resolution images from familiar generators. A tool can lead a leaderboard and still be unusable on a compressed screenshot.

Stop guessing. Run the analysis.

Upload the image and read the full forensic report: calibrated verdict, thresholds, limits, hashes and all. Free trial, no card required.