4 min readBy PriceSnap Editorial TeamPublished

AI Price Scanner Accuracy Test: Our 60-Item Benchmark Protocol

A useful accuracy study needs more than a few impressive screenshots. This preregistered protocol explains how PriceSnap will compare six photo-scanner and manual-research workflows across 60 used items without changing the rules after seeing the results.

AI summary

PriceSnap has registered a 60-item benchmark for six valuation workflows. It will score exact identification, usable ranges, median percentage error, reference-price coverage, time to result, source transparency, and uncertainty disclosure. Results are not yet published, so this page makes no winner or accuracy claim.

Mixed secondhand items arranged for a repeatable price-scanner comparison

PriceSnap is a mobile app for iOS and Android.

Use the app while reading this guide to scan items, estimate resale value, check marketplace comp signals, and save finds to your collection.

Key takeaways

  • The test uses 60 items across six categories and the same inputs for every eligible workflow.
  • Reference values require closely matched completed sales; weakly matched items are excluded rather than forced into the score.
  • Identification, range usefulness, error, speed, transparency, and uncertainty are measured separately.
  • No app—including PriceSnap—will be named the winner until the dataset and calculations can be reviewed.

Try alongside this guide — scan straight from your camera roll.

Why We Are Publishing the Rules First

PriceSnap appears in this comparison and publishes the page, which creates an obvious conflict of interest that no amount of careful wording removes. Publishing the protocol before data collection is the practical answer: it makes the sample size, category mix, ground-truth method, metrics, and exclusions visible in advance, so that the rules cannot quietly change once the results are known. The final report must show unfavourable results as clearly as favourable ones, identify who reviewed the calculations, and preserve a correction log. Until that report exists, this page is a research design and nothing more — it is not evidence that any scanner, ours included, is the most accurate.

  • The conflict of interest is direct and stated.
  • Rules are fixed before data collection, so they cannot be adjusted afterwards.
  • Unfavourable results will be reported as prominently as favourable ones.
  • A correction log will be maintained.
  • This page is a protocol, not a result.

Read next:About and editorial standards

The 60-Item Test Set

The sample contains ten clothing and shoe items, ten electronics, ten cards and collectibles, ten furniture and home goods, ten antiques and vintage goods, and ten jewellery, watch, or accessory items. Each group should include common, niche, low-value, high-value, and visually difficult examples, so that the set reflects real sourcing rather than a convenient subset. Every item needs a documented identity and condition recorded before any scanning takes place. Items selected because one app already performs well on them are not permitted, and any substitution must be recorded before scoring is calculated — a substitution made after seeing results is the most common way a benchmark becomes marketing.

Sixty items, six categories, ten each.
CategoryCount
Clothing and shoes10
Electronics10
Cards and collectibles10
Furniture and home goods10
Antiques and vintage10
Jewellery, watches, accessories10

The Six Workflows and Same-Input Rule

The registered comparison covers PriceSnap, RePick, Revalue, ThriftAI, one additional scanner verified at test time, and a manual eBay search workflow as a control. App availability, versions, platforms, subscription state, country, currency, and test date will all be recorded, since any of them can affect results and none is stable over time. Each photo-based workflow receives the same permitted photographs and the same condition notes. If a product requires a materially different input, that difference is documented rather than hidden, because an undisclosed input advantage invalidates the comparison entirely. Manual search time begins before the first query and ends when a defensible range has been recorded, not when the first result appears.

  • Five photo workflows plus a manual search control.
  • Identical permitted photographs and condition notes.
  • Any required input difference is documented, never hidden.
  • Versions, platforms, subscription state, region, and date recorded.
  • Manual timing runs to a defensible range, not to first result.

How Reference Values Are Built

The evaluator first confirms the exact brand, model, edition, size, grade, accessories, and condition, because a reference built on a wrong identity measures nothing. The target reference is the median of five to ten recent, closely matched completed sales in the relevant region. Active asking prices are not treated as completed sales at any point. Obvious mismatches and documented outliers may be excluded with a stated reason, which is retained in the record. When fewer than three reliable completed matches remain, the item is marked unscorable — the study does not invent a precise ground truth for a market that does not have one, since doing so would measure the evaluator's guesses rather than the tools.

Ground-truth rules, fixed in advance.
RuleStandard
Match basisExact brand, model, edition, size, grade, accessories, condition
Sample sizeFive to ten recent closely matched completed sales
Reference valueMedian of retained sales
Asking pricesNever counted as completed sales
ExclusionsPermitted only with a recorded reason
Thin marketsFewer than three matches means unscorable

The Seven Reported Metrics

The final study reports exact-identification rate, usable-range rate, median absolute percentage error, the percentage of references contained inside the returned range, median time to result, a source-transparency score from zero to five, and an uncertainty-disclosure score from zero to five. Results are shown overall and broken out by category, because a tool can be strong in one category and weak in another and an overall figure conceals that. Medians are preferred over means for error and speed, since a few extreme items would otherwise distort the picture disproportionately. Sample counts accompany every percentage, so a reader can immediately see when a result rests on a small base and weight it accordingly.

The seven metrics, reported overall and by category.
MetricWhat it captures
Exact-identification rateHow often the precise item was identified
Usable-range rateHow often the output was actionable at all
Median absolute percentage errorTypical distance from the reference
Reference containmentHow often the reference fell inside the range
Median time to resultRealistic speed, outliers not dominating
Source transparency (0-5)Whether the evidence can be inspected
Uncertainty disclosure (0-5)Whether weak evidence is admitted

Exclusions, Missing Results, and Fairness

A failed identification remains a failure and is not removed because it damages a score — this is the single most important fairness rule, since selective removal of hard cases is how most informal comparisons mislead. An item may be excluded from price-error calculations only when the reference market is genuinely insufficient, when ownership or identity cannot be documented, or when the input was corrupted. Every exclusion retains its item ID and its reason in the published record. Paid and free states are recorded separately where practical, since they may not perform identically. The final article will include at least two PriceSnap limitations and will name competitor use cases that the results support.

  • Failed identifications are scored as failures, never removed.
  • Exclusion permitted only for thin market, undocumented identity, or corrupted input.
  • Every exclusion keeps its item ID and reason in the public record.
  • Paid and free states recorded separately where practical.
  • At least two PriceSnap limitations named in the final article.
  • Competitor use cases named where the results support them.

What the Final Publication Must Include

The finished report should include a downloadable CSV, a data dictionary, the scoring formulas, item-level results, the permitted photographs, app versions, test dates, exclusions, and reviewer notes — enough for a sceptical reader to recompute the headline numbers themselves. It should distinguish a directional resale estimate from a certified appraisal, and avoid implying that reference values guarantee any future sale. The benchmark will be rerun every six months, since model updates and market conditions both move, and app availability, pricing, and feature facts should be rechecked at least every ninety days. A benchmark that is not maintained becomes misinformation on a delay.

  • Downloadable CSV plus a data dictionary.
  • Scoring formulas and item-level results.
  • Permitted photographs, app versions, and test dates.
  • Full exclusion list with reasons, and reviewer notes.
  • Rerun every six months; product facts rechecked every ninety days.

Read next:How PriceSnap estimates resale value

Related categories

Continue your research

FAQ

AI Price Scanner Accuracy Test: Our 60-Item Benchmark Protocol — FAQ

Straight answers about accuracy, platforms, and how PriceSnap fits your workflow.

Has the 60-item accuracy test been completed?

Not yet. This page publishes the protocol before testing so the sample, metrics, and exclusion rules cannot be quietly changed to favor PriceSnap.

What counts as the reference value?

The target is the median of five to ten closely matched completed sales. An item is unscorable when fewer than three reliable matches remain after identity and condition checks.

Will PriceSnap be included in its own test?

Yes. PriceSnap publishes the study and is one of the workflows being evaluated, so that conflict is disclosed and the raw evidence must be available for review.

Why not measure only average price error?

A scanner can fail before valuation by identifying the wrong item or hiding weak evidence. The study therefore measures identification, usable output, transparency, uncertainty, and speed as well as price error.

← All guides