How the Trust Score works

Scoring version v2 · Last updated · Thresholds on this page are the current published configuration and are re-calibrated against hand-rated products.

A Trust Score rates the reviews on a listing — not the product, not the seller, not whether you should buy it. This page sets out exactly what goes into the number, how each of the five dimensions is computed, where the data comes from, and the places where the method is known to fail.

  • The Trust Score is a 0–100 score with an A–E grade describing how trustworthy a listing's review set looks.
  • It combines three things: the star rating after suspected fake reviews are removed, the share of reviews that failed our authenticity checks, and five review-quality dimensions.
  • Under 20 reviews we publish "Not Enough Evidence" instead of a grade. Under 5, no score at all. Scores from small samples are pulled toward the middle, so a listing with 6 reviews cannot earn an A or an E.
  • Every score is a snapshot of the reviews visible on the date shown on the page. Reviews the platform has already removed are invisible to us.
  • A high score does not mean the product is safe, the seller is reliable, the item is authentic, or that it is right for you.

1. What the Trust Score measures — and what it doesn't

The single most common misreading of this site is treating a Trust Score as a product rating. It isn't one. The score answers one narrow question: if you read this listing's reviews to decide whether to buy, how much can you rely on what you read?

In scope — properties of the review set

  • Whether star ratings and review text say the same thing
  • Whether reviews contain first-hand, checkable detail
  • Whether discussion covers the product broadly or circles one talking point
  • Whether the star distribution has the shape real products produce
  • Whether the writing reads like an owner or like a product page
  • What share of reviews match known manipulation patterns

Out of scope — we do not measure these

  • Product quality, durability, safety, or recall status
  • Whether the item is authentic or counterfeit
  • Seller conduct: shipping, returns, warranty, responsiveness
  • Price, value for money, or whether a deal is real
  • Whether any individual reviewer was paid or refunded
  • Whether this product suits your budget, size, or use case

How the parts combine

Three independent inputs feed the score, and each is computed by ordinary code from per-review annotations — no model is asked to produce the final number.

The first is the adjusted rating: the listing's star average recalculated with the reviews we labelled as fake taken out, and with small-sample shrinkage applied so that removing a handful of reviews cannot swing the average on its own. It is the input closest to the number the platform already shows you.

The second is authenticity — the share of classified reviews that did not fail our checks, entering as a penalty rather than a bonus. Reviews we could not confidently classify either way are excluded from both sides of that fraction rather than counted as clean, so an ambiguous review never earns a listing credit.

The third is the five dimensions in section 2, blended into a single review-quality figure that acts as a correction on the first two. How much each input contributes is calibrated against hand-rated products and pinned to the scoring version rather than chosen by intuition.

One step then sits on top of all three. Confidence — driven by how many reviews we have and how many of them we could confidently classify — pulls the result toward a neutral baseline when the evidence is thin. A listing with six reviews therefore lands near the middle of the range no matter how flattering or how damning those six reviews are. Certainty has to be earned by sample size, and this is the step most scoring systems skip.

Letter grades are percentiles, not fixed cutoffs

Grades are cut at quantiles of the score distribution across everything we have recently analysed, not at round numbers someone chose. A grade is a statement about rank: how this listing's reviews compare with the rest of the listings we've looked at.

Target grade distribution across the analysed corpus
GradeShare of analysed listingsReads as
Atop 15%Review set looks clean on every dimension we measure
Bnext 30%Normal; minor weaknesses, nothing that changes the picture
Cnext 30%Mixed signals; read the dimension detail before relying on the reviews
Dnext 18%Multiple dimensions are off; treat the star average as unreliable
Ebottom 7%The review set matches manipulation patterns on several fronts

Cutoffs are re-derived from a pool of at least 500 recently scored listings, checked against a hand-rated sample, and pinned to a scoring version. If the boundaries move, the version number moves with them, and every product page states the version its score was computed under. Two consequences worth knowing: a grade is relative to the corpus, and a listing sitting one point from a boundary could reasonably have been placed either side — treat adjacent grades as near-equivalent.

Below 20 reviews, no letter grade is issued at all. The page shows Not Enough Evidence and the underlying figures instead. Four five-star reviews on a low-priced accessory is one of the most common shapes manipulation takes; reporting that as an A would be the one failure a review detector cannot survive. Not Enough Evidence is a result, not an error.

2. The five dimensions

Each dimension is a statistic computed over the listing's reviews and mapped onto a −1 to +1 scale, where −1 is the pattern we associate with manipulated review sets and +1 the pattern healthy ones show. A language model annotates individual reviews — sentiment, evidence types, topics mentioned, tone; it never computes a rate, an average or a score. All arithmetic is deterministic code, which is why the same reviews produce the same numbers on re-run. How the five are blended is calibrated against hand-rated products and versioned with the score; we publish the thresholds each dimension turns on, not the blend, because the blend is the part worth gaming.

RTA Rating–Text Agreement

Does the star rating match what the review actually says?

Real reviewers rate what they write. A five-star rating attached to a description of the product failing after two days means the rating and the text were produced separately.

Computed
Text sentiment is annotated on a −1 to +1 scale and the star rating placed on the same scale with (star − 3) / 2, so 1★ → −1 and 5★ → +1. The dimension is driven by the mean absolute gap between the two. Perfect agreement scores +1, an average gap of half a star scores 0, and a full star or more scores −1.
Where it fails
Sarcasm, very short text, and reviews that list problems before concluding the writer is happy anyway all read as disagreement, and cost a genuine listing a little RTA.

SPEC Specificity & Evidence

How many of these reviews say something checkable?

"Great product, works as described" can be written by someone who never opened the box. "Nine hours of battery, about two more than the model it replaced" cannot.

Computed
Every review is tagged for five evidence types — a measurement or quantity, a named entity, a use context, a comparison, or a process or duration — and counts as evidence-bearing if it carries at least one. 20% of the set evidence-bearing or fewer scores −1; 70% or more scores +1; linear in between.
Where it fails
This is the dimension most exposed to AI-written reviews, which are specific and moderately detailed by construction. SPEC alone cannot separate a well-written machine review from a well-written real one — see what this analysis cannot detect.

TCBR Topic Coverage, Balance & Redundancy

Are people discussing the whole product, or repeating one talking point?

Genuine use generates a spread of subjects — the thing it does well, the annoying part, how it fits into a routine. A campaign generates one message in different words.

Computed
Three measurements combined. Coverage: the share of reviews mentioning at least one substantive topic. Balance: the normalised entropy of the topic distribution — 0 when every mention lands on a single topic, 1 when mentions are spread evenly. Redundancy: the mean overlap between the reviewer sets for each pair of topics, which catches two labels that are really one topic and would otherwise inflate balance. Note what redundancy does not do: it never asks whether two words look similar, only whether the same people raised them.
Where it fails
Single-purpose items — a charging cable, a replacement filter — legitimately produce narrow discussion, which is why a low TCBR alone never determines a grade. When fewer than two topics can be extracted, the dimension is dropped rather than guessed at.

EXT Rating Distribution Health

Does the star distribution have the shape real products produce?

Genuine listings tend toward a J: five stars most common, one star the second peak, the middle bands thin but present. Almost every real product disappoints somebody mildly. The signature of a purchased review set is the missing middle.

Computed
Two independent penalties on the star histogram. Concentration: the largest single star band above 55% starts costing points, reaching the full penalty at 90%. Missing middle: the 2–4 star bands totalling under 15% starts costing points, full penalty at zero. Where the platform publishes a complete histogram we use that rather than our own sample, and the product page records which was used.
Where it fails
Excellent products in small niches do produce concentrated distributions honestly, which is why the penalty is graded rather than binary and why no listing is downgraded on this dimension alone.

HYPE Promotional Language vs Authentic Voice

Does this read like an owner, or like the product page?

People who actually own things qualify what they say. They admit a downside, they mention how long they have had it, they say who it isn't for. Marketing copy does none of those, because none of them help sell.

Computed
Two per-review flags. Promotional tone: superlatives, product-description phrasing, no first-person detail. Authentic voice, derived in code from three separate annotations — the review concedes a downside, gives first-hand usage detail, or states a limit on where it applies; any one of the three qualifies. A promotional review moves the number further than an authentic one does, deliberately: writing like a person is merely normal, while writing like an advertisement is an active signal.
Where it fails
Genuine enthusiasts write like advertisements, and some categories — beauty, supplements — borrow marketing vocabulary as ordinary speech. And, as with SPEC, a model writing a review will avoid exactly the patterns this dimension penalises.
When a dimension can't be computed

A dimension that cannot be measured returns nothing and its share is redistributed across the others; the confidence figure falls accordingly. It is never filled in with a neutral value. "We could not measure this" must never be able to look like "this was average" — that substitution is how scoring systems quietly manufacture middling grades out of missing data. Every product page states which dimensions were available.

3. Data, minimum sample size, and refresh cadence

Where the data comes from

Everything we analyse is public information on the retailer's own product and review pages — primarily Amazon, with support for Target and Best Buy. We hold no purchase records, no reviewer identity beyond the public display name, no seller-supplied data, and no privileged platform access.

Fields collected per analysis
LevelFieldsUsed for
Per review Star rating, title, body text, date posted, verified-purchase flag, variation or style selected, helpful votes, attached photos or video, public reviewer name All five dimensions; the rule layer that labels individual reviews
Per product Title, brand, category, current price, the platform's own rating histogram, total ratings vs total written reviews, variation family, first-available date EXT distribution, sample-vs-population checks, detecting reviews older than the listing

How many reviews we read

Up to 100 reviews per listing. The ceiling is a statistical judgement, not a technical one: for the ratio-type measures this method depends on — evidence rate, promotional rate, suspect rate — the confidence interval narrows sharply up to roughly 50 reviews and reaches about ±10% at 100. Past that, more fetching buys very little precision.

What we publish at each sample size
Reviews availableWhat the page shows
n < 5No score. The dimensions are not computed at all.
5 ≤ n < 20Not Enough Evidence. Dimension figures and the reviews themselves are shown; no letter grade is issued.
n < 50, all five-starGrade issued but marked down, with an explicit note that manipulation cannot be ruled out at this sample size.
n ≥ 20Full score and grade, with the confidence figure shown alongside.

Confidence rises with sample size and with the share of reviews we could classify one way or the other, reaching its maximum at around 50 usable reviews. It is displayed on every product page and it is what drives the shrinkage toward the baseline in section 1 — the mechanism that stops a thin sample from producing a confident-looking grade.

When scores update

Every analysis is a snapshot, and every product page prints the date it was taken. Nothing on this site is a live feed.

Because each analysis stands alone, a score can move between snapshots for reasons that have nothing to do with manipulation: the platform removed reviews, new reviews arrived, the seller edited the listing. Where a score changes materially we keep the previous result rather than overwriting it silently.

4. Known errors and limits

These are failure modes of the method as built, distinct from section 5, which covers things the method cannot see even when working perfectly.

Where the method is known to go wrong
LimitWhat it means for a score you're reading
Sampled distributions When the platform's full rating histogram isn't available, EXT runs on our sample — and platform review ordering is not random, so the sample skews toward reviews others found helpful. Product pages record which source was used.
Annotation is not perfectly repeatable The per-review annotations come from a language model. Prompts are fixed and versioned, and a set of fixed reference products is re-scored on every change to catch drift, but two runs on the same reviews can differ slightly. Differences of a point or two are noise.
Thresholds are calibrated, not derived Numbers like "55% concentration" and "20% evidence rate" come from fitting against hand-rated products, not from theory. They are the current best fit for the categories we have rated, and they will be wrong at the edges of categories we have rated less.
Category effects Books, supplements and fashion generate review sets that look structurally different from electronics. A single set of thresholds across all categories systematically flatters some and penalises others.
Grade boundary sensitivity Grades are cut at quantiles. A listing one point either side of a cut receives a different letter for a difference that is not meaningful. Adjacent grades should be read as the same result.
Unclassifiable reviews Reviews we cannot confidently call real or fake are excluded from the ratio entirely rather than assumed clean. This is the conservative choice, but it means a listing with many ambiguous reviews has a less informative authenticity term than the number alone suggests.
Non-English reviews Analysis runs on the original review text. Translated pages translate our summary, not the underlying judgement, and annotation quality does vary by language.
If you think a score is wrong

Every judgement we publish is stated as a countable fact with its evidence attached — how many reviews matched which signal, quoted — rather than as a verdict about a product or a seller. If you are a brand or seller and believe an analysis misreads your listing, contact us with the listing URL and we will re-run it and publish a correction where the analysis was wrong.

5. What this analysis cannot detect

Two kinds of blind spot are listed below, and they are labelled separately on purpose. Some things are undetectable in principle from review text — no future version fixes them. Others are simply not covered in this version, and we expect to remove them from this list. Mixing the two together would be the easier thing to write and the less useful thing to read.

We cannot tell if a reviewer was paid or refunded Undetectable in principle

The most common form of review manipulation is also the hardest to see. A seller refunds the purchase or ships a free unit in exchange for a positive review. The buyer is real, the purchase is real, the product was really used — and the review carries a genuine "Verified Purchase" badge. Nothing in the text is anomalous. Incentivized reviews are often more detailed than organic ones, because sellers coach them.

Note the consequence for how you read a listing: a high share of verified purchases is not evidence of a clean review set, and we do not treat it as one. What we can see is shape at the group level — an unusual gap in the four-star band, a week where review volume jumps far above the listing's own baseline, disclosure phrasing that survived editing. An individual reviewer's motive is not in the text; the statistical footprint of a campaign sometimes is. Both of those statements are true and we publish both.

An AI-written review may score as trustworthy — or higher Open weakness

Reviews generated by large language models are specific, moderately detailed, and often include a small criticism to sound balanced. Those are the same properties our Specificity & Evidence and Hype & Promotional Bias dimensions use to judge a review as credible. An AI-written review can therefore score higher than a lazily written real one. We treat this as an open weakness, not a solved problem.3

We cannot tell if a review was written about a different product Not covered in this version

Reviews can carry over from another item — through variation listings, or when a seller repurposes an old listing for new merchandise. Amazon began separating reviews across functionally different variations on February 12, 2026, with the rollout completing by May 31, 2026, but reviews are still shared across color, size, pack quantity, scent, and device-fitment variations. We analyze the reviews displayed on the listing today; we cannot verify which variant each one actually describes.1, 2

We only see the reviews that are visible right now Single snapshot

This analysis is a snapshot taken on the date shown above. Reviews the platform has already removed are invisible to us. So are patterns that only appear over time — a sudden burst of reviews in a single week, or a rating that drifts downward months after launch. Those require review history we do not yet collect.

We do not assess the product itself Out of scope

A Trust Score is about the reviews, not the item. A product can have completely genuine reviews and still be a counterfeit, poorly made, or unsafe. We do not test products, verify authenticity, or check safety recalls.

We do not assess the seller Out of scope

Whether the seller ships on time, honors returns, replies to messages, or is still in business next month is not something review text reliably reveals. Our score says nothing about the post-purchase experience.

We do not know whether this product is right for you Out of scope

Budget, use case, compatibility, sizing, skin sensitivity — none of that is in scope. "These reviews look genuine" is not the same statement as "you should buy this."

6. A high score is not a recommendation

If you take one thing from this page, take this. A Trust Score describes the evidence available to you on a listing. It says nothing about what happens after you click buy.

An A grade is not a statement that the product is safe or well made

We never open, test, or inspect a product. A well-made product and a badly-made one can both have completely honest reviews — and the honest reviews of a bad product will say so, which is why the dimension detail and the quoted reviews matter more than the letter.

An A grade is not a statement that the seller is reliable

Shipping times, returns handling, warranty support and whether the storefront still exists next quarter are not visible in review text in any dependable way. A clean review set tells you nothing about the post-purchase experience.

An A grade is not a guarantee of authenticity

Counterfeits sell alongside genuine stock on the same listings and accumulate real reviews from buyers who never realised. We have no way to verify that the unit shipped to you is the product the listing describes.

An A grade is not advice that you should buy it

Fit, size, compatibility, budget, sensitivity, whether you'll use it — none of this is in scope, and a trustworthy review set for a product that is wrong for you is still the wrong purchase.

"These reviews look genuine" is not the same statement as "you should buy this."
Affiliate disclosure

Some outbound links to retailers earn us a commission. The commission is identical regardless of the score we publish — there is no arrangement under which a higher grade pays us more, and sellers cannot pay for a score, a re-analysis, or a removal. We are an independent tool and are not affiliated with, endorsed by, or sponsored by any retailer whose listings we analyse.

See the method applied to a listing you care about.

Analyze a product

7. Questions

Why does a product with 4.8 stars get a C?

The star average is one of three inputs, and it is the one the platform already shows you. A C means that while the average is high, something about how those reviews are written or distributed does not match what genuine review sets look like — the dimension detail on the product page names which one.

Why won't you grade a product with 8 reviews?

Because with 8 reviews we cannot tell a good product from a manipulated one, and saying so is more useful than guessing. Not Enough Evidence is a result, not a failure — we still show the reviews and the underlying figures so you can read them yourself.

Do you say which individual reviews are fake?

We label reviews internally and show the evidence behind group-level findings, including quoted excerpts. We state countable facts — how many reviews matched which signal — rather than declaring that any particular person wrote a fake review, because for any single review that claim is not one the text can support.

How is this different from reading the reviews myself?

Mostly in what a person cannot do by hand: comparing every review against every other for near-duplicate text, checking whether the star distribution has a missing middle, measuring how concentrated the topics are, and computing what share of the set contains any first-hand evidence at all. You are better than us at judging whether a specific complaint matters to you.

Can a seller pay to change a score?

No. Sellers cannot pay for a score, a re-analysis, or a removal, and our affiliate commission is identical regardless of the grade we publish. If an analysis is factually wrong, contact us and we will re-run it and publish a correction.

8. Sources and citation

This page describes our own method, so most of it cites nothing but itself. Where it relies on an external fact, that fact is listed below. The wider research on how much of the review economy is fabricated is collected separately in how many online reviews are fake in 2026.

  1. PPC Land. Amazon splits variation reviews in bid to stop misleading ratings. 8 January 2026. Source for the 12 February 2026 start and 31 May 2026 completion of the variation review split.
  2. Ecomclips. Amazon Review Sharing Variation Update 2026: What Sellers Must Fix Before Reviews Split. 24 February 2026. Source for which variation dimensions — colour, size, pack quantity, scent, device fitment — continue to share a review pool after the split.
  3. Trust at risk: Detecting misinformation in LLM-generated product reviews and its implications for consumer behavior and platform governance. ScienceDirect, 2025. Support for treating machine-written reviews as an open weakness rather than a solved detection problem.
Cite this page. Methodology changes, and a score is only interpretable against the version that produced it — so cite the version and date you read, not just the URL. Review Detector. "How the Trust Score works." Scoring version v2, 27 August 2026. https://reviewdetector.ai/en/methodology
ReviewDetector Methodology: How the Trust Score Works