How We Rate
Every score is reproducible and derived from real customer data. Here's exactly how.
1. Scoring from the rating distribution
For each product we collect the full distribution of customer star ratings (how many 1-, 2-, 3-, 4-, and 5-star ratings it has). The score is the weighted average of that distribution, normalized to a 1–10 scale:
score = (1·n₁ + 2·n₂ + 3·n₃ + 4·n₄ + 5·n₅) / total ÷ 5 × 10
This means a product's score reflects what thousands of real buyers reported — not a single average star figure, and never an editor's opinion.
A worked example
Suppose a product has 50 one-star, 20 two-star, 30 three-star, 100 four-star, and 300 five-star ratings (500 total). The weighted average is (1·50 + 2·20 + 3·30 + 4·100 + 5·300) ÷ 500 = 4.16 stars, which normalizes to 8.3 / 10.
2. The 100-rating minimum
The formula above is a pure average, which means it is blind to how many people rated a product: one five-star rating also averages 5.0. So a product must have at least 100 customer ratings before it is eligible to be ranked at all. Below that we don't consider the signal strong enough to publish, and the product is left out of the category entirely.
Beyond that floor, we do not weight by volume — a product with 300 ratings and one with 300,000 are scored on the same scale. We show every product's rating count next to its score so you can judge for yourself how much data a given ranking rests on.
3. Curating the top picks
We gather more candidates than we publish, then select the best of them for each category — weighing overall quality, value, and how credible and positive the reviews are. Products with no scraped review text are excluded, as are products below the 100-rating floor.
4. Pros, cons, and summaries
The summary, pros, and cons shown for each product are written by a large language model reading the actual customer review text we collected for that product. They reflect recurring themes in what buyers reported, rather than manufacturer claims — but they are machine-generated, and we label them as such on every product.
The model never sees or influences the numeric score. Scoring is the deterministic arithmetic above, computed before any text is generated. A model is also used to narrow a category's candidate list down to the shortlist we publish.
5. Freshness
Rankings are regenerated when we re-scrape a category, so scores track the latest customer sentiment. Every category page shows the date its rating data was last refreshed — if that date looks old, the ranking is old, and we'd rather show you that than hide it.
Limitations & transparency
Our scores reflect what customers reported, so they inherit the limits of that data: rating volumes vary between products, review text can be noisy, and sentiment can lag a product's most recent revision. Amazon also pools ratings across a product's variants, so a rating count can cover more than the exact model listed. We surface each product's rating count so you can weigh how much data a score rests on, and we exclude products below the 100-rating floor. If you spot a ranking that looks wrong, tell us and we'll re-check it.
What we don't do
- We don't accept payment to rank or score a product higher.
- We don't include sponsored or ad listings in our rankings.
- We don't hide our scoring — the formula above is the whole story.
- We don't bench-test products, and we don't claim to. See about us.
- We don't pass off machine-written summaries as human editorial — they're labelled on every product.