Research Note

More Reviews, More Sites Below Fifty

645 dealers, grouped into quartiles by the number of reviews Google reported for them on one day in September 2026. In the most-reviewed quartile, 45.7% of sites score below 50 in a throttled lab test of the site taken months earlier, against 26.7% in the least-reviewed.

Cody Stidham · September 12, 2026 · 11 min read

A dealer’s Google Business Profile and a dealer’s website are two separate records of the same business, kept by different parties and read by different people. Across 645 dealers, sorting by the first orders the second, and it orders it downward. The median mobile performance score falls at every step: 60, 57, 55, 51.5. Two instruments are being read side by side there and neither is a version of the other: the review count is a number Google reported on 2026-09-09, and the score is a throttled lab measurement of a page taken between January and August 2026, at a median of March 4.

Nothing here says a dealer with more reviews ought to have a faster site, and this note does not argue that one produces the other. What the base settles is narrower and more useful to anyone reading a dealer off a screen: the review count is not a proxy for the website. The two records do not agree, and they disagree consistently enough to see at every cut. What a review count is also not is a measure of trading: it says nothing about how busy, how large or how long-established a dealer is, and nothing below rests on it doing so.

The gradient

645 dealers whose Google Business Profile resolves from their own website domain, each carrying both a review count and a scored mobile audit, ordered by review count and split into four quartiles.

QuartilenMedian reviewsMedian mobile scoreMedian mobile LCP, secondsMedian desktop score
Fewest reviews16112608.6078
16130579.6879
16161559.0173
Most reviews16215851.59.9370

The quartile means move the same way, 58.9 to 55.1 to 54.4 to 51.0, and so do both ends of the spread. The lower quartile of scores runs 48, 43, 42, 36. The upper quartile runs 68, 65, 65, 62. The whole distribution shifts down rather than a tail dragging the median. The distributions still overlap across most of their range; what moves is the middle of each.

Desktop moves too, 78, 79, 73, 70, with the first two quartiles effectively level. Mobile LCP in seconds runs 8.60, 9.68, 9.01 and 9.93, which moves in the same direction across the ends and is not monotonic in the middle. The performance score is the measure that steps at every quartile, and it is the one the table is ordered on.

How the quartiles were cut, because 645 does not divide by four. The base is sorted ascending on review count and then on domain, and cut into blocks of 161, 161, 161 and 162, with the remaining member landing in the most-reviewed quartile. The domain key is a tie-break and means nothing on its own: 10 dealers tie at 19 reviews across the first boundary, 5 at 43 across the second and 3 at 91 across the third, and which of them lands where is decided alphabetically. Neither the allocation nor the tie-break is forced by the data, so the point values above are contingent on both. The direction is not. Under an array_split allocation of 162, 161, 161 and 161, under a tie-break on the internal dealer id, and under a descending-domain tie-break, the gradient runs the same way and the ratio between the extreme quartiles’ shares below 50 runs between 1.6 and 1.75.

The share of sites that score below 50

The clearest statement of the gradient is not a median. It is how many dealers in each quartile have a site scoring below 50. The source for that line is Google’s own Lighthouse performance scoring documentation, which labels the 0 to 49 band Poor. That is the scale owner’s label rather than a verdict reached here.

QuartileBelow 50Share95% interval
Fewest reviews43 of 16126.7%20.5 to 34.0
56 of 16134.8%27.9 to 42.4
59 of 16136.6%29.6 to 44.3
Most reviews74 of 16245.7%38.2 to 53.4

The line is not a sharp one. Across verified cohort sites audited more than once inside a single month, close enough together that the site itself is unlikely to have changed, the median absolute difference a site shows against itself is 4.0 points on this score, so a site sitting within four points of that line changes side of it as a matter of instrument noise. What the table reports is where a population sits. It reaches no verdict on any dealer in it.

The share rises at every step, and the interval on the most-reviewed quartile does not reach the interval on the least-reviewed. The ratio between the two ends is 1.71. That ratio is a point estimate computed from two figures that each carry an interval, and it is quoted here beside both rather than on its own, because a ratio stated alone reads as more precise than the two numbers it came from.

Whether the gradient survives being checked

Two checks, and both are apparatus rather than presentation.

The Spearman rank correlation between review count and mobile performance score across the 645 is minus 0.1811. That is a weak correlation, and it is reported as one. It says the ordering is real and it does not say the relationship is strong. It is also the one figure here computed from both sides at once, which the section below on the two instruments returns to.

The second check matters more. Business type could produce this on its own if the categories with many reviews happened to be the categories with heavy websites. Within the 401 dealers carrying the single Google category Building Materials Store, the same split gives median review counts of 16, 38, 69 and 173 against median mobile performance scores of 58, 56, 55 and 49, on quartiles of 100, 100, 100 and 101. The rank correlation inside that one category is minus 0.1590. The gradient is not an artifact of comparing different kinds of business.

What this does not settle, and it is the obvious thing

Business size would produce this gradient without any relationship between reviews and websites at all. A larger dealer accumulates more reviews over more years and also runs a larger site with more images, more locations and more third-party tags, and a larger site scores worse. That explanation fits every figure above.

This read cannot separate it from any other explanation. There is no size, age, revenue or page weight measure joined here to test it with. The confound is recorded as unresolved. It is not argued against, and nothing in this note should be read as ruling it out.

Direction of cause is not established in either direction and is not claimed. Two independently measured quantities vary together across a population. That is the whole of it.

The part that does not vary

Sorting by review count separates how bad the sites are. It does not separate out any good ones.

QuartileScoring 90 or aboveShare95% interval
Fewest reviews4 of 1612.5%1.0 to 6.2
3 of 1611.9%0.6 to 5.3
5 of 1613.1%1.3 to 7.1
Most reviews5 of 1623.1%1.3 to 7.0

Every interval overlaps every other. Across the base as a whole, 17 of 645 reach 90 or above, which is 2.6% with an interval of 1.7 to 4.2. The quartile a dealer falls in tells you how much worse than poor the site is likely to be, and it does not tell you whether the site is good, because on this evidence the good end is nearly empty whichever quartile you stand in.

Separately, and on a different metric: fewer than 4% meet Google’s 2.5 second mobile LCP threshold in a standardized lab test across the whole measured cohort. That is a different measurement from the score used here, on a different and much larger population, and the two are not evidence of each other. On these 645 dealers the sites passing the LCP threshold and the sites scoring 90 or above are not the same sites.

The two other things this base was sorted by, both of which found nothing

Review-count quartiles are one of three splits run on this base. The other two are reported at the same size, carrying the same group counts, medians, below-50 shares and intervals, because they were found the same way and reporting only the one that produced a result is how a post-hoc finding gets its edge.

Google primary category does not separate performance. The five categories carrying twenty or more dealers are Building Materials Store with 401, General Contractor with 70, Furniture Store with 65, Manufacturer with 37 and Suppliers with 20, and their median mobile performance scores are 56, 59, 57, 58 and 59.5. Their shares scoring below 50 are 38.9% [34.3, 43.8], 32.9% [23.0, 44.5], 32.3% [22.2, 44.4], 29.7% [17.5, 45.8] and 10.0% [2.8, 30.1]. The remaining 52 dealers sit outside those five: 42 of them spread across 11 further categories, none of which reaches 10 dealers, and 10 carrying no category at all.

Rating does not track performance either, and its bands are not ordered. The median mobile performance score is 57.5 for the 30 dealers rated below 4.0, 55 for the 136 rated 4.0 to 4.5, 56 for the 190 rated 4.5 to 4.8 and 57 for the 289 rated 4.8 or above. The shares below 50 are 20.0% [9.5, 37.3], 39.0% [31.2, 47.4], 38.4% [31.8, 45.5] and 34.6% [29.4, 40.3]. The rank correlation between rating and mobile performance score is plus 0.0014, which is nothing. It is worth saying plainly that the Google signal a reader is most likely to read as reputation is the rating, and the rating has no relationship with website performance in this base at all.

What the base is, and what it leaves out

The 645 are drawn from a single partial tranche, one of eight, of a cohort of 7,700+ verified dealer sites measured. The run read 740 of that tranche’s 964 members before the month’s call budget was spent, and 739 of the 740 produced a record. Of those 739, 693 produced a resolvable profile and 653 have one that resolves from the dealer’s own website domain. The 645 are those of the 653 carrying both a review count and a scored mobile audit. This is a single-tranche estimate and not a census, which is why every share above carries an interval.

The narrower base is a deliberate choice and the wider one is why. The wider base includes dealers matched to a profile on business name alone, and one of those matches is a metropolitan daily newspaper, whose matched profile Google marks permanently closed. A name match can attach a profile that does not belong to the business at all. That organisation sits outside the base used here and outside every figure in this note.

The narrower base is not clean either, and the impurity is disclosed rather than removed: 7 of the 653 belong to organisations that are not flooring dealers, a count that is a floor found by reading the category tail rather than an audit.

Both dates, because there are two instruments here

This is a join, and the two sides were not measured at the same time.

The review counts are a single read taken on 2026-09-09. The performance scores are the most recent PageSpeed audit on record for each dealer, and those audits were taken between 2026-01-17 and 2026-08-02, with a median date of 2026-03-04. The site measurements therefore precede the review reading by one to eight months, unevenly across the base.

That gap cannot be closed by collecting more, because measurement collection against cohort sites is frozen until 2026-11-01. It is stated rather than worked around, and both dates travel with every figure above.

The audit side is also pinned. “The most recent audit on record” is a query, and the audit table is written continuously by the live product outside research control, so the same query run next month would select different rows and move the window and the median date silently. The 645 audits read here are recorded as a fixed set, with the values read from each and the date they were read, and every figure above is recomputed from that set.

What this does not establish

Review count is the number of reviews Google reports for a business. It is not a count of customers, visits, calls or sales, and nothing here describes a dealer as busier, larger, better trading or better regarded. The performance scores are throttled lab measurements of a web page. They are observations of an instrument rather than of any real visit, and no figure here says what a person searching would see or whether any page produced anything.

Nothing in this note measures rankings, traffic, impressions or visibility, for any dealer or for the population.

The review count and the rating are values Google reports rather than values this instrument measures. A future reading that differs will mean Google’s reported figure changed, which is a different thing from a dealer improving or declining.

All research notes