Research

The CrUX Field-Data Pre-Specification

What was fixed before the first field record was requested.

Showroom Standard measures dealer websites in a laboratory. A throttled run against a single page, under fixed conditions, repeated the same way across thousands of sites. That instrument has never measured what a real visitor’s browser experienced, and it was never built to.

The Chrome UX Report is the other kind of reading. Google publishes it from real Chrome sessions, for origins where it observes enough of them to meet its own reporting threshold. Setting one beside the other is the obvious next question and it is also the easiest place in this programme to say something false, so the rules were written down before the first record was requested.

This page is that specification, written out for a reader outside the project. It is published so that anyone reading a figure from this instrument can see which decisions were taken before the first record was requested, and check them.

The population is a pinned base and it is not the cohort

This pull runs against a pinned set of dealer domains, fixed by name before the first request: the dealers whose Google Business Profile resolves from their own website domain and which carry both a review count and a scored mobile audit. It does not run against the verified dealer cohort, which is very much larger.

That is a deliberate narrowing and the reason is worth stating. Keeping the field population identical to the population already read for profiles and for laboratory performance means one set of businesses described by three instruments, rather than three overlapping sets whose differences would have to be disentangled afterwards.

The consequence binds every figure the pull produces. Every coverage denominator is the pinned base. A cohort coverage rate is a different figure on a different population and this pull does not produce one, and none is inferable from it. The size of the pinned base and every coverage figure computed on it are reported in the research note that carries the finding, under a claim row whose publication scope does not reach this page.

The unit is an origin, and the base is keyed by something else

A CrUX record is keyed by an origin: scheme, host and port. The base is keyed by a domain. Mapping one onto the other is a decision rather than an implementation detail, so it is fixed here.

For each pinned domain: lowercase it, strip any leading www., strip any path, query, fragment and port, then query https:// plus the result. One origin per pinned dealer, one query per origin, one row per query.

https://example.com and https://www.example.com are different origins to the Chrome UX Report. This specification queries the apex form. That is a choice with a cost, and the cost is recorded below rather than discovered later.

Nothing is requested except what is named

The form factor is PHONE on every request and is written on every row. There is no all-form-factor row, no desktop row, and no row where the field is absent, empty or inferred. The programme’s published performance figures are mobile figures, the question is what a person reaching a dealer on a phone arrived into, and a mixed aggregate answers neither. Desktop is a different question and needs its own specification.

Three metrics are requested and no others: Largest Contentful Paint, Interaction to Next Paint, and Cumulative Layout Shift, each as its published histogram densities and its published 75th percentile value. The source of that percentile is Google’s Chrome UX Report, which publishes it, and this specification, fixed on 2026-09-12, requires every row to store it exactly as received. It is not a percentile this programme chose, and nothing here recomputes one.

They are recorded as published. The bucket boundaries, the densities and the percentile values are written to storage as the API returns them. This pull does no bucketing of its own, applies no threshold of its own, and derives no pass rate at collection time. The collection period the API reports is written on every row.

No metric is added to that list after the first row lands. Adding one is an amendment carrying its own date and reason, and an amendment after collection has begun is a new question wearing the word amendment.

Absence is a measurement, and it is the first output

This is the load-bearing rule of the whole document.

The Chrome UX Report holds a record for an origin only where Google observes enough real Chrome visits to meet its reporting threshold, and the dealer layer is largely small local businesses. If absence were stored as a missing value, every figure computed afterwards would quietly become a figure about the covered subset, and the covered subset is selected on traffic. The population finding and the selection bias are the same fact. Storing absence as a null throws away the first and keeps the second.

So coverage takes exactly one of three values on every member, and the three are never collapsed into two:

ValueMeaning
presentThe API returned a record for this origin.
absentThe API answered, and its answer was that it holds no record for this origin. A measured answer about the population.
could_not_measureThe pull could not obtain an answer: transport failure, quota stop, tripwire stop, malformed response. Not a statement about the origin.

absent and could_not_measure are never merged. A pull that failed has learned nothing about the origin, and folding those in with the origins Google genuinely holds no record for turns a figure about the population into a figure about the pull.

The first output of this pull is the coverage table, reported before any other figure is computed from it, with its denominator stated, its read date stated, and an interval on every share.

The apex rule has a known cost, and the cost is measured rather than disclosed

A dealer whose traffic lands on the www host can hold a record at the form this specification does not query, and would be recorded as absent. That is a false negative of the apex-only definition, and it was observed before the pull rather than theorised after it.

It is not random, either. A site that canonicalises to www is a site with enough history to have made that choice, and those are disproportionately the sites the Chrome UX Report covers. The apex-only rate is biased downward, and biased selectively, against the segment most likely to be present.

A sentence cannot say how large that is, so the class is measured. Every domain recorded absent at the apex is queried once more at its www form. The result is stored as a separate row carrying its own origin form, never merged into the first row and never overwriting it.

That makes two definitions of one metric, and both travel. The apex-only rule is one definition and the rule accepting either form is another, each carrying its own definition version in the metric registry. A coverage figure computed across both forms is labelled as the wider definition, on every appearance, beside the apex-only figure rather than instead of it. Results are comparable only within a definition version.

Why not simply adopt the wider definition. Choosing the definition that produces the better coverage number, after seeing that it produces the better number, is the exact move a pre-specification exists to prevent. Both are collected and both are reported.

Known false negatives, recorded before the data

  • Apex-only querying, as above.
  • An origin is not a URL. An origin record describes all traffic to a host. It does not describe the page any laboratory audit measured, and where the audited URL is an interior page the two instruments are not looking at the same object at all. No output may treat them as interchangeable.
  • The threshold is Google’s. Absence means Google publishes no record, which is a statement about a reporting threshold as much as about a dealer’s traffic. It is not a measurement of traffic and is never reported as one.

Every row is a dated snapshot, and the dates are not ours

A row carries three dates and is not written without all three: the date the pull ran, and the first and last dates of the collection period the API reported for that record. A CrUX record describes a trailing window this programme does not set, so the read date alone does not date the figure.

A later pull writes new rows. It never updates an existing row. Two readings of one origin are two rows, and a difference between them is a difference between two dated readings.

A figure from this instrument is a value Google published. This instrument did not measure it. A later reading that differs means Google’s published figure changed, which is a different thing from a dealer’s site improving or declining. A published sentence names the figure as Google-reported in the same breath as the figure, and a limitations section further down does not discharge that.

What may not be converted into what

Field data is closer to a real visit than a throttled laboratory test. Closer is not the same.

No figure from this instrument becomes a description of what a visitor did. An absent record is an absent record: it is not a customer who did not arrive, a call not placed, or demand that went elsewhere. A published percentile for interaction delay does not become abandonment, a call, or a sale. This is the same rule that governs the laboratory instrument and it is not relaxed because the reading is nearer to a person.

Setting field beside lab, and the gate that decides whether it publishes

The chapter’s opening question is whether the slowness the laboratory measures shows up for real visitors, and the comparison that would answer it was specified before the pull rather than after.

It may say whether the two instruments place the same origins on the same side of a published threshold, and how often they do not, on a named and dated base.

It may not say that one instrument is right and the other wrong, that a dealer’s real visitors experienced anything in particular, or that laboratory measurement is validated or invalidated.

Four points travel with any such comparison, in the same breath as the comparison itself rather than in a later section, and one of the four is a match rather than a difference. The metrics differ: a percentile over a trailing field window against a single throttled laboratory reading. The regimes differ: field against laboratory. The populations match, by construction, because both instruments read the same pinned base. And the windows do not match.

Three limits bind every sentence it could produce. An origin is not a page, and where the audited URL is an interior page the two instruments are not looking at the same object. The two windows are one to eight months apart, the laboratory audits having run between 2026-01-17 and 2026-08-02 at a median of 2026-03-04, and a site may have changed in between. And disagreement is not error: an origin passing on one instrument and failing on the other is the expected consequence of two instruments measuring different things under different conditions.

Above all it is gated on coverage. Below a pre-specified coverage line, the covered subset is selected on traffic and a comparison drawn from it describes the busiest dealer websites rather than the pinned base. Below that line the comparison does not publish anywhere. It may be computed and held, and the coverage rate and the field distribution publish alone.

That gate was written before anyone knew which side of it this pull would land on. It is the part of a pre-specification that only means anything if it is allowed to bind when it is inconvenient.

What this specification does not authorize

It authorizes no re-audit, no laboratory call, no write to the audit record, and no change to the measurement roster. It fixes one pull, at one form factor, against one pinned base, and questions it does not name are not answerable from it: whether either side moved over time, anything about desktop, and any figure about an individual dealer. Each is a real question and each needs a specification written before the pull that answers it.

No dealer is named in anything built on this instrument.

More research

View all research