01 · Two lenses

Controlled runs explain; real visits validate

Lab tools load a page under defined device, network and CPU assumptions. This makes before/after debugging and like-for-like experiments practical. Lighthouse can expose the LCP element, request chains, render-blocking resources, long tasks and layout-shift sources in one reproducible report.

Field tools collect eligible experiences from real visitors. The Chrome UX Report (CrUX) aggregates opted-in Chrome users who meet its reporting criteria. Traffic spans devices, networks, locations, cache states and user behavior. Field data therefore answers whether actual users receive a good experience—not whether one test run is clean.

LabField
PopulationSimulated single loadMany eligible real visits
ConditionsDefined and repeatableNaturally variable
InteractionUsually limitedReal user interactions
Best useDiagnosis and regression testsOutcome monitoring
Blind spotTraffic diversity and long sessionsExact cause of a problem
Do not say “real-world Lighthouse score”

Lighthouse is lab tooling. CrUX and properly instrumented Real User Monitoring provide field measurements.

02 · Know the vocabulary

Map lab diagnostics to Core Web Vitals carefully

The current Core Web Vitals are Largest Contentful Paint (LCP), Interaction to Next Paint (INP) and Cumulative Layout Shift (CLS). Google recommends assessing “good” at the 75th percentile of page loads, segmented by mobile and desktop. The good thresholds are LCP at or below 2.5 seconds, INP at or below 200 milliseconds and CLS at or below 0.1.

Lighthouse reports LCP and CLS for its controlled load. It generally cannot produce representative INP because there is no population of real interactions. Total Blocking Time (TBT) is a useful lab diagnostic correlated with main-thread responsiveness, but it is not INP and should not be relabelled.

LCPLoading

Time until the largest qualifying visible element renders.

INPResponsiveness

Field latency across click, tap and keyboard interactions.

CLSStability

Unexpected layout-shift accumulation within session windows.

TBTLab diagnostic

Blocking time from long main-thread tasks between FCP and TTI.

First Contentful Paint and Time to First Byte are supporting metrics, not Core Web Vitals. They can still reveal early rendering and server-response constraints.

03 · Investigate disagreement

Why a page can score 100 and fail users

A clean lab run may start with an empty cache but omit consent flows, personalization, logged-in tooling or interactions that happen later. Field visitors may use slower phones, arrive from farther away, encounter cold server caches, expand menus, add products or trigger delayed scripts. The lab page may also differ from the URL group represented in field reporting.

The reverse can happen too: a throttled lab profile might appear slow while a site’s actual audience uses fast devices and repeat caches. Neither result invalidates the other. Their disagreement is diagnostic.

A five-question conflict check

  1. Same URL? Page-level data, origin-level data and a single test URL are not equivalent.
  2. Same device class? Compare mobile with mobile and desktop with desktop.
  3. Same time window? CrUX is aggregated over a rolling period; a lab fix is immediate.
  4. Same page state? Consent, experiments, authentication and cache state alter execution.
  5. Same metric? TBT is not INP; one-run CLS is not a full visit’s layout-shift experience.

Use lab traces to form a cause hypothesis, then segment real-user data to verify it. For example, slow field LCP with good lab LCP may cluster on product pages, a geography, uncached responses or low-end devices.

04 · Read the dataset

CrUX is powerful, aggregated and sometimes unavailable

CrUX includes origins and URLs with enough eligible traffic to meet publication thresholds. Smaller sites and low-traffic pages may have no report. Absence is not a performance result. Public tools may fall back from page-level to origin-level data; check the displayed scope before attributing an origin result to one template.

Percentiles matter. Passing Core Web Vitals means the 75th-percentile experience is within the good threshold for each metric. An average can conceal a harmed tail of users. Segment mobile and desktop because their device and network populations differ substantially.

Theme-level field claims require caution

Live WordPress sites using the same theme can differ in hosting, builders, plugins, versions, content, optimization and audience. A theme cohort can reveal ecosystem outcomes, but it does not isolate theme code. Detection errors and survivorship should be disclosed. Treat this evidence as observational, not a controlled theme benchmark.

05 · Close the loop

A practical measurement workflow

  • Before launch: set budgets and test representative templates in a controlled lab.
  • During development: use Lighthouse traces and browser profiling to diagnose regressions.
  • At release: verify production URLs from more than one region/device profile.
  • After launch: instrument Web Vitals or use CrUX/Search Console when eligible.
  • When field data fails: segment by template, device, version and geography; reproduce in lab.
  • After fixing: confirm immediately in lab, then wait for sufficient field observations.

Keep versioned baselines instead of replacing old results. A performance history makes regressions, seasonality and delayed field improvements visible. Record deployment markers alongside the time series.

Reference desk

Primary sources