01 · Two lenses
Controlled runs explain; real visits validate
Lab tools load a page under defined device, network and CPU assumptions. This makes before/after debugging and like-for-like experiments practical. Lighthouse can expose the LCP element, request chains, render-blocking resources, long tasks and layout-shift sources in one reproducible report.
Field tools collect eligible experiences from real visitors. The Chrome UX Report (CrUX) aggregates opted-in Chrome users who meet its reporting criteria. Traffic spans devices, networks, locations, cache states and user behavior. Field data therefore answers whether actual users receive a good experience—not whether one test run is clean.
| Lab | Field | |
|---|---|---|
| Population | Simulated single load | Many eligible real visits |
| Conditions | Defined and repeatable | Naturally variable |
| Interaction | Usually limited | Real user interactions |
| Best use | Diagnosis and regression tests | Outcome monitoring |
| Blind spot | Traffic diversity and long sessions | Exact cause of a problem |
Lighthouse is lab tooling. CrUX and properly instrumented Real User Monitoring provide field measurements.
02 · Know the vocabulary
Map lab diagnostics to Core Web Vitals carefully
The current Core Web Vitals are Largest Contentful Paint (LCP), Interaction to Next Paint (INP) and Cumulative Layout Shift (CLS). Google recommends assessing “good” at the 75th percentile of page loads, segmented by mobile and desktop. The good thresholds are LCP at or below 2.5 seconds, INP at or below 200 milliseconds and CLS at or below 0.1.
Lighthouse reports LCP and CLS for its controlled load. It generally cannot produce representative INP because there is no population of real interactions. Total Blocking Time (TBT) is a useful lab diagnostic correlated with main-thread responsiveness, but it is not INP and should not be relabelled.
Time until the largest qualifying visible element renders.
Field latency across click, tap and keyboard interactions.
Unexpected layout-shift accumulation within session windows.
Blocking time from long main-thread tasks between FCP and TTI.
First Contentful Paint and Time to First Byte are supporting metrics, not Core Web Vitals. They can still reveal early rendering and server-response constraints.
03 · Investigate disagreement
Why a page can score 100 and fail users
A clean lab run may start with an empty cache but omit consent flows, personalization, logged-in tooling or interactions that happen later. Field visitors may use slower phones, arrive from farther away, encounter cold server caches, expand menus, add products or trigger delayed scripts. The lab page may also differ from the URL group represented in field reporting.
The reverse can happen too: a throttled lab profile might appear slow while a site’s actual audience uses fast devices and repeat caches. Neither result invalidates the other. Their disagreement is diagnostic.
A five-question conflict check
- Same URL? Page-level data, origin-level data and a single test URL are not equivalent.
- Same device class? Compare mobile with mobile and desktop with desktop.
- Same time window? CrUX is aggregated over a rolling period; a lab fix is immediate.
- Same page state? Consent, experiments, authentication and cache state alter execution.
- Same metric? TBT is not INP; one-run CLS is not a full visit’s layout-shift experience.
Use lab traces to form a cause hypothesis, then segment real-user data to verify it. For example, slow field LCP with good lab LCP may cluster on product pages, a geography, uncached responses or low-end devices.
04 · Read the dataset
CrUX is powerful, aggregated and sometimes unavailable
CrUX includes origins and URLs with enough eligible traffic to meet publication thresholds. Smaller sites and low-traffic pages may have no report. Absence is not a performance result. Public tools may fall back from page-level to origin-level data; check the displayed scope before attributing an origin result to one template.
Percentiles matter. Passing Core Web Vitals means the 75th-percentile experience is within the good threshold for each metric. An average can conceal a harmed tail of users. Segment mobile and desktop because their device and network populations differ substantially.
Theme-level field claims require caution
Live WordPress sites using the same theme can differ in hosting, builders, plugins, versions, content, optimization and audience. A theme cohort can reveal ecosystem outcomes, but it does not isolate theme code. Detection errors and survivorship should be disclosed. Treat this evidence as observational, not a controlled theme benchmark.
05 · Close the loop
A practical measurement workflow
- Before launch: set budgets and test representative templates in a controlled lab.
- During development: use Lighthouse traces and browser profiling to diagnose regressions.
- At release: verify production URLs from more than one region/device profile.
- After launch: instrument Web Vitals or use CrUX/Search Console when eligible.
- When field data fails: segment by template, device, version and geography; reproduce in lab.
- After fixing: confirm immediately in lab, then wait for sufficient field observations.
Keep versioned baselines instead of replacing old results. A performance history makes regressions, seasonality and delayed field improvements visible. Record deployment markers alongside the time series.
Reference desk
Primary sources
- web.dev: Web Vitals — definitions, thresholds and the 75th-percentile recommendation.
- Chrome UX Report documentation — eligibility, dimensions and aggregation.
- web.dev: Why lab and field data can be different.
- Lighthouse: Total Blocking Time.
- web.dev: Interaction to Next Paint.