01 · Test a proposition

Start with a narrow question

“What is the fastest theme?” is underspecified. Fast for which page type, configuration, device, network, host and metric? A repository-preview benchmark can compare vendor-maintained baselines. A representative-install study can compare how several stacks implement the same design. Field research can observe live outcomes. These are different studies.

Write a one-sentence hypothesis before installing anything. For example: “Under identical hosting, content and optimization, which candidate produces the lowest mobile LCP for this editorial homepage?” Define the primary metric in advance and name secondary diagnostic metrics. This reduces the temptation to switch winners after seeing the output.

Our FLT-WP-001 question

How do ten official WordPress.org preview pages behave under one repeatable simulated-mobile Lighthouse setup? It does not answer how completed production sites behave.

02 · Match the use case

Use a representative scenario—not just a blank install

Blank themes isolate baseline overhead, but visitors do not use blank pages. Build around a stable content fixture: the same heading hierarchy, local hero image, card count, navigation depth, footer, type scale and responsive breakpoints. For commerce, include product images, variation controls, cart fragments and the same WooCommerce version.

There are two defensible comparison designs:

DesignStrengthUse it for
Same exact markup/contentStrong isolation of styles and assetsTheme-level engineering comparison
Same visual/functional outcomeCaptures real implementation costBuyer decision support

The second is usually more useful to buyers but harder to standardize. Record every judgement: which native pattern was used, whether a companion plugin was required, and how much custom CSS was added. Do not silently strip a dependency from one theme while leaving another complete.

Plan more than one template

A homepage alone can conceal archive, article or product costs. Select two to four templates based on the target audience. Define a primary template before testing so the benchmark does not cherry-pick each theme’s best page.

03 · Reduce confounding

Control what you can and document what you cannot

Pin WordPress, PHP, theme, plugin and database versions. Use cloned infrastructure, identical caching rules and a resettable database fixture. Disable unrelated background work. Warm or clear caches according to a declared policy; either can be valid if it reflects the question. Record server region and test region.

  • Use the same host, resource limits, CDN and TLS behavior.
  • Use the same image files, dimensions and delivery settings.
  • Use the same font files and loading strategy.
  • Set a fixed viewport, network profile and CPU slowdown.
  • Test logged-out pages in a fresh browser context.
  • Preserve raw reports and the exact tested URL.

Optimization plugins can overwhelm theme differences. A baseline phase without them helps diagnose native behavior; a second production-configured phase shows attainable results. Never optimize one candidate more aggressively unless the study explicitly compares recommended stacks rather than themes.

04 · Expect variation

Collect repeated runs and disclose spread

Lab tests vary because scheduling, server response, cache layers and the test machine vary. Three runs with a median is a practical minimum for a small exploratory study; five or more better expose instability. Keep failed runs and define exclusion rules before collection. Do not rerun only the scores you dislike.

The median resists one extreme result, but it can hide variation. Publish per-run values or at least a range/interquartile spread alongside the median. Keep raw Lighthouse JSON so audits, trace evidence and tool settings remain inspectable.

Collect metrics that diagnose the system

MetricRoleInterpret carefully
LCPLoading experienceElement and resource discovery matter more than score alone
CLSVisual stabilityLab runs may miss shifts after interaction
TBTLab main-thread blocking proxyNot a substitute for field INP
TTFBInitial server responseHosting/cache signal more than theme-only signal
Bytes / requestsPayload and connection cluesComposition and priority also matter
DOM sizeMarkup complexity clueNo universal pass/fail count

05 · Make limitations visible

Report direct metrics before editorial conclusions

Publish the collection date, tool and browser versions, emulation settings, tested URL, theme versions, hosting setup, cache state, run count, aggregation rule and exclusions. Label a lab result as lab data every time it appears in a buying context. Never attach the result to a different configuration without saying so.

A ranking needs a predeclared sort metric. Do not create a composite score that blends Lighthouse, ratings, install counts and opinions unless the weighting has a defensible user model. Direct leaderboards—LCP, transfer size, requests—are easier to audit. Keep recommendation fit separate.

Describe uncertainty in plain language

State what could reverse the result: different content, a required builder, uncached server response, logged-in tools, WooCommerce, third-party consent or real traffic. A benchmark is evidence for a bounded question, not a permanent property of a theme.

06 · Copy this

A compact reusable test protocol

  1. Pre-register: question, candidates, primary metric, templates and exclusions.
  2. Freeze: software versions, content fixture, media, server and optimization.
  3. Verify: visual equivalence, required features, responsive states and no failed assets.
  4. Collect: fresh sessions, fixed device/network, at least three runs in rotated order.
  5. Aggregate: median plus variation; retain every valid raw report.
  6. Diagnose: identify LCP element, blocking tasks, shifts, critical requests and server delays.
  7. Report: conditions, versions, date, raw data, limitations and corrections path.
  8. Repeat: after material software or template changes; do not silently overwrite history.

For a real purchase decision, add a short editorial observation: implementation effort, accessibility, dependency cost and maintainability. Keep it visibly separate from the measurements.

Reference desk

Primary sources