Open methods / revision 1.0

Trust the method—or challenge it.

A score without conditions is marketing. This page records what we ran, what we did not run, and exactly how each evidence type is allowed to influence a recommendation.

LAST REVISED16 Aug 2026Study ID: FLT-WP-001Use the testing playbook →

01 / Principles

What a fair theme test must protect

Theme performance is easy to oversimplify. A minimal theme can look unbeatable on an empty page, while a polished demo can be penalized for carrying the very design and content a buyer wants. Hosting, cache warmth, traffic, plugins, image choices and measurement variance can overwhelm a small theme-level difference. Our method therefore uses five guardrails.

  1. Compare like with like. Themes in the same table run against the same platform, preview service, test tool, form factor and collection window.
  2. Repeat and preserve. We run three times, report the median for each metric and publish the individual JSON reports in the raw archive.
  3. Label the evidence. Lab measurements, field data, official facts and editorial interpretation are never silently merged.
  4. Prefer representative scenarios. A content-rich business page, article and store route are more useful than an empty install. When that is not available, the result is explicitly named a preview baseline.
  5. Do not reward unusable minimalism. A theme that needs a builder to become useful is evaluated as a dependency shell, not crowned on its empty-theme metric.

02 / Evidence labels

Every datum has a provenance class

Measured

Measured by Fast Loading Themes. A dated result from a documented run we executed. The label always names the scenario. “Measured preview” must not be rephrased as “production tested.”

Official repository

Official repository. Metadata retrieved from WordPress.org: version, update date, install bucket, rating, requirements, tags and screenshot. Official does not mean independently verified.

Editorial

Editorial analysis. Our interpretation of fit, workflow, design range, dependency risk and trade-offs. It can use measured and official inputs, but it is not a measured fact.

Not available

Not available. We could not obtain a reliable, comparable datum. A visible gap is better than an inferred number.

Vendor claims would receive a separate “Vendor claim” label if quoted. No current profile relies on a vendor speed claim.

03 / Current study

Study FLT-WP-001: official preview baseline

The launch dataset measures ten themes on wp-themes.com, the official WordPress.org preview service. These pages use the service’s shared preview content and infrastructure. That reduces cross-host variation, but it does not produce a completed representative business, publication or WooCommerce site.

Collection date
16 August 2026
Theme count
10
Runs
3 per theme / 30 total
Tool
Lighthouse 12.8.2
Browser
Headless Chrome 151.0
Form factor
Simulated mobile
Viewport
412 × 823 at 1.75×
Storage
Reset between runs

Throttling configuration

SettingValueWhy it matters
MethodSimulatedLighthouse models the trace rather than applying the network profile live.
RTT150 msModels mobile network latency.
Throughput1,638.4 KbpsSets the Lantern simulation input.
Download throughput1,474.56 KbpsEffective downstream throttle setting recorded in the report.
Upload throughput675 KbpsEffective upstream setting.
CPU slowdown4×Models a less capable mobile processor.

Theme selection

The sample is purposeful, not exhaustive. It covers popular multipurpose themes (Astra, GeneratePress, Kadence, Blocksy, Neve and OceanWP), a native block-theme control (Twenty Twenty-Five), an Elementor shell (Hello Elementor), and two WooCommerce-oriented options (Storefront and Botiga). Popularity, workflow variety, official metadata availability and direct user decision relevance drove inclusion.

Run and aggregation procedure

  1. Resolve the official preview URL over HTTPS.
  2. Launch a fresh Lighthouse navigation with mobile form factor and storage reset.
  3. Collect performance category results and network records.
  4. Repeat two more times in the same collection window.
  5. Take the median of each individual metric, rather than selecting one whole “best” run.
  6. Store tested theme version and date next to the aggregate.

We do not discard slow runs merely because they look inconvenient. A failed run would be reported and repeated with the failure retained in the study log; this collection completed all 30 navigations.

Per-run variation

Aggregate cards can conceal normal run noise. The table exposes all three captured values as score / LCP / transferred bytes. “LCP spread” is the slowest minus the fastest run; it is descriptive, not a confidence interval.

ThemeRun 1Run 2Run 3Median LCPLCP spread
GeneratePress100 / 1.50s / 25 KB100 / 0.93s / 25 KB100 / 0.93s / 25 KB0.93s568ms
Astra95 / 2.77s / 148 KB91 / 3.21s / 148 KB91 / 3.21s / 148 KB3.21s440ms
Kadence100 / 1.08s / 43 KB100 / 1.08s / 43 KB100 / 1.08s / 43 KB1.08s0ms
Blocksy100 / 1.11s / 45 KB100 / 1.08s / 45 KB100 / 1.08s / 45 KB1.08s27ms
Neve88 / 3.11s / 310 KB87 / 3.21s / 310 KB88 / 3.04s / 310 KB3.11s168ms
Twenty Twenty-Five95 / 2.97s / 120 KB95 / 2.97s / 120 KB97 / 2.66s / 120 KB2.97s309ms
Hello Elementor100 / 0.81s / 18 KB100 / 0.85s / 18 KB100 / 0.81s / 18 KB0.81s44ms
Storefront89 / 3.60s / 166 KB89 / 3.57s / 166 KB89 / 3.59s / 166 KB3.59s28ms
Botiga100 / 1.23s / 91 KB100 / 1.23s / 91 KB100 / 1.23s / 91 KB1.23s0ms
OceanWP99 / 1.76s / 340 KB99 / 1.74s / 340 KB99 / 1.72s / 340 KB1.74s44ms

GeneratePress shows the widest LCP spread in this small collection: 568 ms between its first and fastest runs. That is why differences of only a few hundred milliseconds should not override stack and workflow fit. All 30 raw Lighthouse JSON reports are retained in the study archive.

PLANNED / NOT YET MEASURED

Next: representative installations

FLT-WP-002 is intended to compare a narrower theme set on matched WordPress installations with production-like content, pinned versions, required companion plugins, identical media/fonts and at least one commerce route. The protocol will be finalized before collection, use rotated repeat runs and publish variation. No result from that future study is claimed on this site today.

04 / Metrics

What each number can—and cannot—say

LCPLargest Contentful Paint

Lab estimate of when the largest visible content element renders. It is not a CrUX field percentile.

FCPFirst Contentful Paint

When the browser first paints text, an image or another content element.

CLSCumulative Layout Shift

Unexpected visual movement observed during the lab navigation. A zero lab value does not prove all templates are stable.

TBTTotal Blocking Time

Lab measure of long main-thread tasks after FCP. It is a useful diagnostic, not a substitute for field INP.

TTFBTime to First Byte

Preview-host response timing. It mostly describes the shared service and run conditions, not theme PHP efficiency in isolation.

TransferCompressed network bytes

Page total observed by Lighthouse, including content assets. It is not the install ZIP size.

RequestsNetwork records

Count of requests in the trace. Fewer can reduce overhead, but caching and resource size still matter.

DOMRendered elements

Element count on the tested page. Depth and interaction cost matter in addition to the count.

Core Web Vitals are field metrics judged at the 75th percentile. Google’s current “good” thresholds are LCP ≤ 2.5 s, INP ≤ 200 ms and CLS ≤ 0.1. A Lighthouse run is not field compliance.

Source: Google’s Web Vitals guidance and threshold methodology.

05 / Recommendation rules

Fit and evidence remain separate

We do not calculate a universal Fast Loading Themes score. The Lighthouse performance score is shown as produced by Lighthouse, but rankings use named raw metrics such as LCP or transferred bytes. This avoids burying weighting choices inside a proprietary number.

Editorial recommendations consider:

  • site and content type;
  • editing workflow and builder requirement;
  • design quality and how much work is needed to reach an acceptable result;
  • dependency surface and likely companion plugins;
  • current maintenance signals from the repository;
  • measured baseline results, with their scenario limits;
  • commercial constraints only after usefulness and fit.

The finder uses visible fit points to order candidates. Measured metrics are displayed beside the match result; they are not secretly folded into that editorial score. A hard mismatch—such as selecting full-site editing for a classic theme—carries a larger penalty than a small LCP difference.

Affiliate independence

The launch dataset links to WordPress.org and contains no affiliate links. If affiliate relationships are added later, they will not determine sample inclusion, ranking or outcome. A visible disclosure will appear beside commercial links and on the disclosure page.

06 / Sources

Current data provenance

DataSourceRetrievedRefresh rule
Theme metadata and screenshotsWordPress.org Themes API 1.2 (example source record)16 Aug 2026Before each published study
Preview measurementsFast Loading Themes Lighthouse run against wp-themes.com16 Aug 2026After meaningful theme updates or quarterly
Core Web Vitals definitionsweb.dev16 Aug 2026Review quarterly
Field-data definitionChrome UX Report documentation16 Aug 2026Review quarterly
Fit and trade-offsFast Loading Themes editorial analysis16 Aug 2026Review with profile updates

07 / Limitations

Known constraints in this dataset

  • Preview content is not a full representative site. It offers cross-theme control but limited real-world depth.
  • No WooCommerce route was measured. Store-oriented profiles clearly state this gap.
  • No field data is attributed to individual WordPress themes. Reliable detection and sufficient CrUX coverage are not available in this release.
  • INP is unavailable. A scripted lab interaction would still not be field INP; we show TBT as a diagnostic and label it accordingly.
  • Third-party and font analysis is scenario-bound. The preview does not activate the starter sites or production integrations many users will install.
  • Lighthouse variance remains. Three medians reduce—but do not eliminate—environmental noise. Differences of a few hundred milliseconds should not dominate a decision.
  • Repository install counts are buckets. “500K+” is not an exact current count.
  • Pricing is omitted. We will not publish pricing until a repeatable freshness and region/tax disclaimer process is implemented.

08 / Updates

Freshness, corrections and version history

Each measured profile stores test date, theme version, scenario and run count. A new test adds a history row; it does not silently rewrite the old result. Official metadata is re-retrieved before a study. Profiles older than 90 days or superseded by a major version should be visibly flagged before appearing in a current leaderboard.

Factual errors are corrected promptly and recorded in a study changelog when they affect a metric, rank or recommendation. Send a reproducible correction through the contact page. Theme vendors may point out errors but cannot pay to remove a limitation or alter a ranking.

16 Aug 2026
Method revision 1.0

Published FLT-WP-001 with ten themes, thirty Lighthouse runs and four evidence labels.