01 / Principles
What a fair theme test must protect
Theme performance is easy to oversimplify. A minimal theme can look unbeatable on an empty page, while a polished demo can be penalized for carrying the very design and content a buyer wants. Hosting, cache warmth, traffic, plugins, image choices and measurement variance can overwhelm a small theme-level difference. Our method therefore uses five guardrails.
- Compare like with like. Themes in the same table run against the same platform, preview service, test tool, form factor and collection window.
- Repeat and preserve. We run three times, report the median for each metric and publish the individual JSON reports in the raw archive.
- Label the evidence. Lab measurements, field data, official facts and editorial interpretation are never silently merged.
- Prefer representative scenarios. A content-rich business page, article and store route are more useful than an empty install. When that is not available, the result is explicitly named a preview baseline.
- Do not reward unusable minimalism. A theme that needs a builder to become useful is evaluated as a dependency shell, not crowned on its empty-theme metric.
02 / Evidence labels
Every datum has a provenance class
Measured by Fast Loading Themes. A dated result from a documented run we executed. The label always names the scenario. “Measured preview” must not be rephrased as “production tested.”
Official repository. Metadata retrieved from WordPress.org: version, update date, install bucket, rating, requirements, tags and screenshot. Official does not mean independently verified.
Editorial analysis. Our interpretation of fit, workflow, design range, dependency risk and trade-offs. It can use measured and official inputs, but it is not a measured fact.
Not available. We could not obtain a reliable, comparable datum. A visible gap is better than an inferred number.
Vendor claims would receive a separate “Vendor claim” label if quoted. No current profile relies on a vendor speed claim.
03 / Current study
Study FLT-WP-001: official preview baseline
The launch dataset measures ten themes on wp-themes.com, the official WordPress.org preview service. These pages use the service’s shared preview content and infrastructure. That reduces cross-host variation, but it does not produce a completed representative business, publication or WooCommerce site.
- Collection date
- 16 August 2026
- Theme count
- 10
- Runs
- 3 per theme / 30 total
- Tool
- Lighthouse 12.8.2
- Browser
- Headless Chrome 151.0
- Form factor
- Simulated mobile
- Viewport
- 412 × 823 at 1.75×
- Storage
- Reset between runs
Throttling configuration
| Setting | Value | Why it matters |
|---|---|---|
| Method | Simulated | Lighthouse models the trace rather than applying the network profile live. |
| RTT | 150 ms | Models mobile network latency. |
| Throughput | 1,638.4 Kbps | Sets the Lantern simulation input. |
| Download throughput | 1,474.56 Kbps | Effective downstream throttle setting recorded in the report. |
| Upload throughput | 675 Kbps | Effective upstream setting. |
| CPU slowdown | 4× | Models a less capable mobile processor. |
Theme selection
The sample is purposeful, not exhaustive. It covers popular multipurpose themes (Astra, GeneratePress, Kadence, Blocksy, Neve and OceanWP), a native block-theme control (Twenty Twenty-Five), an Elementor shell (Hello Elementor), and two WooCommerce-oriented options (Storefront and Botiga). Popularity, workflow variety, official metadata availability and direct user decision relevance drove inclusion.
Run and aggregation procedure
- Resolve the official preview URL over HTTPS.
- Launch a fresh Lighthouse navigation with mobile form factor and storage reset.
- Collect performance category results and network records.
- Repeat two more times in the same collection window.
- Take the median of each individual metric, rather than selecting one whole “best” run.
- Store tested theme version and date next to the aggregate.
We do not discard slow runs merely because they look inconvenient. A failed run would be reported and repeated with the failure retained in the study log; this collection completed all 30 navigations.
Per-run variation
Aggregate cards can conceal normal run noise. The table exposes all three captured values as score / LCP / transferred bytes. “LCP spread” is the slowest minus the fastest run; it is descriptive, not a confidence interval.
| Theme | Run 1 | Run 2 | Run 3 | Median LCP | LCP spread |
|---|---|---|---|---|---|
| GeneratePress | 100 / 1.50s / 25 KB | 100 / 0.93s / 25 KB | 100 / 0.93s / 25 KB | 0.93s | 568ms |
| Astra | 95 / 2.77s / 148 KB | 91 / 3.21s / 148 KB | 91 / 3.21s / 148 KB | 3.21s | 440ms |
| Kadence | 100 / 1.08s / 43 KB | 100 / 1.08s / 43 KB | 100 / 1.08s / 43 KB | 1.08s | 0ms |
| Blocksy | 100 / 1.11s / 45 KB | 100 / 1.08s / 45 KB | 100 / 1.08s / 45 KB | 1.08s | 27ms |
| Neve | 88 / 3.11s / 310 KB | 87 / 3.21s / 310 KB | 88 / 3.04s / 310 KB | 3.11s | 168ms |
| Twenty Twenty-Five | 95 / 2.97s / 120 KB | 95 / 2.97s / 120 KB | 97 / 2.66s / 120 KB | 2.97s | 309ms |
| Hello Elementor | 100 / 0.81s / 18 KB | 100 / 0.85s / 18 KB | 100 / 0.81s / 18 KB | 0.81s | 44ms |
| Storefront | 89 / 3.60s / 166 KB | 89 / 3.57s / 166 KB | 89 / 3.59s / 166 KB | 3.59s | 28ms |
| Botiga | 100 / 1.23s / 91 KB | 100 / 1.23s / 91 KB | 100 / 1.23s / 91 KB | 1.23s | 0ms |
| OceanWP | 99 / 1.76s / 340 KB | 99 / 1.74s / 340 KB | 99 / 1.72s / 340 KB | 1.74s | 44ms |
GeneratePress shows the widest LCP spread in this small collection: 568 ms between its first and fastest runs. That is why differences of only a few hundred milliseconds should not override stack and workflow fit. All 30 raw Lighthouse JSON reports are retained in the study archive.
Next: representative installations
FLT-WP-002 is intended to compare a narrower theme set on matched WordPress installations with production-like content, pinned versions, required companion plugins, identical media/fonts and at least one commerce route. The protocol will be finalized before collection, use rotated repeat runs and publish variation. No result from that future study is claimed on this site today.
04 / Metrics
What each number can—and cannot—say
Lab estimate of when the largest visible content element renders. It is not a CrUX field percentile.
When the browser first paints text, an image or another content element.
Unexpected visual movement observed during the lab navigation. A zero lab value does not prove all templates are stable.
Lab measure of long main-thread tasks after FCP. It is a useful diagnostic, not a substitute for field INP.
Preview-host response timing. It mostly describes the shared service and run conditions, not theme PHP efficiency in isolation.
Page total observed by Lighthouse, including content assets. It is not the install ZIP size.
Count of requests in the trace. Fewer can reduce overhead, but caching and resource size still matter.
Element count on the tested page. Depth and interaction cost matter in addition to the count.
Core Web Vitals are field metrics judged at the 75th percentile. Google’s current “good” thresholds are LCP ≤ 2.5 s, INP ≤ 200 ms and CLS ≤ 0.1. A Lighthouse run is not field compliance.
Source: Google’s Web Vitals guidance and threshold methodology.
05 / Recommendation rules
Fit and evidence remain separate
We do not calculate a universal Fast Loading Themes score. The Lighthouse performance score is shown as produced by Lighthouse, but rankings use named raw metrics such as LCP or transferred bytes. This avoids burying weighting choices inside a proprietary number.
Editorial recommendations consider:
- site and content type;
- editing workflow and builder requirement;
- design quality and how much work is needed to reach an acceptable result;
- dependency surface and likely companion plugins;
- current maintenance signals from the repository;
- measured baseline results, with their scenario limits;
- commercial constraints only after usefulness and fit.
The finder uses visible fit points to order candidates. Measured metrics are displayed beside the match result; they are not secretly folded into that editorial score. A hard mismatch—such as selecting full-site editing for a classic theme—carries a larger penalty than a small LCP difference.
Affiliate independence
The launch dataset links to WordPress.org and contains no affiliate links. If affiliate relationships are added later, they will not determine sample inclusion, ranking or outcome. A visible disclosure will appear beside commercial links and on the disclosure page.
06 / Sources
Current data provenance
| Data | Source | Retrieved | Refresh rule |
|---|---|---|---|
| Theme metadata and screenshots | WordPress.org Themes API 1.2 (example source record) | 16 Aug 2026 | Before each published study |
| Preview measurements | Fast Loading Themes Lighthouse run against wp-themes.com | 16 Aug 2026 | After meaningful theme updates or quarterly |
| Core Web Vitals definitions | web.dev | 16 Aug 2026 | Review quarterly |
| Field-data definition | Chrome UX Report documentation | 16 Aug 2026 | Review quarterly |
| Fit and trade-offs | Fast Loading Themes editorial analysis | 16 Aug 2026 | Review with profile updates |
07 / Limitations
Known constraints in this dataset
- Preview content is not a full representative site. It offers cross-theme control but limited real-world depth.
- No WooCommerce route was measured. Store-oriented profiles clearly state this gap.
- No field data is attributed to individual WordPress themes. Reliable detection and sufficient CrUX coverage are not available in this release.
- INP is unavailable. A scripted lab interaction would still not be field INP; we show TBT as a diagnostic and label it accordingly.
- Third-party and font analysis is scenario-bound. The preview does not activate the starter sites or production integrations many users will install.
- Lighthouse variance remains. Three medians reduce—but do not eliminate—environmental noise. Differences of a few hundred milliseconds should not dominate a decision.
- Repository install counts are buckets. “500K+” is not an exact current count.
- Pricing is omitted. We will not publish pricing until a repeatable freshness and region/tax disclaimer process is implemented.
08 / Updates
Freshness, corrections and version history
Each measured profile stores test date, theme version, scenario and run count. A new test adds a history row; it does not silently rewrite the old result. Official metadata is re-retrieved before a study. Profiles older than 90 days or superseded by a major version should be visibly flagged before appearing in a current leaderboard.
Factual errors are corrected promptly and recorded in a study changelog when they affect a metric, rank or recommendation. Send a reproducible correction through the contact page. Theme vendors may point out errors but cannot pay to remove a limitation or alter a ranking.
Published FLT-WP-001 with ten themes, thirty Lighthouse runs and four evidence labels.