Metricum Lab

Web Performance · Core Web Vitals · Measurement

How to Test Core Web Vitals Correctly: CrUX vs PageSpeed Insights vs Lighthouse vs RUM

CrUX, PageSpeed Insights, Lighthouse and RUM are not four competing ways to measure the same thing. I use field data to establish the user problem, then calibrate lab profiling to the affected audience, trace the cause, ship the smallest justified change and return to field data for validation.

Yurii Pekach Technical SEO & Web Performance ConsultantPublishedUpdated28 min read
Core Web Vitals measurement workflow showing CrUX field evidence, PageSpeed Insights, calibrated Lighthouse profiling and first-party RUM feeding a validation loop.

The short answer

Start with field evidence, not with the Lighthouse score. I use CrUX as the public baseline for whether eligible real Chrome users have a Core Web Vitals problem, then use first-party RUM or analytics to understand the affected cohort. Lighthouse and local DevTools profiling are controlled debugging environments: configure them to approximate those users, reproduce the symptom, inspect the trace, and validate the change in production. A green lab run is useful evidence, but it is not a substitute for field data.

≤2.5s · ≤200ms · ≤0.1

Good LCP · INP · CLS

Current good thresholds for LCP, INP and CLS respectively.

3

CrUX form factors

CrUX exposes phone, tablet and desktop—not exact device models—so audience calibration needs additional first-party context.

Measurement model

A Core Web Vitals test is not one test

The first decision is not which tool is best. It is which question you are trying to answer.

The mistake I see most often is putting a CrUX percentile, a PageSpeed Insights lab result, a local Lighthouse run and a RUM percentile into one mental bucket called the score. They describe different populations, time windows and levels of control. Field data tells you what users actually experienced; lab data gives you a controlled environment for reproduction and diagnosis. Google documents the same division: CrUX and PSI field data are real-user distributions, while Lighthouse is a simulated load designed for diagnostics.

What each Core Web Vitals measurement source is best at
SourceWhat it representsBest questionMain limitation
CrUXAggregated eligible Chrome user experiences; page or origin; rolling 28 daysDoes a real-user problem exist, and on which form factor?Public aggregate; Chrome-only eligible population; limited segmentation
PageSpeed InsightsCrUX field data plus a Lighthouse lab analysis in one interfaceCan I inspect field status and get an immediate lab diagnostic from one URL?The field and lab panels are different datasets and may legitimately disagree
Lighthouse / DevToolsOne controlled or simulated test environmentCan I reproduce a symptom and inspect why it happens?A run is not the distribution of your real users and can vary with conditions
First-party RUMMeasurements collected from visits to your own productWhich routes, cohorts, releases or environments are actually affected?Implementation, consent, sampling and metric methodology become your responsibility

Custom diagram

Use field evidence to choose the lab experiment

CrUX and RUM establish the production problem. PSI can surface both the field signal and a Lighthouse snapshot. Lighthouse and DevTools then turn that signal into a controlled experiment rather than replacing it.

Use field evidence to choose the lab experimentCrUX and RUM establish the production problem. PSI can surface both the field signal and a Lighthouse snapshot. Lighthouse and DevTools then turn that signal into a controlled experiment rather than replacing it.REAL USERS / FIELDCONTROLLED / LABCrUXp75 · 28 daysRUMcohorts · releasesPageSpeed InsightsCrUX field + Lighthouse labDefine the field problemmetric · grain · cohort · windowLighthouse + DevTools tracereproduce → explain → validate
Metricum Lab measurement model: field evidence defines the problem; lab tooling tests a falsifiable explanation; field measurement validates the release.

Field baseline

Start with CrUX—but read URL, origin and form factor before you act

A real-user number is only useful after you know exactly what population and aggregation it represents.

CrUX is where I start because it answers the most important first question: is the problem visible in real-user data? But I never copy the first number I see. PSI can show page-level CrUX when that URL has enough data, then fall back to origin-level data when it does not. Search Console goes one level further and groups similar URLs. Those are three different grains, and confusing them can send an investigation toward a page that was never actually measured on its own.

Build a CrUX baseline you can defend

  1. 1

    Confirm the grain

    In PageSpeed Insights, verify whether the field panel is showing the requested URL or an origin fallback. Do not attribute an origin-wide p75 directly to one page.

    PathPageSpeed Insights → Discover what your real users are experiencing → This URL / Origin

    Pass criterionThe investigation notes state URL-level or origin-level explicitly.

  2. 2

    Split mobile and desktop

    Read the affected form factor separately. The Core Web Vitals target is evaluated at p75 by device category, and Search Console also separates Mobile and Desktop views.

    PathPSI / CrUX → form factor · Search Console → Core Web Vitals → Mobile or Desktop

    Pass criterionThe failing metric and p75 are recorded for the relevant form factor, not only as an all-device aggregate.

  3. 3

    Record the collection window

    Save the first and last dates or at least the access date. A CrUX API or PSI field number represents a rolling 28-day period, not the performance of the current deployment alone.

    Pass criterionThe baseline includes the measurement window and verification date.

  4. 4

    Use the distribution, not just the badge

    Capture p75 together with the good / needs-improvement / poor distribution when available. A threshold classification hides how close the cohort is to a boundary.

    Pass criterionThe evidence package preserves the percentile and distribution rather than only a green/amber/red label.

  5. 5

    Query explicitly when the UI is ambiguous

    For reproducible checks, use the CrUX API with a specific URL or origin and formFactor. The API supports DESKTOP, PHONE and TABLET dimensions.

    Pass criterionThe query target and form factor are explicit and repeatable.

How to interpret common CrUX surfaces
SurfaceGranularityTime behaviorUse it for
PSI field panelPage when eligible; otherwise origin fallbackTrailing 28 days, updated dailyFast URL investigation—but verify the shown grain
CrUX APIPage or origin; optional form factorTrailing 28 days, updated dailyReproducible queries and automation
CrUX History API / CrUX VisPage or origin; form factor40 weekly points, each based on a 28-day collection periodTrend context without pretending each point is an independent week
Search Console Core Web VitalsGroups of similar URLs; Mobile / Desktop28-day field measurements surfaced as issue groupsPrioritize site-wide cohorts rather than diagnose one exact URL

PSI interpretation

PageSpeed Insights is two measurements in one report

The field panel and the Lighthouse panel are adjacent in the UI, but they are not observations from the same run.

PageSpeed Insights is useful precisely because it places two evidence layers together: the top field section is CrUX over the previous 28 days, while the lab diagnostics are produced by Lighthouse in a simulated environment. That convenience also creates a common interpretation bug: people compare a field p75 LCP with the current Lighthouse LCP as if one should reproduce the other exactly. It should not. One is a historical distribution across many real-user conditions; the other is a controlled single test.

Custom diagram

Read PageSpeed Insights as field evidence plus a lab experiment

The PSI interface can show a CrUX distribution and a Lighthouse diagnostic for the same requested URL, but their populations and time horizons differ.

Read PageSpeed Insights as field evidence plus a lab experimentThe PSI interface can show a CrUX distribution and a Lighthouse diagnostic for the same requested URL, but their populations and time horizons differ.PSI · Field dataCrUX · p75 · rolling 28 daysURL or Origin fallbackPSI · Lab diagnosticsLighthouse · simulated loadone test environmentDo not force the values to matchField: problem · Lab: mechanismCompare evidence roles, not one scoreOne interface ≠ one dataset
Do not average, subtract or directly reconcile the two values. Use the field layer to define the production problem and the lab layer to investigate a reproducible mechanism.

Three checks before you quote a PSI number

  • Is the field value for This URL or Origin?
  • Are you quoting the field p75 or the Lighthouse lab metric?
  • Which form factor, collection period and lab environment does the number represent?

Lab variability

Why the same page can show 2.0s, then 4.0s, then 3.5s LCP in Lighthouse

Repeated lab runs are samples from a test environment. They become evidence only when that environment is controlled and recorded.

In day-to-day profiling, I routinely expect repeated Lighthouse runs to move. A simple illustrative sequence is LCP 2.0 s → 4.0 s → 3.5 s on the same page without a code change. Those values are not client measurements; they illustrate why I never choose the fastest run and call the page fixed. Google lists network routing, different hardware, browser extensions, antivirus and other resource contention among common variability sources. PageSpeed Insights also notes that local network availability, client hardware and client resource contention can change results.

Illustrative Lighthouse run sequence — hypothetical values, not measured client data
RunLCPWhat it proves
12.0 sOnly that this test execution produced a good LCP under its conditions
24.0 sThe test environment or page behavior is variable enough to change the conclusion
33.5 sA single best run is not a stable baseline; investigate spread and conditions

Make lab measurements comparable before using them as evidence

  1. 1

    Freeze the tested state

    Use the same route, build, feature flags, consent state and user journey. If ads or experiments are nondeterministic, record that as a limitation or control them in a dedicated test environment.

    Pass criterionEach run starts from the same documented application state.

  2. 2

    Remove avoidable machine noise

    Use a clean browser profile without extensions where possible and close heavy local workloads that compete for CPU, memory or network resources.

    Pass criterionThe machine and browser state are stable enough that obvious local contention is not changing between runs.

  3. 3

    Choose cache state deliberately

    Do not mix cold-load and warm-cache runs in one comparison. Use the state that matches the question and keep it consistent across baseline and candidate build.

    Pass criterionEvery compared run uses the same cache/storage rule.

  4. 4

    Lock CPU and network settings

    Use one documented throttling profile for a comparison. The profile should approximate the affected audience rather than whichever preset happens to be convenient.

    Pass criterionCPU and network settings are recorded with the result.

  5. 5

    Run a small repeatable series

    For this workflow, my recommendation is five equivalent runs when time permits. Keep every result, use the median as the compact summary and inspect the spread instead of discarding inconvenient runs. Five is a practical comparison protocol for this guide, not a Google threshold or a claim about an official standard.

    Pass criterionBaseline and candidate have the same run count, conditions, median and visible spread.

  6. 6

    Save the trace for the slow cases

    The outlier is often diagnostically useful. If one run regresses, inspect the Performance trace and network activity rather than averaging the symptom away.

    PathChrome DevTools → Performance → Record and reload / Record

    Pass criterionAt least the representative slow trace can be tied to a concrete loading, scripting or rendering mechanism.

Audience calibration

Calibrate local profiling to your users, not to a generic 4G preset

A reproducible lab is useful; a reproducible lab that resembles the affected cohort is much more useful.

My normal sequence is field baseline → audience profile → local reproduction. I first determine whether the problem is primarily mobile or desktop in CrUX. Then I use first-party analytics or RUM to understand what those visitors actually use: browser mix, viewport/device class, route, geography or release cohort where available and privacy-safe. Only then do I choose CPU and network constraints. 4G is an example, not a universal preset. If it does not approximate your affected users, it is the wrong model for the investigation.

Build a local test profile from audience evidence

  1. 1

    Locate the failing field cohort

    Start with the CrUX or Search Console device category and metric that is actually failing at p75.

    Pass criterionYou can name the field cohort and the metric before opening DevTools.

  2. 2

    Add first-party audience context

    Use your analytics/RUM to learn which route types, screen/device classes, browsers or releases dominate the affected traffic. Do not assume CrUX can tell you an exact phone model; its public form-factor dimension is only phone, tablet or desktop.

    Pass criterionThe local profile is justified by first-party audience evidence, not a generic device persona.

  3. 3

    Use the Performance panel field overlay

    Enable CrUX Field metrics in Chrome DevTools Performance. The Live Metrics screen can show local metrics next to field metrics and its environment settings can recommend device, CPU and network throttling based on CrUX data.

    PathChrome DevTools → Performance → Live metrics → Field metrics / Environment settings

    Pass criterionLocal and field context are visible and the chosen environment is documented.

  4. 4

    Calibrate CPU carefully

    DevTools can calibrate custom CPU throttling presets for low- and mid-tier mobile approximation, but the slowdown is relative to your own computer. Chrome explicitly notes that a desktop CPU cannot truly simulate a mobile CPU architecture.

    PathPerformance → Capture settings → CPU → Calibrate

    Pass criterionThe report records the host machine and selected/calibrated CPU slowdown instead of calling it an exact device simulation.

  5. 5

    Reproduce the real journey

    For LCP, test the navigation state that produces the slow load. For INP and CLS, include the interactions or post-load behavior implicated by field data or RUM. A load-only Lighthouse audit cannot expose every interaction-driven issue.

    Pass criterionThe test journey can trigger the same class of symptom that motivated the investigation.

First-party field data

Use RUM when CrUX is too coarse for the decision

CrUX is an excellent public baseline, but first-party RUM can expose the cohorts and release context needed to act.

CrUX is not a census of every visitor. Its current eligibility rules include supported Chrome platforms and opted-in users; Chrome on iOS, Android WebView and other Chromium browsers such as Edge do not contribute. A trustworthy RUM implementation can therefore answer questions CrUX cannot: which route template regressed after release X, whether a browser cohort is worse, whether signed-in journeys differ, or which element is associated with a poor metric. That does not make RUM automatically more correct—its own consent, sampling and implementation choices can change the population too.

CrUX and first-party RUM solve different field-data problems
DimensionCrUXFirst-party RUM
PopulationEligible Chrome experiences under CrUX methodologyVisitors your implementation is permitted and able to measure
Time windowFixed rolling 28-day view in common CrUX surfacesCan be release-, day- or cohort-specific if sample size supports it
Device detailPhone / tablet / desktop public form factorCan add privacy-safe device, viewport or browser cohorts you control
Route / release contextPage/origin aggregates; Search Console URL groupsCan attach normalized route templates, build/release IDs and journeys
Debug attributionLimited public aggregatesCan collect targeted LCP/INP/CLS attribution with suitable instrumentation

STATICALLY VERIFIED · Minimal first-party Web Vitals attribution payload

import {onCLS, onINP, onLCP} from 'web-vitals/attribution';

function sendWebVital(metric) {
  const attribution = metric.attribution || {};
  const target =
    metric.name === 'LCP' ? attribution.target :
    metric.name === 'INP' ? attribution.interactionTarget :
    metric.name === 'CLS' ? attribution.largestShiftTarget :
    null;

  const payload = {
    name: metric.name,
    value: metric.value,
    rating: metric.rating,
    id: metric.id,
    path: location.pathname,
    target: target || null
  };

  navigator.sendBeacon(
    '/rum/web-vitals',
    new Blob([JSON.stringify(payload)], {type: 'application/json'})
  );
}

onLCP(sendWebVital);
onINP(sendWebVital);
onCLS(sendWebVital);

Syntax checked on 2026-09-08. Requires the web-vitals package and a site-specific /rum/web-vitals endpoint. Keep the payload intentionally small; normalize routes and remove identifiers or sensitive data before collection.

Metricum Lab method

My workflow: field signal → cohort → lab → trace → release → field

The measurement chain should make every transition explicit: what was observed, what is hypothesized, how it is reproduced and how the release will be validated.

I use CrUX as the first public baseline, not the final diagnostic. Once the field signal is clear, I identify the affected audience, make the lab look enough like that audience to reproduce the symptom, capture a trace, test one falsifiable explanation at a time, and then go back to production. This ordering prevents two common failures: optimizing a Lighthouse number that no user cohort actually has, and shipping a plausible fix that was never connected to the field regression.

Custom diagram

The field-to-lab-to-field validation loop

Each lab action is anchored to a production signal, and every release returns to production measurement rather than ending at a synthetic score.

The field-to-lab-to-field validation loopEach lab action is anchored to a production signal, and every release returns to production measurement rather than ending at a synthetic score.CrUXfield baselineRUMaffected cohortLocal profileCPU · network · journeyPerformance traceobserved mechanismMinimal changeone causal leverRelease + RUMproduction validationCrUX confirms over timerolling 28-day public evidence
Metricum Lab workflow: observe → segment → reproduce → trace → change → validate in RUM → confirm in CrUX over its rolling window.

Run the investigation as an evidence chain

  1. 1

    1 · Establish the field problem

    Record CrUX p75, distribution, grain, form factor and collection window for the failing metric.

    Pass criterionThe problem statement can be written without mentioning Lighthouse.

  2. 2

    2 · Identify the affected cohort

    Use RUM/analytics when available to narrow route, browser, viewport/device class, geography or release differences that CrUX cannot expose publicly.

    Pass criterionThe hypothesis names a cohort rather than an abstract average user.

  3. 3

    3 · Reproduce under controlled conditions

    Configure the browser environment to approximate that cohort and execute the same navigation or interaction journey repeatedly.

    Pass criterionThe symptom appears often enough that a trace can be captured under documented conditions.

  4. 4

    4 · Trace the mechanism

    Use the Performance panel to connect the metric to network, main-thread, rendering or layout evidence. Separate the observed trace from the hypothesis about root cause.

    Pass criterionA proposed change points to an observed mechanism in the representative trace.

  5. 5

    5 · Change one causal lever

    Prefer the smallest change that can falsify or support the current hypothesis. Avoid broad bundles of performance edits that make attribution impossible.

    Pass criterionThere is a clear before/after expectation for the same controlled test.

  6. 6

    6 · Validate in production RUM

    After release, compare the affected cohort against the baseline with the same metric definition and enough samples. Watch for regressions in the other Core Web Vitals as well.

    Pass criterionThe production distribution moves in the expected direction without a material new regression.

  7. 7

    7 · Let CrUX confirm the durable outcome

    Use CrUX/History data as slower public confirmation. Because common CrUX surfaces use a rolling 28-day window, a real release improvement should not be expected to replace the whole field distribution the next morning.

    Pass criterionThe longer-window field trend becomes consistent with the production RUM result, or the remaining disagreement is investigated.

Decision support

When CrUX, PSI, Lighthouse and RUM disagree, diagnose the mismatch

Disagreement is not a reason to pick your favorite tool. It is evidence that population, grain, timing or environment differs.

Core Web Vitals disagreement matrix
Observed patternMost likely interpretationNext action
CrUX poor · Lighthouse goodYour lab did not reproduce the affected field conditions or journeyCheck URL vs origin, form factor, RUM cohorts, throttling and interaction path; then trace a representative case
CrUX good · Lighthouse poorThe constrained lab case may be rarer or harsher than the current field p75Keep the lab issue if it exposes meaningful risk, but do not report it as a current field failure
PSI shows origin data · one URL looks poorThe field number belongs to the origin aggregate, not necessarily that URLQuery URL-level CrUX if eligible or use RUM before assigning root cause to the page
RUM poor · CrUX goodThe measured populations or metric implementations differ, or the issue is concentrated in a cohort CrUX smooths outAlign Chrome-only / device / 28-day comparisons first; then audit RUM instrumentation and cohort mix
Search Console group poor · sampled URL goodA URL group can contain similar pages with a shared status while one example is an outlierInvestigate the template/cohort and additional group members rather than treating one URL test as a group verdict
Lighthouse baseline and candidate overlap heavilyThe apparent change is smaller than test variabilityIncrease control/repetition, inspect traces and do not claim improvement from the best individual run

For CLS specifically, lab/field gaps are especially common when shifts happen after user interaction. The same principle is covered in the Metricum Lab CLS debugging guide: reproduce the user journey, capture evidence and validate in the field instead of stopping at a load-only audit.

Release gate

A green Lighthouse run is not my definition of fixed

A fix is complete when the mechanism improves under controlled conditions and the affected production distribution moves in the same direction.

The lab is where I want fast feedback. The field is where I want the final verdict. A strong performance change should survive both: the same controlled test should improve for an explainable reason, then the production cohort should move after release. CrUX is deliberately slower because its common views are 28-day rolling aggregates, so I use RUM for earlier release feedback and CrUX for durable public confirmation.

Release validation contract

  1. 1

    Controlled comparison improves

    Repeat the same lab contract on baseline and candidate build; compare median, spread and traces rather than one best run.

    Pass criterionThe candidate improves under equivalent conditions and the mechanism is visible in trace evidence.

  2. 2

    No Core Web Vitals trade-off is hidden

    Check LCP, INP and CLS together. A change that improves one metric by materially worsening another is not a clean win.

    Pass criterionNo new material regression appears in the other Core Web Vitals.

  3. 3

    Production RUM moves for the affected cohort

    Compare the same route/device/browser/release cohort and metric definition after enough real-user samples accumulate.

    Pass criterionThe field distribution shifts in the expected direction with sufficient sample context.

  4. 4

    CrUX catches up over the rolling window

    Track PSI/CrUX or the History API without expecting an overnight reset. Historical weekly points still represent overlapping 28-day collection periods.

    Pass criterionThe longer public field trend confirms the improvement or identifies a remaining cohort mismatch.

  5. 5

    Keep the evidence package

    Store the field baseline, environment contract, representative traces, code change, release ID and post-release field comparison. This turns the fix into a regression test rather than a one-off success.

    Pass criterionA future regression can be compared against the same evidence chain.

The practical principle is simple: use CrUX to decide whether you have a real-user problem; use RUM to understand who has it; use Lighthouse and DevTools to reproduce and explain it; then return to field data to prove the change worked. If the issue is visual instability, continue with the CLS debugging workflow. The same field-first logic will also underpin our LCP diagnosis workflow.

Primary sources and documentation

Sources

Every changing search, browser, interface or technical-behavior claim in this guide is tied to a current primary source.

  1. web.dev / Chrome team Web Vitals (opens in a new tab)Current Core Web Vitals definitions, thresholds and 75th-percentile guidance.
  2. Google for Developers About PageSpeed Insights (opens in a new tab)Primary reference for PSI field versus lab data, 28-day CrUX window, URL-to-origin fallback and run variability.
  3. Chrome for Developers CrUX Tools (opens in a new tab)Current comparison of CrUX API, History API, PSI, Search Console and BigQuery time windows and granularity.
  4. Chrome for Developers CrUX methodology (opens in a new tab)Eligibility, page/origin aggregation and supported-user limitations for CrUX.
  5. Chrome for Developers CrUX dimensions (opens in a new tab)Form-factor definitions: phone, tablet and desktop.
  6. Chrome for Developers CrUX API (opens in a new tab)Programmatic page/origin queries, p75 metrics and form-factor filtering.
  7. Google Search Console Help Core Web Vitals report (opens in a new tab)Current Search Console grouping, device views, CrUX source and 28-day p75 interpretation.
  8. Chrome for Developers Performance features reference (opens in a new tab)Live Metrics, CrUX field overlay, field-aware environment guidance and CPU/network throttling behavior.
  9. Chrome for Developers Lighthouse performance scoring (opens in a new tab)Primary documentation on Lighthouse metric and score variability.
  10. web.dev / Chrome team Why lab and field data can be different (and what to do about it) (opens in a new tab)Explains controlled lab conditions versus distributions of real-user experiences.
  11. web.dev / Chrome team Why is CrUX data different from my RUM data? (opens in a new tab)Current comparison of CrUX and first-party RUM populations, aggregation and measurement differences.
  12. web.dev / Chrome team Best practices for measuring Web Vitals in the field (opens in a new tab)Why production field measurement is required to validate whether performance changes work for users.
  13. web.dev / Chrome team Debug performance in the field (opens in a new tab)Field-to-lab debugging model and real-user attribution guidance.
  14. GoogleChrome / GitHub web-vitals library README (opens in a new tab)Current library usage and attribution build for LCP, INP and CLS diagnostics.

Need help diagnosing and implementing the fix?

Turn Core Web Vitals measurements into a reproducible performance backlog

Metricum Lab can connect CrUX and first-party field signals to controlled browser traces, implementation hypotheses and post-release validation—without optimizing for a Lighthouse number in isolation.

Explore Core Web Vitals Optimization