Skip to main content

Website Performance Monitoring: Core Web Vitals, Page Load Time, and Performance Regressions

· 20 min read
Avi Stramer
Founder, Testable

A website does not have to be offline to be unavailable in practice. A page that takes ten seconds to reveal its main content, ignores a tap while JavaScript runs, or moves the checkout button as someone tries to click it is technically up but functionally broken.

Website performance monitoring turns those experiences into measurable signals. The challenge is choosing signals that represent users, interpreting them in the right context, and alerting on meaningful regressions without treating every millisecond of normal variation as an incident.

This guide explains how to combine Core Web Vitals, supporting page-load metrics, synthetic browser checks, and real-user data into a practical monitoring strategy.

Website performance is not one number​

“How fast is the site?” sounds like a simple question, but a page loads and responds in stages:

  1. The browser resolves the hostname and establishes a connection.
  2. The server begins returning the HTML document.
  3. The browser discovers and downloads styles, scripts, fonts, images, and data.
  4. Something visible appears.
  5. The primary content becomes visible.
  6. The page becomes responsive to interactions.
  7. Late content, ads, banners, and application state may continue changing the layout.

Different metrics describe different parts of that experience. A fast server response does not guarantee that the main content renders quickly. A good initial render does not guarantee that the page responds when clicked. A single “page load time” can hide all of those distinctions.

Use a small set of metrics with defined roles:

  • User-experience metrics describe loading, responsiveness, and visual stability.
  • Diagnostic metrics help explain which stage is slow.
  • Operational metrics track complete journeys, availability, errors, and dependencies.

The goal is not to collect every metric a browser exposes. It is to detect when an important experience gets worse and retain enough context to find the cause.

What are the Core Web Vitals?​

Core Web Vitals are Google's stable, user-centered metrics for three aspects of page experience: loading, interactivity, and visual stability. The current set is:

MetricWhat it representsGoodNeeds improvementPoor
Largest Contentful Paint (LCP)Loading of the main visible content≤ 2.5 s> 2.5 s and ≤ 4.0 s> 4.0 s
Interaction to Next Paint (INP)Responsiveness to user interactions≤ 200 ms> 200 ms and ≤ 500 ms> 500 ms
Cumulative Layout Shift (CLS)Unexpected visual movement≤ 0.1> 0.1 and ≤ 0.25> 0.25

These thresholds apply to the 75th percentile of page visits, evaluated separately for mobile and desktop. A page provides a good Core Web Vitals experience when all three metrics are in the good range at that percentile. Google's current Web Vitals guidance documents the metrics, thresholds, and percentile requirement.

The percentile matters. An average can look healthy while a significant group of visitors has a poor experience. At the 75th percentile, at least three out of four measured visits must be at or below the threshold.

Core Web Vitals are not a complete monitoring system. They do not tell you whether login works, an API is returning errors, or checkout can complete. They are focused signals for page experience and should sit alongside availability and functional monitoring.

Largest Contentful Paint: when does the main content appear?​

Largest Contentful Paint measures when the largest eligible image, video frame, or text block visible in the viewport renders. It is intended to approximate the moment the page's primary content becomes useful.

LCP is more meaningful than the browser's load event for many modern sites. A page can finish loading small technical resources while its hero image or main application content is still missing. The reverse can also happen: useful content may be visible while non-critical resources continue downloading.

When LCP gets worse, investigate its major components:

  • Server and connection delay: Redirects, DNS, TLS, cache misses, backend work, and network latency contribute before the document arrives.
  • Resource discovery delay: The browser may not discover the LCP image until JavaScript runs or a stylesheet is downloaded.
  • Resource loading time: A large image, slow origin, or missing CDN cache can delay the asset.
  • Render delay: The resource may be available but blocked by CSS, fonts, main-thread work, or client rendering.

Record the LCP element as well as the value when possible. A regression is much easier to diagnose when you know that the LCP changed from a text heading to a newly introduced hero image.

Google's LCP reference notes that field measurements include navigation delays such as redirects, connection setup, and Time to First Byte. That makes LCP an end-to-end user signal rather than merely an image-download timer.

Interaction to Next Paint: does the page respond promptly?​

Interaction to Next Paint measures page responsiveness throughout a visit. It observes clicks, taps, and keyboard interactions, then reports a representative value based on the slowest interactions in the page lifecycle.

An interaction includes:

  • The delay before the event handler begins.
  • The time spent running event handlers.
  • The presentation delay before the browser paints the next frame.

A page can have excellent LCP and still have poor INP. This often happens when the browser's main thread is busy parsing or executing JavaScript, rendering a large update, or handling an expensive event. The button is visible, but the interface appears to ignore the user.

INP is especially important for long-lived applications. Problems may appear after initial load when a user opens a complex menu, filters a large table, edits a document, or changes routes within a single-page application.

Standard Lighthouse runs do not generate real user interactions and therefore cannot directly measure field INP. Lighthouse reports Total Blocking Time (TBT) as a lab diagnostic that can reveal main-thread work likely to harm responsiveness. A scripted browser journey can also exercise representative interactions, but it still covers only the actions and environment you define.

Use field INP to understand actual user responsiveness. Use TBT, long-task data, browser traces, and repeatable interactions to investigate and prevent regressions.

Cumulative Layout Shift: does the page stay visually stable?​

Cumulative Layout Shift measures unexpected movement of visible content. Common causes include:

  • Images, videos, ads, or embeds without reserved dimensions.
  • A cookie banner or promotion inserted above existing content.
  • Web fonts that significantly change text dimensions.
  • Client-rendered content appearing after the initial layout.
  • Animations that trigger layout rather than using transforms.

CLS is unitless. It combines how much of the viewport was affected with how far elements moved, grouped into windows of layout shifts. A score of 0 means no unexpected movement; larger scores represent greater instability.

A short synthetic page-load test may miss layout shifts that occur later in a real session. A user might scroll, open a menu, wait for an advertisement to refresh, or keep a dashboard open while new data arrives. Field measurement covers the page lifecycle and is therefore essential for understanding post-load CLS.

When monitoring CLS synthetically, allow enough observation time for known late content and exercise interactions that reveal dynamic sections. Do not extend every check indefinitely; use representative scripts and rely on field data for the long tail of real behavior.

Supporting metrics that explain the experience​

Core Web Vitals answer important user-centered questions, but supporting metrics make changes easier to interpret.

Time to First Byte (TTFB)​

TTFB measures the time from navigation until the browser receives the first byte of the main document. It includes redirects and connection setup in addition to server processing.

It is a useful early signal for:

  • Slow application or database work.
  • Origin saturation.
  • CDN cache misses.
  • Geographic latency.
  • DNS, connection, and TLS overhead.
  • Unexpected redirect chains.

TTFB is not a Core Web Vital. Google's TTFB guidance suggests 0.8 seconds or less as a rough target for most sites, but emphasizes that it should be interpreted in the context of later user-facing metrics. A server-rendered page can have a somewhat slower TTFB yet render efficiently, while a client-rendered application may need a very fast document response before substantial browser work even begins.

First Contentful Paint (FCP)​

FCP measures when the browser first renders content from the document, such as text, an image, or a canvas. It answers “When did the blank screen begin to change?”

The gap between FCP and LCP can be revealing. A fast FCP with a slow LCP often means a shell or placeholder appeared quickly while the primary content was delayed. A slow FCP points earlier in the chain toward server response, render-blocking resources, fonts, or initial JavaScript.

Total Blocking Time (TBT)​

TBT is a lab metric that totals the blocking portions of long main-thread tasks between FCP and Time to Interactive. It is useful for identifying JavaScript and rendering work that may lead to poor responsiveness.

Treat TBT as a diagnostic and regression signal, not as a substitute for field INP. It measures a particular lab window and does not observe the variety of interactions users perform over a complete visit.

Speed Index​

Speed Index estimates how quickly visible content is populated during page load. It can reveal a page that gradually fills in even if its final load milestone looks acceptable.

Because it summarizes visual progress under a particular test condition, it is most useful for comparing repeatable synthetic runs. Do not interpret one Speed Index value as a universal user experience.

Page weight and request count​

Transferred bytes and request count are not direct experience metrics, but they are valuable leading indicators. A new JavaScript bundle, uncompressed image, font family, tag manager, or third-party integration can increase cost before field percentiles visibly cross a threshold.

Break totals down by resource type and first-party versus third-party origin. A 300 KB increase in render-blocking JavaScript has a different effect than a lazily loaded image far below the fold.

Complete journey duration​

For an application, the important performance unit may be a workflow rather than a page load. Measure how long it takes to sign in, search, add an item to a cart, or open a report. Preserve step timing so a regression can be isolated to navigation, API response, rendering, or a third-party action.

Why “page load time” is not enough​

Traditional page-load time often means the duration until the browser fires the load event. That event waits for many dependent resources, including some that do not affect the initial experience, and it does not describe responsiveness or later layout movement.

It remains operationally useful when defined consistently. A sharp increase may indicate more resources or a slower dependency. But it should not be the only performance objective.

A practical dashboard might show:

  • TTFB for the document and critical APIs.
  • FCP and LCP for initial rendering.
  • INP from real interactions.
  • CLS across the page lifecycle.
  • TBT and Speed Index from controlled lab runs.
  • Transferred bytes and requests by resource type.
  • Total duration and step timing for critical journeys.

Each metric has a question attached to it. That makes the dashboard actionable instead of decorative.

Synthetic monitoring vs. real-user monitoring​

Website performance has two primary data sources, and neither replaces the other.

DimensionSynthetic or lab monitoringReal-user or field monitoring
TrafficGenerated on a scheduleActual visits
ConditionsControlled device, network, location, and cacheReal devices, networks, locations, and behavior
RepeatabilityHigh when configuration is fixedLow at the individual-visit level
CoverageSelected pages and scripted actionsPages and actions users actually reach
Before releaseWorks in CI, staging, and preview environmentsUsually requires deployed traffic
Low-traffic pagesCan always collect a sampleMay take a long time to collect enough data
Regression speedCan detect a change on the next runDepends on traffic and aggregation window
Best useBaselines, release comparison, diagnosis, proactive alertsReal-world distribution, segmentation, business impact

Synthetic monitoring answers: Did this page become slower under the same test conditions?

Field monitoring answers: What performance did real visitors experience across their actual conditions?

Google's explanation of lab and field data differences is useful when the numbers disagree. A synthetic test may use a cold cache, one device profile, and one location. Real users have different hardware, networks, cache states, navigation paths, viewport sizes, and interaction patterns.

Do not average the two datasets together. Preserve their labels and use each for the question it can answer.

Understand the limitations of CrUX data​

The Chrome User Experience Report (CrUX) provides public field data from eligible Chrome users. It powers the field portion of tools such as PageSpeed Insights and the Core Web Vitals report in Search Console.

CrUX is an excellent starting point, but it has operational limits:

  • It represents a subset of Chrome visits, not every browser or user.
  • Low-traffic pages may not have enough samples for URL-level reporting.
  • Origin-level data can combine fast and slow page types.
  • Public tools expose fewer dimensions than a first-party RUM implementation.
  • Results use a rolling 28-day window, so a new regression or improvement appears gradually.

That 28-day window makes CrUX good for tracking broad user experience and search-related assessment, but too slow for detecting that this morning's deployment added two seconds to product-page LCP.

Collect your own field data when you need faster visibility and detailed segmentation. The official web-vitals library provides a production-ready way to observe LCP, INP, and CLS using browser APIs. Send the measurements to an analytics endpoint with page, device, navigation, release, and diagnostic context—while respecting privacy and keeping cardinality under control.

Build a repeatable synthetic baseline​

Synthetic results are valuable only when the test conditions are known and reasonably stable.

Define these dimensions for every monitored page:

  • Browser and version.
  • Viewport and device profile.
  • CPU and network conditions, including whether throttling is simulated or applied.
  • Geographic region and runner type.
  • Cold or warm browser cache.
  • Authenticated or anonymous state.
  • Cookie, consent, experiment, and localization state.
  • URL parameters and application data.

A baseline is a distribution, not the fastest run you have ever seen. Run enough checks to understand normal variation by page and region. Record medians and higher percentiles, then compare like with like.

For example, do not compare a cold-cache mobile simulation in Virginia with a warm-cache desktop browser in Frankfurt. Both may be valid tests, but the difference says nothing about a release.

Use representative pages rather than monitoring only the homepage:

  • Landing or marketing page.
  • Product, article, or documentation detail page.
  • Search or listing page.
  • Login page and an authenticated application route.
  • Checkout or conversion page.
  • A route that depends heavily on third-party code.

Pages that share a template can often be monitored as a group, with one stable representative URL plus field data across the complete page family.

Set performance budgets before a regression​

A performance budget defines the boundary a page or workflow should not cross. It turns “keep the site fast” into a decision rule.

Budgets can cover:

  • Core Web Vitals thresholds.
  • TTFB, FCP, TBT, or Speed Index.
  • Total JavaScript, image, font, or page bytes.
  • First-party and third-party request counts.
  • Duration of a critical journey or step.
  • Number or duration of long tasks.

Use two kinds of guardrail together:

  1. Absolute budget: The maximum acceptable value, such as LCP under 2.5 seconds in field data or checkout under 4 seconds in a controlled journey.
  2. Regression budget: The maximum acceptable change from a stable baseline or previous release, such as no more than a 15% increase in median synthetic LCP.

An absolute threshold alone may miss a serious change that remains just inside the limit. A relative threshold alone can create noise when a fast metric changes by a small absolute amount. Requiring a meaningful relative and absolute change is often more reliable.

Keep budgets specific to page type and test profile. A content page, analytics dashboard, and checkout flow have different work and user expectations.

Detect performance regressions without alert fatigue​

Browser measurements vary even under controlled conditions. Network scheduling, shared runners, CDN state, background work, and third-party services introduce noise.

Avoid paging on one slightly slower sample. Instead:

  • Compare a short rolling window with a longer baseline.
  • Require a minimum absolute change as well as a percentage change.
  • Confirm a threshold breach with an immediate rerun.
  • Require multiple consecutive degraded checks for non-critical metrics.
  • Compare regions before declaring a global incident.
  • Separate warnings from failures.
  • Preserve every sample for trends even when it does not alert.

Alert severity should reflect user impact. A 10% increase in below-the-fold image bytes may create a ticket. Checkout timing out in two regions may page the service owner immediately.

Do not hide persistent degradation behind repeated retries. A monitor that crosses its budget every morning and passes on confirmation is still reporting a pattern worth investigating.

Connect regressions to deployments and third parties​

A performance graph becomes far more useful when it includes release, configuration, feature-flag, and infrastructure-change markers. If LCP increases immediately after a deployment, responders can compare bundles, templates, and resource waterfalls before searching every dependency.

Record enough evidence with each synthetic run to explain the change:

  • The final LCP element and timing components.
  • Document and API response timing.
  • Request waterfall and failed resources.
  • Resource sizes by type and origin.
  • Long tasks and main-thread activity.
  • Screenshots or filmstrips at useful milestones.
  • Browser console errors.
  • Region, device profile, cache state, and release identifier.

Third-party code deserves its own visibility. Tag managers, consent tools, analytics, chat widgets, advertisements, fonts, payment components, and personalization services can change outside your deployment process.

Track third-party bytes, requests, errors, and duration by hostname. A regression with no internal release may line up with a new vendor script or a slower external API. Where the product permits it, compare a controlled run with optional third-party code blocked to estimate its contribution—but continue monitoring the real customer configuration as the primary signal.

Use the metric pattern to narrow the cause​

No metric proves a root cause, but combinations can prioritize investigation:

PatternLikely area to inspect first
TTFB and LCP rise togetherOrigin, backend, CDN cache, redirects, or regional network path
TTFB stable, LCP risesLCP resource discovery/loading, render-blocking CSS, fonts, or client rendering
FCP stable, LCP risesHero content, primary data, image priority, or later rendering work
LCP healthy, field INP poorMain-thread JavaScript, event handlers, rendering after interaction, or long tasks
Lab TBT rises and field INP later worsensNew JavaScript or main-thread work is a strong suspect
Synthetic CLS healthy, field CLS poorPost-load shifts, ads, personalized content, long sessions, or untested interactions
One region slows while others remain stableCDN, routing, DNS, regional origin, or local third party
Page weight rises but timings remain stableLeading indicator; caching or network conditions may be masking future impact

Use the request waterfall, browser trace, application telemetry, and deployment diff to confirm the hypothesis. Performance monitoring should shorten investigation, not replace it with guesswork.

A practical monitoring cadence​

Different signals operate at different speeds:

On every meaningful change​

Run controlled lab checks in CI or against a preview environment. Enforce resource and diagnostic budgets that are stable enough for the build pipeline. Compare several runs when timing affects the decision.

After deployment​

Run the production synthetic checks immediately. Verify important pages and journeys from the primary region, then compare the measurements with the pre-deployment baseline.

Continuously​

Schedule browser checks often enough to detect an operational regression within the response time the business requires. Use lightweight HTTP timing more frequently for broad endpoint coverage, and deeper browser checks for representative pages and critical workflows.

Daily and weekly​

Review synthetic distributions, RUM percentiles, Core Web Vitals pass rates, page families, regions, devices, browsers, and release versions. Look for slow trends that remain below immediate alert thresholds.

After an incident​

Add or refine a metric, page, device profile, segment, or assertion that would have detected the failure earlier. Remove measurements that produced noise without changing a decision.

Common website performance monitoring mistakes​

Treating a Lighthouse score as the objective​

The aggregate Lighthouse performance score is useful for a quick audit, but its component weights and scoring curves can change. Store raw metrics and diagnostics. Optimize the user experience, not only the score.

Monitoring only the homepage​

The homepage may be static and heavily cached while product, search, checkout, or authenticated pages perform very differently. Choose pages by template, traffic, revenue, and business risk.

Using averages alone​

Averages hide slow user segments and outliers. Track distributions and percentiles, especially the 75th percentile used for Core Web Vitals and higher percentiles for operational diagnosis.

Comparing unlike test conditions​

Changing the region, browser, cache, device, or throttle profile can look like an application regression. Version and label the test configuration.

Expecting lab INP to equal field INP​

INP depends on real interactions throughout a visit. Use field measurement as the primary signal, supported by TBT and scripted interaction diagnostics.

Ignoring single-page application navigations​

Initial document load is only one part of a long-lived application. Instrument route changes and important interactions, and use scripted journeys to time the complete workflow.

Alerting at the exact threshold​

Normal variation around a hard line creates flapping. Use rolling windows, confirmation checks, recovery thresholds, and material-change rules.

Collecting metrics without ownership​

Every alert and budget needs a responsible team and an expected action. A large dashboard nobody reviews does not protect performance.

Website performance monitoring checklist​

Use this as a starting point:

  • Identify representative pages and critical user journeys.
  • Collect field LCP, INP, and CLS at the 75th percentile for mobile and desktop.
  • Segment RUM by page family, device, geography, browser, and release.
  • Use CrUX as a public benchmark, understanding its coverage and 28-day window.
  • Run repeatable synthetic checks with documented browser, network, device, region, and cache settings.
  • Track TTFB, FCP, LCP, TBT, CLS, page weight, request count, and journey timing where relevant.
  • Record the LCP element, request waterfall, long tasks, screenshots, and release context for diagnosis.
  • Define absolute performance budgets and regression budgets by page type.
  • Add CI checks to catch regressions before release.
  • Run production checks after deployment and on a continuous schedule.
  • Use confirmation, rolling windows, and separate warning and failure thresholds.
  • Track first-party and third-party resources separately.
  • Route alerts to a named owner with a diagnostic link and runbook.
  • Review performance trends regularly, not only when an alert fires.

Monitor the experience, not merely the load event​

A credible performance program combines three perspectives:

  1. Field data shows what real users experienced across their actual devices, networks, and behavior.
  2. Synthetic monitoring provides a fast and repeatable signal for regressions under controlled conditions.
  3. Diagnostic evidence connects a changed metric to a resource, dependency, main-thread task, release, or region.

Core Web Vitals provide a useful common language for loading, responsiveness, and visual stability. TTFB, FCP, TBT, resource size, request timing, and complete journey duration explain more of the path. Performance budgets and scheduled checks turn those measurements into protection against regressions.

Testable Monitoring can pair frequent HTTP response-time checks with scheduled browser workflows on hosted or self-hosted runners. Scripted checks capture timing, assertions, screenshots, logs, traces, and request details as available, while metrics, incidents, notifications, maintenance windows, and status pages connect the measurement to an operational response. Learn more about synthetic browser and API monitoring.

For related guidance, see our website monitoring checklist, introduction to synthetic monitoring, RUM, and APM, and guide to turning Playwright tests into production monitors.