When a checkout button changes colour, your software tests mostly stay green. Because tests verify that the button exists — not that it looks right. What is on screen is colour, position, typography and whitespace, and the only thing software reliably knows is the DOM.
Visual regression testing closes that gap: it takes a screenshot of a page, compares it against a reference image, finds the difference, and tells you about it.
In this post we cover what the technique actually does, where it earns its cost, the common methods, and — most usefully — where you should not bother with it. Because this is not a tool you bolt onto every project, and putting it in the wrong place produces noise for weeks.
Why DOM tests are not enough
Most software tests verify that an element is present:
expect(page.locator('[data-testid="checkout-button"]')).toBeVisible()
That test passes while any of the following is true:
- The button has moved off screen (overflow does not break visibility)
- Another layer sits on top of it (unclickable, still "visible")
- The background colour has dropped below the contrast threshold
- The font failed to load and rendered in a fallback
- Elements overlap on mobile
None of these break a DOM test. Visual regression testing does, because it compares what the page genuinely looks like.
How it works, in four steps
1. Capture a reference. The page is saved in a known-good state. That is the reference.
2. Re-capture. The same page, same browser, same viewport, same wait conditions.
3. Compute the pixel difference. The two images are compared, and the tool reports which pixels changed and by how much.
4. Apply a threshold. Small differences are ignored; large ones fail the build.
The hard part is step four. Too loose and you miss real bugs; too tight and every re-render produces a broken test. Animation, clocks, random content and font loading are all masked for exactly this reason.
The common methods
Pixel diff. Images compared pixel by pixel. Simple and fast, but brittle under small shifts — move an element one pixel and the test fails.
Perceptual hash. The image is reduced to a numeric summary that approximates human perception, and two numbers are compared. Far more tolerant of shifts; can miss small but meaningful changes.
Structural similarity (SSIM). Measures how structurally similar two images are. Cares more about layout than colour and texture. More stable when tests break because text shifted.
DOM and visual together. Many teams combine them: DOM tests give fast feedback, visual tests verify deeply. The cost is the sum of both.
Where it genuinely earns its cost
Where visual regression testing wins:
- Design system changes. One component update can affect forty pages. Checking forty pages by hand does not happen in practice.
- High-traffic conversion paths. Checkout, signup, cart. A break here is directly revenue.
- Consistency across browsers and viewports. How the same page looks at different sizes.
- Corporate content. Pricing tables, legal copy, campaign imagery.
Where it does not:
- Simple CRUD screens with flat text and backgrounds. The cost exceeds the rate of bugs you will actually catch.
- Constantly changing personal content. A timestamp, counter or ad slot means the test breaks on every run, and the team learns to ignore it — which is worse than having no test.
- One-off pre-launch checks. Opening a browser is faster.
A different use for live sites
Everything above is pre-launch testing: verifying what your code produces.
An agency or site owner has a different problem: monitoring the published site from the outside. Nobody is testing it, and nobody wrote the code. A CSS change hides an element, the server throws, a form silently stops submitting.
That is visual regression testing, run from outside. Same comparison — you just hold the reference image, and the comparison runs on its own.
The practical difference for an agency: checking thirty client sites by hand every day is not possible. Once it is automated, "something changed on your site" reaches you before the client has to ask.
Common mistakes
Setting the threshold too tight. Tests break, the team learns to treat them as noise, and eventually stops reading them. A broken test is worse than no test.
Skipping masks. If the clock, date and ad slot are not masked, every run fails. Starting without masks is harder than adding them later.
Unbounded test count. Writing a test per page instead of per user journey needs constant maintenance and catches nothing extra.
Not pinning the viewport. If the test passes at 1280px, you know nothing about 375px.
Summary
- Visual regression testing covers what the DOM cannot see: the actual appearance.
- Its biggest win is catching changes no single person can check by hand.
- Pixel diff is fast but brittle; perceptual hash and SSIM are more tolerant.
- A test in the wrong place is worse than no test: it produces noise and blinds the team.
- Monitoring live sites uses the same technique, run by an external tool.
Step-by-step setup is in our technical documentation, and you can compare plans or start free.