CRO

Ad to Landing Page Message Match: How to Test It

Message match: why comparing landing pages by ad misleads, how to design the right factorial test, and what to measure after the click.

Flat illustration of a small card on the left joined by a curved arrow to a tall page on the right, the same leaf shape appearing small on the card and large on the page, on a mint green background

Message match is how much of the ad’s promise reappears on the page that receives the click. Almost every team measures it the wrong way: they look at landing page conversion rate by ad and conclude which page is better. That does not work. The ad chooses who clicks, so each ad delivers a different population to the page, and the comparison mixes the page effect with a difference in audience intent. In this guide’s worked example, the aggregate read shows page X winning by 74.5 percent with a p-value under 0.0001, while inside each ad, separately, page Y is the one ahead. The design that fixes it is simple and almost nobody uses it: randomize the page inside each ad and read the effect within each block before aggregating. This guide is part of our complete A/B testing guide.

What message match is and why it matters

The idea is old and has a name in the usability literature: information scent. In the Nielsen Norman Group’s explanation of information foraging, a theory developed by Peter Pirolli and Stuart Card at PARC in the late 1990s, “each source of information thus emits a ‘scent’, a signal that tells the forager how likely it is that it contains what she needs”. People keep going while they feel they are getting warmer, and leave when they estimate that continuing costs more than it will return. The article frames the decision as a ratio between the value of the information and the cost of obtaining it.

Applied to paid media, that means the following. The ad creates a specific expectation: a price, a promise, a problem solved. The landing page either confirms that expectation in the first seconds, or forces the person to go looking. Looking is cost. High cost on a weak scent is an exit.

There is a second channel that almost nobody connects to the first: message match moves the auction price. In the Google Ads documentation, Quality Score has three components, one of which is landing page experience, defined as “how relevant and useful your landing page is to people who click your ad”. The Ad Rank documentation lists among its factors “the quality of your ads and landing page”, and records that “even if your competition has higher bids than yours, you can still win a higher position at a lower price by using highly relevant keywords and ads”. The same page decision can make the click cheaper and make that click convert better, which is rare in CRO.

Two paths from ad to conversion, with and without message matchTwo horizontal tracks. On the top track, with message match, the ad’s promise appears at three consecutive stops: in the ad, in the page headline and in the form offer, with the scent staying strong all the way to conversion. On the bottom track, without message match, the promise appears only in the ad, disappears in the page headline and in the offer, and the track fades out to an exit. A caption says the cost of looking is what decides the exit.a promise that vanishes midway turns into search costwith message matchadpage headlineoffer in the formconversionfree shippingfree shippingfree shippingwithout message matchadpage headlineoffer in the formconversionfree shippingpremium qualitysign upthe scent is lost herepeople keep going while they feel they are getting warmer, and leave when looking costs more than it returns
A reading of the information scent described by the Nielsen Norman Group from Pirolli and Card’s foraging theory. The promise does not need to be everywhere, it needs to be where the person looks first.

The measurement mistake almost everyone makes

Here is the problem that gives this guide its name. The natural question is “which page converts better”, and the natural data is each page’s conversion rate by ad. The problem is that in most accounts each ad sends most of its traffic to one specific page. The page was not randomized: it was chosen together with the ad.

And because the ad decides who clicks, each page receives a different mix of intent. Compare the two and you are comparing audiences, not pages.

An example with numbers. An account has two ads and two pages. Ad 1 is high-intent (brand search, the person already knows what they want, 6 percent baseline). Ad 2 is discovery (2 percent baseline). Historically, most of ad 1’s traffic went to page X and most of ad 2’s went to page Y.

slice page X page Y who is ahead
ad 1 (high intent) 16,000 clicks, 960 orders, 6.0000% 4,000 clicks, 252 orders, 6.3000% Y, plus 5.00%
ad 2 (discovery) 4,000 clicks, 80 orders, 2.0000% 16,000 clicks, 344 orders, 2.1500% Y, plus 7.50%
aggregate 20,000 clicks, 1,040 orders, 5.2000% 20,000 clicks, 596 orders, 2.9800% X, plus 74.5%

Inside each ad, page Y is ahead. In aggregate, page X wins by a landslide: the z-test returns minus 42.69 percent for Y, with a p-value under 0.0001 and a confidence interval from minus 2.6076 to minus 1.8324 percentage points, nowhere near zero.

Simpson reversal between the per-ad read and the aggregate readThree groups of two bars. In the first group, the high-intent ad, page X marks 6.00 percent and page Y marks 6.30 percent, with Y ahead. In the second group, the discovery ad, page X marks 2.00 percent and page Y marks 2.15 percent, again with Y ahead. In the third group, the aggregate, page X marks 5.20 percent and page Y marks 2.98 percent, now with X far ahead. A note explains that the reversal comes from the uneven allocation of each ad’s traffic across the pages.Y wins inside both ads and loses in the aggregatead 1, high intent6.00%6.30%XYY aheadad 2, discovery2.00%2.15%XYY aheadaggregate5.20%2.98%XYX aheadpage X received 80% of the high-intent traffic and page Y received 80% of the discovery trafficthe aggregate difference measures the audience mix, not the page
The numbers from the previous table, drawn. The reversal is not an arithmetic error: 5.20 and 2.98 percent are both real. They just do not answer the question the report claims to answer.

The cause is allocation: page X received 80 percent of the high-intent traffic and page Y received 80 percent of the discovery traffic. It is the classic structure described in Simpson’s paradox in A/B testing. And note that the p-value protects nobody here: the aggregate result is extremely significant and extremely misleading, because the problem is identification, not sample size.

The design that fixes it: the ad becomes a block, not a treatment

The correction is cheaper than it sounds. Instead of choosing one page per ad, randomize the page inside each ad, with fixed weights, and analyze the page effect within each ad before aggregating.

design what you learn what you miss
one page per ad, comparison across ads nothing causal about the page everything
page randomized across all traffic, aggregate analysis only the average page effect under the current traffic mix whether the effect depends on the ad
page randomized inside each ad, analyzed per block then aggregated the page effect inside each ad plus the weighted average nothing relevant

The third design is a simple factorial: two ads times two pages, four cells, with the page randomized inside each ad. It has three advantages and no practical downside.

First, the page comparison stays clean inside each block, because whoever clicked that ad was randomized across the two pages. Second, the aggregate stops depending on the traffic mix, because the mix becomes identical across arms by construction. Third, and most interesting, you measure whether the page effect depends on the ad, which is literally the message match question. That is pre-declared heterogeneity, not segment mining after the fact, as we explain in heterogeneous treatment effects.

On the practical side, the randomization can happen on your site, at the moment the person arrives, with the campaign parameter in the URL acting as the block. That is better than randomizing inside the platform for two reasons: you control the weights and the randomization unit, and you do not depend on the platform’s split mechanism, whose limits are covered in ad platform split testing.

Worked example: the factorial that reveals the fit

A store runs two paid search ads. Ad 1 promises free shipping above a threshold. Ad 2 promises 20 percent off the first order. Two landing pages:

The page is randomized 50/50 inside each ad. Ad 1 brings 26,000 clicks over the period and ad 2 brings 54,000.

block page clicks orders conversion rate
ad 1 (free shipping) X 13,000 780 6.0000%
ad 1 (free shipping) Y (aligned) 13,000 920 7.0769%
ad 2 (discount) X 27,000 540 2.0000%
ad 2 (discount) Y (misaligned) 27,000 524 1.9407%

The results, computed with the same two-proportion z-test as the calculator:

comparison relative lift p-value 95% CI of the difference reading
inside ad 1 plus 17.95% 0.0004 0.4761 to 1.6777 percentage points the aligned page wins
inside ad 2 minus 2.96% 0.6203 minus 0.2937 to 0.1752 percentage points a wash
aggregate (equal mix) plus 9.39% 0.0164 0.0569 to 0.5631 percentage points a real gain, but diluted
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste 13000 and 780 on side A and 13000 and 920 on side B to reproduce the first row: p-value 0.0004 and a 17.95 percent lift. Switch to 27000 and 540 against 27000 and 524 and the calculator returns 0.6203, which is the second row. And 40000 with 1320 against 40000 with 1444 returns 0.0164, the aggregate.

The reading. The 9.39 percent aggregate gain is real and statistically significant, but it is the weighted average of a large gain in one block and nothing in the other. Stop at the aggregate and the decision becomes “switch the page for everyone”, which applies to 54,000 clicks per period a change that does nothing for them (and, at the point estimate, slightly less than nothing). The right decision from this result is to route by ad: whoever comes from the shipping ad sees the shipping page; whoever comes from the discount ad stays on the generic page, or better, gets their own page aligned with the discount, which becomes the next test.

Note also what is not proven. It is not proven that “message match increases conversion by 18 percent”. It is proven that in this pair, aligning the top of the page with the shipping promise increased conversion for that traffic. Message match is not one effect, it is a family of effects that depends on which promise, which page and which audience.

How much traffic this design takes

The good news is that the high-intent block is usually cheap to test, because the baseline rate is high and so is the expected effect. The bad news is that the low-intent block is expensive, and that is usually where most of the traffic lives.

block and baseline plus 5% rel. plus 8% rel. plus 10% rel. plus 15% rel.
high intent, 6.0% baseline 100,670 39,861 25,740 11,693
low intent, 2.0% baseline 315,206 124,891 80,682 36,693

Values per cell, 95 percent confidence, 80 percent power, two-sided. For the large effect message match tends to produce in the aligned block, 18 percent relative on a 6 percent baseline, you need 8,225 clicks per cell, which on an ad producing 6,500 clicks a week takes 8 days with two arms.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

Three practical decisions come out of that table:

  1. Start with the highest-intent block. That is where the effect is largest and the sample is smallest. The result arrives in a week or two.
  2. Do not expect the same precision in the discovery block. In the example, the ad 2 cells with 27,000 clicks each can only see 16.88 percent relative. Declaring “no difference” there is imprecise; the correct wording is “we did not detect a difference above 17 percent”.
  3. Do not pool arms from different ads to gain sample. That is exactly what creates the reversal in the previous section.

What to align, in order of return

Message match is not copying the ad across the whole page. It is repeating enough of it where the person looks first.

page element how much it weighs what to do
main headline (H1) high repeat the ad’s central promise, in nearly the same words
first line under the H1 high resolve the condition of the promise: above what amount, for how long, for whom
hero image or video medium-high show the same object or scene as the creative, when there is a visual creative
primary button label medium use the verb of the promise, not a generic one
price or offer visible without scrolling medium when the ad gives a number, the same number has to appear
navigation and menu low, but negative on a campaign page, a full menu is usually an exit rather than a help
social proof low it helps, but it does not replace the promise; see social proof A/B testing

Two specific traps deserve attention.

The first is promising on the page what the product does not deliver. Alignment is not a license to repeat an exaggerated promise; if the ad promises unconditional free shipping and a condition exists, aligning the page only brings the disappointment forward. In that case the problem is the ad, not the page.

The second is confusing alignment with cloning. Repeating the ad copy literally in a long H1, with every keyword in it, usually hurts readability and does not improve conversion. What the person needs in two seconds is to recognize the promise, not to reread the ad.

Dynamic text replacement: automatic alignment has a price

There is a shortcut where the page swaps its headline based on a URL parameter, mirroring the keyword or the ad copy. It works, and it has three costs that need to be declared before you switch it on.

  1. It creates many loosely controlled content versions. A page that changes its H1 by parameter is, in practice, N pages. If they are indexable, that becomes a duplicate content problem; the minimum is blocking the parameter variants from indexing and keeping a canonical.
  2. It introduces a window where the headline changes after the page appears. That is the flicker effect, and it is not cosmetic: it interrupts reading and can bias the measurement. The treatment is in flicker effect in A/B testing.
  3. It mixes treatment and content at analysis time. If the injected text varies by query, each visitor received a slightly different treatment, and the “arm” stops being homogeneous. For the test, freeze the set of possible variants and log which one was served.

If the goal is just to validate the message match hypothesis, prefer two well-written static pages for the first test. Dynamic injection is a scaling optimization, not a measurement instrument.

How to read the result without fooling yourself

  1. Was the page randomized, or chosen together with the ad? If it was chosen, you do not have a test, you have an observation.
  2. Was the effect read inside each ad before aggregating? Without that, the audience mix can flip the sign.
  3. Is the traffic mix the same across arms? In the correct factorial it is, by construction. Check anyway.
  4. What was each block’s minimum detectable effect? “No difference” in a block that only sees 17 percent is not the same sentence as “no difference” in a block that sees 5.
  5. Is the metric conversion or revenue per click? If the discount ad brings smaller orders, conversion alone misleads; use revenue per visitor.
  6. Did cost per click move during the test? If so, comparing periods is contaminated by the auction on top of the page effect.
  7. Did you change only the top, or the whole page? Swapping the whole page answers “this page is better”, not “aligning the promise helps”.
  8. Does the result survive a second promise? A match effect that only appears for one ad and page pair is a small, useful finding, not a law.

Pre-launch checklist

  1. List the ads that qualify as blocks, with click volume and baseline rate for each.
  2. Pick the promise to align, one of them, and write the hypothesis: if the H1 repeats promise X, conversion for that ad’s traffic goes up because the information scent holds.
  3. Build the factorial: page randomized 50/50 inside each ad, with the campaign parameter defining the block.
  4. Freeze the ad for the whole period. If the creative changes, who clicks changes, and the block is no longer the same.
  5. Compute each block’s minimum detectable effect before launch, and accept up front that the small block may not conclude.
  6. Declare the primary metric, with the denominator in block clicks, and the guardrail (revenue per click, return rate).
  7. Make sure the aligned page promises nothing false.
  8. Run whole weeks, as in weekly cycle and test duration.
  9. Plan the routing, not just the winner: the likely result is “it depends on the ad”, and that is a campaign architecture decision, not a single-page one.

Automate this with Donnu

The specific pain of testing message match is that the variable lives on one side of the click and the measurement on the other, and the easy way to look at the data is exactly the one that flips the sign.

In Donnu, an experiment can be restricted by traffic source, which lets you run the page test only for visitors arriving from a specific campaign, which is exactly the block in this guide. Randomization is per visitor, with fixed weights you define, and a conversion goal can be confirmed from your own server, with the order value, which lets you read revenue per click instead of conversion rate alone. The report is Bayesian, warns when the visitor split drifts from what you configured, and only declares a winner with at least 200 visitors per variation and 7 days of testing.

What stays on you: choosing which promise to align, freezing the ad during the test, and deciding the routing at the end. For the math, the sample size calculator, the significance calculator and the landing page grader work with your numbers, for free.

References

Read next: Ad creative testing · Ad platform split testing · Simpson’s paradox · Heterogeneous treatment effects · Flicker effect · Social proof A/B testing · Incrementality testing for paid media · Leia em português

Frequently asked questions

What is ad to landing page message match?
It is how much of the ad's promise reappears on the landing page, in the copy, the offer and the imagery. The concept behind it is information scent, from information foraging theory: according to the Nielsen Norman Group, each source emits a signal that tells the forager how likely it is to hold what they need, and people keep going as long as they feel they are getting warmer. When the ad promises free shipping and the page talks about craftsmanship, the scent is lost.
Can I compare landing page conversion rate by ad?
Not as causal evidence. The ad decides who clicks, so each ad delivers a different population to the page. Comparing pages across ads mixes the page effect with the difference in audience intent. In this guide's worked example, the aggregate comparison shows page X winning by 74.5 percent with a p-value under 0.0001, while inside each ad page Y is ahead. It is a textbook Simpson reversal.
What is the correct design?
Randomize the page inside each ad, 50/50, and analyze the effect within each ad before aggregating. The ad becomes a block, not a treatment. The comparison between pages stays clean, and you also learn whether the page effect depends on the ad, which is exactly the message match question.
Does message match only affect conversion?
No. It also enters the auction price in paid search. According to Google Ads, Quality Score has three components, one of which is landing page experience, defined as how relevant and useful your landing page is to people who click your ad. The Ad Rank documentation records that even if your competition has higher bids than yours, you can still win a higher position at a lower price by using highly relevant keywords and ads.
Is the message match effect the same for every ad?
No, and that is the point. In this guide's worked example, the page aligned with the shipping promise wins 17.95 percent inside the shipping ad, with a p-value of 0.0004, and loses 2.96 percent inside the discount ad, with a p-value of 0.6203. There is no universally better page: there are ad and page pairs that fit.
How much traffic does this take?
Less than you might think, if you test the match where it should matter most. Detecting 18 percent relative on a 6 percent conversion baseline takes 8,225 clicks per cell. What gets expensive is the low-intent block: detecting 10 percent relative on a 2 percent baseline takes 80,682 clicks per cell.