Pricing

Upsell and Cross-Sell A/B Testing: A Practical Framework

How to A/B test upsell and cross-sell offers: what to measure, where to place them, and why AOV decides the test, not attach rate alone.

Flat illustration in deep forest green and teal on a soft mint background: two shopping bags side by side, the left one receiving one extra item dropping into it and the right one holding several smaller items, representing an upsell next to a cross-sell

Testing whether to upsell or cross-sell, and where to place the offer, is one of the highest-volume experiment categories inside the complete guide to A/B testing pricing, and also one of the most commonly measured wrong. Teams routinely optimize for the number that is easiest to read on a dashboard, the attach rate (what percent of orders included the extra item or upgrade), when the number that actually decides whether the offer was worth showing is average order value (AOV) and total incremental revenue. This article covers what actually separates an upsell from a cross-sell, where in the funnel to place the offer, why attach rate alone can point you at the wrong winner, the guardrail metric that keeps an aggressive prompt from quietly increasing cart abandonment, and how testing a recommendation-engine cross-sell is a fundamentally different exercise from testing one manually curated offer.

Upsell vs Cross-Sell: The Distinction That Decides What You Are Testing

The two terms get used interchangeably in casual conversation, but they point at different customer decisions, and an A/B test needs to know which one it is running.

An upsell keeps the customer inside the same product line and asks them to buy more of it or a better version of it: the next size up, a longer commitment, a higher usage tier, a premium finish on the same item already in the cart. The customer’s basic intent does not change, only the scale or grade of what fulfills it.

A cross-sell introduces a second, different product that complements the first one: a case suggested after a phone, an extended-warranty plan bundled with an appliance, an integration offered alongside the core software plan. The customer’s basic intent stays with the original item; the test is whether a second, separate need can be surfaced and satisfied in the same purchase moment.

UpsellCross-sell
What changesThe same core product, moved to a higher tier or a larger quantityA different, complementary product added alongside it
Ecommerce exampleA larger size, a premium material, an extended warranty on the item already in the cartA phone case suggested right after a phone is added to the cart
SaaS exampleMore seats, a higher usage tier, an annual plan instead of monthlyAn add-on module, a third-party integration, a priority-support tier
What the A/B test actually measuresWhether a different price or framing converts better for the same upgrade pathWhether the suggested item is relevant enough, to that specific customer, to add
The practical consequence: an upsell test is a pricing and framing experiment on one path; a cross-sell test is a relevance experiment on a suggestion, and the two need different hypotheses and often different primary metrics.

The distinction is not academic. A losing upsell test usually means the price or the framing was wrong for an offer the customer already understood. A losing cross-sell test usually means the suggested item was not relevant enough, which is a data and matching problem, not a pricing one.

Where to Place the Offer: Pre-Purchase, In-Cart, In-Checkout, or Post-Purchase

The same upsell or cross-sell idea behaves very differently depending on where it interrupts the customer, because placement changes what is actually at risk if the customer says no.

Four places to test an upsell or cross-sell offerFrom the product page to after payment, there are four placement points: pre-purchase upsell on the product page, in-cart cross-sell, in-checkout cross-sell, and post-purchase upsell. The in-checkout point carries the highest friction risk.Product pagepre-purchase upsellCartin-cart cross-sellCheckout stephighest friction riskPost-purchaseorder already secured
Placement changes what is actually being risked: a rejected offer on the product page or in the cart costs nothing but a click, but an offer injected as a mandatory checkout step risks the sale itself.

The table below applies the same four placements to both ecommerce and SaaS, since the underlying risk logic is identical even though the surface looks different.

PlacementEcommerce shapeSaaS shapePrimary risk if overused
Pre-purchaseBundle or bigger size shown next to the base itemA higher tier highlighted before signupAnchors the wrong price expectation
In-cart / in-app“Frequently bought together” under the cart“Teams like yours also add X” inside the accountAdds clutter if shown on every single visit
In-checkout / in-billingA mandatory extra step before paymentAn upgrade prompt during the billing stepHighest cart or checkout abandonment risk
Post-purchaseA one-click offer on the thank-you pageAn in-product prompt after the first success momentLowest risk to the sale, easy to ignore if mistimed
The recommended default for a first test is in-cart or post-purchase in both business models: enough visibility to matter, without turning the offer into a toll booth.

The Metric That Should Decide the Test: AOV and Total Revenue, Not Attach Rate

Here is the trap most upsell and cross-sell tests fall into: attach rate tells you how many people said yes, and says nothing about what that yes was worth. A variant that gets more people to add something cheap is not automatically better than a variant that gets fewer people to add something expensive. The only way to know which one actually helped the business is to multiply attach rate by the value of what was attached, and compare the totals.

Salesforce’s Shopping Index, based on data from more than 150 million shoppers across roughly 250 million visits, found that visits where a shopper clicked a product recommendation made up only about 7% of all visits, yet those same visits drove roughly 24% of orders and 26% of revenue. A small slice of engaged customers, not the largest attach rate on the page, produced the outsized share of value: the same logic that makes attach rate, on its own, a misleading way to pick a winner inside a single test.

A Worked Example With Real Numbers

Say a store or app tests two versions of a post-purchase (or post-signup) offer against the same traffic, 6,000 sessions on each side:

Running those two attach rates through a two-proportion z-test gives a relative lift of +37.5% for B over A, z is approximately 5.60, and the p-value is far below 0.0001, an emphatically significant result. Read on attach rate alone, B is the clear winner, and by a wide margin.

Now multiply each attach rate by its item value across the same 6,000 sessions: A produces 480 times 45 dollars, or 21,600 dollars in incremental revenue. B produces 660 times 12 dollars, or 7,920 dollars. A generates nearly three times more revenue than B, despite having the statistically significant lower attach rate. Judging this test by attach rate alone would have shipped the variant that made the business less money.

Higher attach rate does not always mean higher revenueIllustrative example: variant B has a higher, statistically significant attach rate than A, 11.0% versus 8.0%, but because A’s item carries a much higher average value, A produces nearly three times more incremental revenue than B.Attach rateA8.0%B11.0%+37.5% relative, p under 0.0001Incremental revenueA$21,600B$7,920lower attach, higher item value wins
B wins decisively on attach rate, but A’s higher-value item still produces close to three times more incremental revenue over the same traffic. Deciding by attach rate alone would have picked the wrong variant.

Check the attach-rate comparison yourself with the numbers above (or your own):

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The lesson generalizes beyond this specific example: whenever two variants offer items of different value, attach rate and revenue can point in opposite directions, and only one of those two numbers should decide the test.

The Cart Abandonment Risk of an Aggressive Prompt

Attach rate and revenue are the upside metrics. There is also a downside metric that has to run as a guardrail on every upsell or cross-sell test: does the offer itself make people abandon the cart or checkout entirely.

The Baymard Institute’s large-scale checkout usability testing, drawn from a benchmark of hundreds of ecommerce sites, found that when a cross-sell offer was injected as its own mandatory step in checkout, one that required an active accept-or-reject decision before the customer could proceed, 66% of tested users who encountered this pattern at Amazon exhibited extreme frustration. The offer itself did not have to be a bad one to trigger that reaction; the friction came from being forced to make an extra decision at the exact moment the customer expected to simply finish paying.

That is the practical argument for treating cart or checkout abandonment as a guardrail metric, tracked on every upsell or cross-sell test, not just reviewed afterward if revenue looks disappointing:

A revenue win that quietly increases abandonment is not a clean win, it is a trade the test needs to expose, not hide.

Recommendation-Engine Cross-Sell vs a Manually Curated Upsell: A Different Test Entirely

A cross-sell surfaced by a recommendation engine, the “frequently bought together,” “customers who bought this also bought,” or “recommended for you” pattern, is not the same kind of test as a single curated upsell offer, even though both can be wired into the same A/B testing tool.

A manually curated offer tests one fixed hypothesis. The same specific item, the same price, the same copy, shown to everyone in that variant. A losing result tells you something concrete: this offer, at this price, with this framing, did not convert well enough.

A recommendation-engine cross-sell tests a model, not a single offer. The item shown in variant B is different for nearly every session, depending on what the model infers is relevant to that specific cart or account. A losing (or winning) result tells you something about the model’s overall relevance across the whole traffic mix, not about any one product pairing. Two consequences follow directly:

Neither approach is universally better: a curated upsell is easier to reason about and to iterate on quickly; a recommendation engine can scale relevance across a catalog far too large to curate by hand. What matters for the test design is knowing which one you are actually running, because the two failure modes, and the two paths to a better result, are not the same.

Sizing an Upsell or Cross-Sell A/B Test

Attach rates for a single upsell or cross-sell offer are often in the single digits to low teens, which means the sample needed to detect a meaningful difference is larger than it looks on paper, the same dynamic covered for freemium conversion in freemium paywall experiments.

With an 8.0% baseline attach rate (variant A in the worked example above) and a minimum detectable effect of 15% relative, the required sample works out to 8,568 sessions per variant. At 6,000 cart or post-purchase sessions per week reaching the offer, that is roughly 20 days to run the test, close to three weeks, which is longer than many teams budget for what looks like a small, low-stakes feature.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

Plug in your own baseline attach rate and the smallest lift that would actually be worth acting on before committing to a launch date for the test.

Common Mistakes in Upsell and Cross-Sell Testing

MistakeWhy it happensFix
Declaring a winner from attach rate aloneIt is the number that updates first and is easiest to readMultiply attach rate by item value and compare total incremental revenue
Comparing an upsell test to a cross-sell test with the same hypothesisBoth get built in the same A/B tool and look like the same kind of experimentWrite a separate hypothesis for each: one is about price and framing, the other about relevance
Forcing the offer as a mandatory checkout stepIt looks like a guaranteed way to raise visibilityTest in-cart or post-purchase first, and track abandonment as a guardrail from day one
Judging a recommendation-engine variant by a single product pairingOne bad-looking suggestion feels like proof the whole variant failedRead the result in aggregate; the unit under test is the model, not one pairing
Sizing the test by gut feelA single-digit attach rate makes teams assume any sample is enoughCalculate the required sample from the real baseline with the calculator above
Every mistake in this table traces back to the same root cause: treating attach rate as the finish line instead of an input to a revenue calculation.

Automate This in Donnu

The hard part of an upsell or cross-sell test is rarely running the split. It is resisting the pull to call a winner the moment attach rate moves, before the same variant has been checked against average order value, total incremental revenue, and cart or checkout abandonment. Donnu A/B runs the same two-proportion significance math on your attach rate while letting you track revenue and completion metrics alongside it, so a variant that wins the click and quietly loses the sale does not get shipped by accident.

Start a 14-day free trial and test your next upsell or cross-sell placement against the metric that actually decides it: revenue, not attach rate alone. See also the complete guide to A/B testing pricing and price anchoring experiments, the sibling pricing experiment on how the first price a visitor sees shapes what they choose next.


Read also: How to A/B Test Pricing: The Complete Guide · Price Anchoring Experiments · Freemium Paywall Experiments · Sample Size Calculator

References

Frequently asked questions

What is the real difference between an upsell and a cross-sell?
An upsell moves the customer to a higher tier or a larger quantity of the same thing they were already buying: a bigger size, more seats, an annual plan instead of monthly. A cross-sell adds a different, complementary item alongside it: a case suggested right after a phone is added to the cart, an integration offered next to the software plan just chosen. The distinction matters for testing because an upsell test is really a price or framing decision on a single upgrade path, while a cross-sell test is testing the relevance of a suggestion, which is a different hypothesis entirely.
Where should I place an upsell or cross-sell offer: before or after purchase?
Test at least three points before deciding: pre-purchase (on the product or plan page), in-cart (alongside the items already chosen), and post-purchase (after payment is captured, on the confirmation page or its equivalent). Avoid making the offer a mandatory step inside checkout itself: research from the Baymard Institute found that when a cross-sell was injected as a separate, unavoidable step in the checkout flow at Amazon, 66% of tested users showed extreme frustration. Post-purchase and in-cart placements let the customer skip the offer without touching the sale that is already secured.
Why should not attach rate alone decide an upsell or cross-sell test?
Because attach rate only counts how many customers said yes, not how much that yes was worth. A variant with a significantly higher attach rate on a low-value add-on can still generate less total revenue than a variant with a lower attach rate on a higher-value item, exactly the scenario in the worked example in this article. The decision metric has to be incremental revenue or average order value, calculated from attach rate times the value of what was attached, never attach rate in isolation.
Does an aggressive upsell or cross-sell prompt increase cart abandonment?
It can, and that risk is exactly why abandonment or checkout completion rate needs to run as a guardrail metric alongside AOV, not get discovered after the test ends. The Baymard Institute checkout research found that when a cross-sell offer required an active accept-or-reject decision as its own step in checkout at Amazon, 66% of tested users reacted with extreme frustration, the kind of friction that shows up later as a quieter completion rate, not a headline number anyone was watching in real time.
Is testing a recommendation-engine cross-sell the same as testing a manually curated upsell?
No, they test different things. A manually curated upsell tests a single fixed hypothesis: this specific offer, this specific price, this specific copy, against one alternative. A recommendation-engine cross-sell, the "frequently bought together" or "customers also chose" pattern, tests the output of a model that selects a different item for nearly every session, so the A/B test is really comparing two algorithms, not two static offers, and the winning variant can look inconsistent at the level of any single product while still winning in aggregate.