Upsell and Cross-Sell A/B Testing: A Practical Framework
How to A/B test upsell and cross-sell offers: what to measure, where to place them, and why AOV decides the test, not attach rate alone.

📚 This article is part of the guide How to A/B Test Pricing: The Complete Guide.
Testing whether to upsell or cross-sell, and where to place the offer, is one of the highest-volume experiment categories inside the complete guide to A/B testing pricing, and also one of the most commonly measured wrong. Teams routinely optimize for the number that is easiest to read on a dashboard, the attach rate (what percent of orders included the extra item or upgrade), when the number that actually decides whether the offer was worth showing is average order value (AOV) and total incremental revenue. This article covers what actually separates an upsell from a cross-sell, where in the funnel to place the offer, why attach rate alone can point you at the wrong winner, the guardrail metric that keeps an aggressive prompt from quietly increasing cart abandonment, and how testing a recommendation-engine cross-sell is a fundamentally different exercise from testing one manually curated offer.
Upsell vs Cross-Sell: The Distinction That Decides What You Are Testing
The two terms get used interchangeably in casual conversation, but they point at different customer decisions, and an A/B test needs to know which one it is running.
An upsell keeps the customer inside the same product line and asks them to buy more of it or a better version of it: the next size up, a longer commitment, a higher usage tier, a premium finish on the same item already in the cart. The customer’s basic intent does not change, only the scale or grade of what fulfills it.
A cross-sell introduces a second, different product that complements the first one: a case suggested after a phone, an extended-warranty plan bundled with an appliance, an integration offered alongside the core software plan. The customer’s basic intent stays with the original item; the test is whether a second, separate need can be surfaced and satisfied in the same purchase moment.
| Upsell | Cross-sell | |
|---|---|---|
| What changes | The same core product, moved to a higher tier or a larger quantity | A different, complementary product added alongside it |
| Ecommerce example | A larger size, a premium material, an extended warranty on the item already in the cart | A phone case suggested right after a phone is added to the cart |
| SaaS example | More seats, a higher usage tier, an annual plan instead of monthly | An add-on module, a third-party integration, a priority-support tier |
| What the A/B test actually measures | Whether a different price or framing converts better for the same upgrade path | Whether the suggested item is relevant enough, to that specific customer, to add |
The distinction is not academic. A losing upsell test usually means the price or the framing was wrong for an offer the customer already understood. A losing cross-sell test usually means the suggested item was not relevant enough, which is a data and matching problem, not a pricing one.
Where to Place the Offer: Pre-Purchase, In-Cart, In-Checkout, or Post-Purchase
The same upsell or cross-sell idea behaves very differently depending on where it interrupts the customer, because placement changes what is actually at risk if the customer says no.
- Pre-purchase, on the product or plan page itself, before anything has been added to the cart or any signup started. Low risk: a rejected offer here costs nothing but a moment of attention.
- In-cart, alongside the items or plan already chosen, most commonly as a “frequently bought together” or “you might also need” module. Still low risk, and the customer can simply ignore it and continue.
- In-checkout, as a dedicated step inside the payment flow itself, one that the customer has to actively accept or dismiss before moving forward. This is the highest-risk placement, covered in detail in the next section.
- Post-purchase, after payment has already been captured (an ecommerce thank-you page, or the equivalent moment in SaaS right after signup or renewal). The sale is already secured, so a rejected offer here costs nothing at all to the original transaction.
The table below applies the same four placements to both ecommerce and SaaS, since the underlying risk logic is identical even though the surface looks different.
| Placement | Ecommerce shape | SaaS shape | Primary risk if overused |
|---|---|---|---|
| Pre-purchase | Bundle or bigger size shown next to the base item | A higher tier highlighted before signup | Anchors the wrong price expectation |
| In-cart / in-app | “Frequently bought together” under the cart | “Teams like yours also add X” inside the account | Adds clutter if shown on every single visit |
| In-checkout / in-billing | A mandatory extra step before payment | An upgrade prompt during the billing step | Highest cart or checkout abandonment risk |
| Post-purchase | A one-click offer on the thank-you page | An in-product prompt after the first success moment | Lowest risk to the sale, easy to ignore if mistimed |
The Metric That Should Decide the Test: AOV and Total Revenue, Not Attach Rate
Here is the trap most upsell and cross-sell tests fall into: attach rate tells you how many people said yes, and says nothing about what that yes was worth. A variant that gets more people to add something cheap is not automatically better than a variant that gets fewer people to add something expensive. The only way to know which one actually helped the business is to multiply attach rate by the value of what was attached, and compare the totals.
Salesforce’s Shopping Index, based on data from more than 150 million shoppers across roughly 250 million visits, found that visits where a shopper clicked a product recommendation made up only about 7% of all visits, yet those same visits drove roughly 24% of orders and 26% of revenue. A small slice of engaged customers, not the largest attach rate on the page, produced the outsized share of value: the same logic that makes attach rate, on its own, a misleading way to pick a winner inside a single test.
A Worked Example With Real Numbers
Say a store or app tests two versions of a post-purchase (or post-signup) offer against the same traffic, 6,000 sessions on each side:
- Variant A, a manually curated, higher-value add-on (an extended warranty, a premium onboarding package): 480 of 6,000 sessions attach it (8.0%), at an average value of 45 dollars per attach.
- Variant B, a recommendation-engine cross-sell surfacing a cheaper, more broadly relevant item: 660 of 6,000 sessions attach it (11.0%), at an average value of 12 dollars per attach.
Running those two attach rates through a two-proportion z-test gives a relative lift of +37.5% for B over A, z is approximately 5.60, and the p-value is far below 0.0001, an emphatically significant result. Read on attach rate alone, B is the clear winner, and by a wide margin.
Now multiply each attach rate by its item value across the same 6,000 sessions: A produces 480 times 45 dollars, or 21,600 dollars in incremental revenue. B produces 660 times 12 dollars, or 7,920 dollars. A generates nearly three times more revenue than B, despite having the statistically significant lower attach rate. Judging this test by attach rate alone would have shipped the variant that made the business less money.
Check the attach-rate comparison yourself with the numbers above (or your own):
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The lesson generalizes beyond this specific example: whenever two variants offer items of different value, attach rate and revenue can point in opposite directions, and only one of those two numbers should decide the test.
The Cart Abandonment Risk of an Aggressive Prompt
Attach rate and revenue are the upside metrics. There is also a downside metric that has to run as a guardrail on every upsell or cross-sell test: does the offer itself make people abandon the cart or checkout entirely.
The Baymard Institute’s large-scale checkout usability testing, drawn from a benchmark of hundreds of ecommerce sites, found that when a cross-sell offer was injected as its own mandatory step in checkout, one that required an active accept-or-reject decision before the customer could proceed, 66% of tested users who encountered this pattern at Amazon exhibited extreme frustration. The offer itself did not have to be a bad one to trigger that reaction; the friction came from being forced to make an extra decision at the exact moment the customer expected to simply finish paying.
That is the practical argument for treating cart or checkout abandonment as a guardrail metric, tracked on every upsell or cross-sell test, not just reviewed afterward if revenue looks disappointing:
- Set the guardrail threshold before the test starts (for example, no statistically meaningful drop in checkout completion rate), the same discipline used for any other guardrail in an A/B test.
- Prefer in-cart or post-purchase placement over a forced in-checkout step whenever the two produce comparable attach rates, since the abandonment risk is structurally lower.
- If abandonment moves in the wrong direction even while revenue looks good, treat that as a real signal to redesign the placement, not as noise to explain away.
A revenue win that quietly increases abandonment is not a clean win, it is a trade the test needs to expose, not hide.
Recommendation-Engine Cross-Sell vs a Manually Curated Upsell: A Different Test Entirely
A cross-sell surfaced by a recommendation engine, the “frequently bought together,” “customers who bought this also bought,” or “recommended for you” pattern, is not the same kind of test as a single curated upsell offer, even though both can be wired into the same A/B testing tool.
A manually curated offer tests one fixed hypothesis. The same specific item, the same price, the same copy, shown to everyone in that variant. A losing result tells you something concrete: this offer, at this price, with this framing, did not convert well enough.
A recommendation-engine cross-sell tests a model, not a single offer. The item shown in variant B is different for nearly every session, depending on what the model infers is relevant to that specific cart or account. A losing (or winning) result tells you something about the model’s overall relevance across the whole traffic mix, not about any one product pairing. Two consequences follow directly:
- You cannot inspect a single “loser” product and conclude the whole variant failed. The unit under test is the algorithm’s aggregate output, so the result has to be read as an average of many different, session-specific suggestions, not as a verdict on any individual pairing.
- Iterating on a recommendation engine means retraining or reconfiguring the model, not editing copy. A curated upsell can be improved with a new headline or a new price point between test rounds. A recommendation-engine result can only really be improved by feeding it better signal (purchase history, catalog metadata, session context) or adjusting its ranking logic, a materially slower iteration loop.
Neither approach is universally better: a curated upsell is easier to reason about and to iterate on quickly; a recommendation engine can scale relevance across a catalog far too large to curate by hand. What matters for the test design is knowing which one you are actually running, because the two failure modes, and the two paths to a better result, are not the same.
Sizing an Upsell or Cross-Sell A/B Test
Attach rates for a single upsell or cross-sell offer are often in the single digits to low teens, which means the sample needed to detect a meaningful difference is larger than it looks on paper, the same dynamic covered for freemium conversion in freemium paywall experiments.
With an 8.0% baseline attach rate (variant A in the worked example above) and a minimum detectable effect of 15% relative, the required sample works out to 8,568 sessions per variant. At 6,000 cart or post-purchase sessions per week reaching the offer, that is roughly 20 days to run the test, close to three weeks, which is longer than many teams budget for what looks like a small, low-stakes feature.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Plug in your own baseline attach rate and the smallest lift that would actually be worth acting on before committing to a launch date for the test.
Common Mistakes in Upsell and Cross-Sell Testing
| Mistake | Why it happens | Fix |
|---|---|---|
| Declaring a winner from attach rate alone | It is the number that updates first and is easiest to read | Multiply attach rate by item value and compare total incremental revenue |
| Comparing an upsell test to a cross-sell test with the same hypothesis | Both get built in the same A/B tool and look like the same kind of experiment | Write a separate hypothesis for each: one is about price and framing, the other about relevance |
| Forcing the offer as a mandatory checkout step | It looks like a guaranteed way to raise visibility | Test in-cart or post-purchase first, and track abandonment as a guardrail from day one |
| Judging a recommendation-engine variant by a single product pairing | One bad-looking suggestion feels like proof the whole variant failed | Read the result in aggregate; the unit under test is the model, not one pairing |
| Sizing the test by gut feel | A single-digit attach rate makes teams assume any sample is enough | Calculate the required sample from the real baseline with the calculator above |
Automate This in Donnu
The hard part of an upsell or cross-sell test is rarely running the split. It is resisting the pull to call a winner the moment attach rate moves, before the same variant has been checked against average order value, total incremental revenue, and cart or checkout abandonment. Donnu A/B runs the same two-proportion significance math on your attach rate while letting you track revenue and completion metrics alongside it, so a variant that wins the click and quietly loses the sale does not get shipped by accident.
Start a 14-day free trial and test your next upsell or cross-sell placement against the metric that actually decides it: revenue, not attach rate alone. See also the complete guide to A/B testing pricing and price anchoring experiments, the sibling pricing experiment on how the first price a visitor sees shapes what they choose next.
Read also: How to A/B Test Pricing: The Complete Guide · Price Anchoring Experiments · Freemium Paywall Experiments · Sample Size Calculator
References
- Salesforce. Personalized Product Recommendations Drive Just 7% of Visits but 26% of Revenue. The original page is no longer live; archived copy: web.archive.org/web/20240901080301/salesforce.com/…/personalized-product-recommendations-drive-just-7-visits-26-revenue.html.
- Baymard Institute. Cross-Sells: Design Examples and Checkout Benchmark. baymard.com/checkout-usability/benchmark/step-type/cross-sell.
Frequently asked questions
- What is the real difference between an upsell and a cross-sell?
- An upsell moves the customer to a higher tier or a larger quantity of the same thing they were already buying: a bigger size, more seats, an annual plan instead of monthly. A cross-sell adds a different, complementary item alongside it: a case suggested right after a phone is added to the cart, an integration offered next to the software plan just chosen. The distinction matters for testing because an upsell test is really a price or framing decision on a single upgrade path, while a cross-sell test is testing the relevance of a suggestion, which is a different hypothesis entirely.
- Where should I place an upsell or cross-sell offer: before or after purchase?
- Test at least three points before deciding: pre-purchase (on the product or plan page), in-cart (alongside the items already chosen), and post-purchase (after payment is captured, on the confirmation page or its equivalent). Avoid making the offer a mandatory step inside checkout itself: research from the Baymard Institute found that when a cross-sell was injected as a separate, unavoidable step in the checkout flow at Amazon, 66% of tested users showed extreme frustration. Post-purchase and in-cart placements let the customer skip the offer without touching the sale that is already secured.
- Why should not attach rate alone decide an upsell or cross-sell test?
- Because attach rate only counts how many customers said yes, not how much that yes was worth. A variant with a significantly higher attach rate on a low-value add-on can still generate less total revenue than a variant with a lower attach rate on a higher-value item, exactly the scenario in the worked example in this article. The decision metric has to be incremental revenue or average order value, calculated from attach rate times the value of what was attached, never attach rate in isolation.
- Does an aggressive upsell or cross-sell prompt increase cart abandonment?
- It can, and that risk is exactly why abandonment or checkout completion rate needs to run as a guardrail metric alongside AOV, not get discovered after the test ends. The Baymard Institute checkout research found that when a cross-sell offer required an active accept-or-reject decision as its own step in checkout at Amazon, 66% of tested users reacted with extreme frustration, the kind of friction that shows up later as a quieter completion rate, not a headline number anyone was watching in real time.
- Is testing a recommendation-engine cross-sell the same as testing a manually curated upsell?
- No, they test different things. A manually curated upsell tests a single fixed hypothesis: this specific offer, this specific price, this specific copy, against one alternative. A recommendation-engine cross-sell, the "frequently bought together" or "customers also chose" pattern, tests the output of a model that selects a different item for nearly every session, so the A/B test is really comparing two algorithms, not two static offers, and the winning variant can look inconsistent at the level of any single product while still winning in aggregate.