CRO

Social Proof A/B Testing: What to Test and How to Measure

Social proof A/B testing: reviews, logos, testimonials and live counters. What the evidence shows, what to test, and the metric that does not lie.

Flat illustration of a dark green shopping bag in the center, surrounded by avatars of people in speech bubbles, a bubble with lines of text, gold stars and hearts, on a mint green background

Social proof is any signal that other people have already chosen, used or approved what a visitor is evaluating: review scores and counts, customer numbers, testimonials, client logos, real-time activity, badges and user-generated content. It works by reducing uncertainty, which is why it is a common CRO hypothesis and an easy one to measure badly. The usual mistake is not picking the element, it is picking the metric: in the worked example in this guide, a review badge on a product page lifts orders by 8.75 percent with a p-value of 0.0010, while orders that are not sent back as returns rise by only 0.99 percent, with a p-value of 0.7280. This guide is part of our complete guide to conversion rate optimization (CRO) and covers what the research shows (and what it does not), what to test in ecommerce and SaaS, the five measurement mistakes, and two worked examples calculated end to end.

What social proof is and why it works

Robert Cialdini’s company sums up the principle in one sentence: especially when they are uncertain, people look to the actions and behaviors of others to determine their own. The word doing the work is uncertain. Social proof does not persuade anyone on its own; it lends other people’s experience to someone who does not yet have enough information to decide.

For testers, that means the expected effect depends on how much uncertainty exists at that point on the page. A visitor who came from a branded search, already decided, has little doubt left to reduce; a visitor looking at an expensive product from an unfamiliar brand has plenty. The same badge can have zero effect on one page and a meaningful one on another.

Goldstein, Cialdini and Griskevicius added a second condition in a 2008 field experiment: the proof worked better when the group described was in the same situation as the reader, which they called a provincial norm. On a page, “customers your size” or “people who bought this size” tends to be a stronger hypothesis than “thousands of customers”.

The seven types of social proof on a page

type what it signals where it usually appears main risk
Review score and count other people bought and rated it near the title and price on a product page a perfect score with few reviews looks too good
Customer or user count scale and staying power home page, landing page, pricing page an inflated number is deceptive advertising
Testimonials concrete experience of someone similar landing page, pricing page, checkout generic or unverifiable testimonials
Client logos known companies trusted you B2B landing page and SaaS pricing page logos the visitor does not recognize or relate to
Real-time activity other people are acting right now product page and cart a randomly generated or hard-coded counter
Badges and certifications a third party verified something checkout, footer, security page a badge with no real verification behind it
User-generated content photos and videos of real use product page gallery, social media page weight and curation that hides the negatives

The last column is almost entirely about credibility. Social proof is a claim about facts, which exposes it to two risks an A/B test dashboard never shows: looking fake and being fake.

What the evidence actually shows

Three sources are commonly cited on social proof, and the quality of evidence behind each is very different. Read each one for its design, not its headline number.

The towel experiment: strong, famous, and a replication that failed

In the first experiment by Goldstein, Cialdini and Griskevicius, the 190 rooms of a hotel belonging to a national US chain were randomly assigned one of two cards. The standard card talked about protecting the environment. The descriptive norm card said almost 75 percent of guests took part in the program. Counting only each guest’s first eligible day, towel reuse was 44.1 percent with the norm versus 35.1 percent with the environmental appeal, with a p-value of .05 in a chi-square test on 433 observations.

In the second experiment, with 1,595 observations at the same hotel and five cards, the four norm messages combined reached 44.5 percent against 37.2 percent for the environmental message, and the “guests who stayed in this room” norm reached 49.3 percent, versus 42.8 percent for the other three norms combined.

Cialdini’s own page presents the result as increases of 26 and 33 percent. The arithmetic matches the paper (44.1 against 35.1 is about 26 percent relative; 49.3 against 37.2 is about 33), but each number comes from a different experiment with a different control. It is a good reminder of how research turns into a headline.

Then there is the part that usually gets left out. In 2014, Bohner and Schlüter repeated the design at two hotels in Germany, published in PLoS ONE. In study 1, with 724 observations, the four norms combined produced 81.9 percent reuse against 83.7 percent for the environmental message, p = .62. In study 2, with 204 observations, the environmental message did best of all (93.3 percent). The authors’ conclusion: descriptive norm messages were not more effective than the standard message, and proximity effects were inconsistent across studies.

Towel reuse in the original experiment and in the replicationHorizontal bars of towel reuse rate. Goldstein and coauthors, experiment 1: environmental message 35.1 percent, descriptive norm 44.1 percent. Experiment 2: environmental message 37.2 percent, four norms combined 44.5 percent, same room norm 49.3 percent. Bohner and Schlüter, study 1: environmental message 83.7 percent, four norms combined 81.9 percent, no significant difference.same idea, two outcomes: the context changes everythingGoldstein and coauthors, 2008, experiment 1 (US)environmental message35.1%descriptive norm44.1%Goldstein and coauthors, 2008, experiment 2 (US)environmental message37.2%four norms combined44.5%same room norm49.3%Bohner and Schlüter, 2014, study 1 (Germany)environmental message83.7%four norms combined81.9%scale from 0 to 100 percent starting at the vertical line. In the replication, p = .62.sources: Journal of Consumer Research, 2008; PLoS ONE, 2014
In Germany, reuse was already high with any card, and the descriptive norm added nothing. A real effect in one context is not a guaranteed effect in another, which is exactly why social proof gets tested instead of copied.

There is also a design detail that matters to anyone running experiments. At the US hotel, randomization was by room (190 rooms in experiment 1) and the analysis was by guest. When the randomized unit is larger than the analyzed unit, observations from the same group tend to resemble each other, and a test that treats each one as independent tends to overstate precision. We cover that math in cluster randomization and in choosing the randomization unit.

The 270 percent figure: observational, with a vendor involved

The report How Online Reviews Influence Sales, from Northwestern’s Spiegel Research Center, produced in partnership with PowerReviews (a company that sells ratings and reviews software), is the source of one of the most repeated numbers about social proof. What it says, read in full:

finding where it comes from how to read it
purchase likelihood with five reviews is 270% greater than with none high-end gift retailer: about 15.5 million page views, 1,800 products, 7.8 million users over a year, tracking products as they accumulated reviews observational; a product that accumulates reviews also accumulates catalog time and sales
displaying reviews lifted conversion by 190% for lower-priced and 380% for higher-priced products same retailer, comparing categories the direction (more effect on risky purchases) transfers, the magnitude does not
purchase likelihood typically peaks between 4.0 and 4.7 stars and falls as ratings approach 5.0 across product categories the same report also says 4.2 to 4.5 in its closing summary; treat it as a range, not a target
nearly all of the gain happens within the first 10 reviews, and the first 5 drive most of it gift retailer an argument for a few reviews on many products
verified buyer reviews raise purchase likelihood by 15% versus anonymous ones about 13,500 products, 57,000 anonymous and 65,000 verified buyer reviews a verified buyer badge is a cheap hypothesis to test

This is not an A/B test. A product that reaches five reviews is, by construction, a product that has already sold and has been in the catalog longer, so the number does not mean “add stars and sell 3.7 times more”. What transfers is the direction: reviews matter more where uncertainty is higher, a perfect score raises suspicion, and the first reviews are worth more than later ones.

Social proof that is a lie

Mathur and coauthors at Princeton crawled roughly 53,000 product pages across roughly 11,000 stores (CSCW 2019) and found 1,818 instances of dark patterns on 1,254 sites. Among the 313 activity notifications (who just bought, how many people are viewing), 29, on 20 sites, were deceptive: most generated the number with a random number generator or showed hard-coded messages. Among the testimonials of uncertain origin, in one case the same set of testimonials appeared on another store, with different customer names attached.

For testers: a fake counter can “win” an A/B test. The dashboard measures behavior, not truthfulness, so “is it true?” comes before “does it convert?”.

What to test in social proof: the ecommerce vs SaaS matrix

Because social proof works on uncertainty, the useful question is not “should we add testimonials?”, it is “what doubt does the visitor have at this point, and what proof answers it?”. The doubt differs between ecommerce and SaaS, and from one step to the next.

The visitor’s doubt and the proof that answers itTwo columns. Ecommerce, product page: doubts “will it fit me”, “does this store deliver”, “did other people regret it”, answered by reviews with photos and size worn, a score with count near the price, and a verified buyer badge. SaaS, pricing page: doubts “does it work for companies like mine”, “what if I subscribe and it does not work”, “is my data safe”, answered by a testimonial from a similar sized company, logos from the visiting segment, and compliance information.the right proof answers the doubt at that exact pointEcommerce: product pagedoubtwill it fit me? does it look like the photo?does this store actually deliver?did other people regret buying it?proof that answers itreviews with photos and size wornscore with review count, near the priceverified buyer badgeSaaS: pricing pagedoubtdoes it work for companies like mine?what if I subscribe and it does not work?is my data safe?proof that answers ittestimonial from a company your sizelogos from the visiting segmentcompliance and where data is stored
The provincial norm from the towel experiment, translated to a page: the more the cited group resembles the visitor, the more plausible the effect. Plausible, not guaranteed, as the replication showed.

The matrix below organizes the most common hypotheses by the lever they pull. The “measurement trap” column is the one that matters most when you read the result.

lever ecommerce: what to test SaaS: what to test measurement trap
Placement score with count right under the title, near the price, versus only at the bottom logos and a testimonial above the plan table versus below it clicks on the score rise just because it became visible
Specificity reviews filtered by size or use case versus the general list testimonial from the same company size or industry versus a generic one segment defined after looking at the data
Score format average with count versus average alone; showing the star distribution customer count versus logos high score with very few reviews
Credibility verified buyer badge; showing negative reviews testimonial with name, role and company versus anonymous a conversion gain that turns into returns or churn
Activity “sold in the last 24 hours”, if real “companies that started this month”, if real a counter that changes on its own during the test
Badges and compliance security badge next to the card field GDPR or LGPD information, data location, real certifications a badge both arms already see through another path
User content customer photo gallery on the product page case study with a verifiable number page weight changes along with the proof

Two hypotheses cut against intuition. Showing negative reviews: the Spiegel report links near-perfect scores to lower purchase likelihood, and our product page A/B testing guide argues that negatives calibrate expectations and tend to reduce returns. Swapping big-brand logos for logos that look like the visitor: the SaaS pricing page guide describes the wall of famous brands that convinces a small company the product is not for them, and the B2B SaaS landing page guide recommends testing which logos, for which segment.

A technical note for ecommerce: Google asks that review content marked up with structured data be readily available to users on the marked-up page. Hiding reviews in the control while keeping the markup conflicts with that. Testing placement, format and emphasis is safer than testing whether reviews exist at all.

How to measure social proof without fooling yourself

Social proof is cheap to implement and expensive to measure. Five problems show up in almost every test of this kind, and each one manufactures a false winner in a different way.

1. The wrong metric: clicks, conversion, revenue or orders that stay

The most tempting metric is the one that moves first: clicks on the stars, opens of the testimonial carousel. It explains why a variation won and does not decide whether it won, because any visual emphasis increases clicks.

The metric ladder for a social proof testFour steps from left to right, each taller. Ecommerce: clicks on stars, add to cart, completed order, order not returned. SaaS: click on logo or testimonial, trial start, paid account, retained account. An arrow shows that climbing the ladder brings the metric closer to money, lowers the baseline rate and raises the sample required.the closer to money, the more sample the metric needsclick on the proofstars, logodiagnosticnext stepcart, trialsecondarytransactionorder, paid accountprimaryvalue that staysnot returned,retained accountguardrail or decisionlower baseline and slower maturation at every step: more visitors and more days for the same read
A social proof test that only measures the first step almost always “wins”. The honest work is sizing for the step that decides and using the lower steps to explain.

The rule of thumb: primary metric on the transaction step, guardrail on the value that stays. In ecommerce, completed orders decide and orders not returned (or revenue net of returns) protect. In SaaS, trial start is usually the feasible primary, and paid accounts decide value, at a much larger sample. The guardrail metrics guide shows how to fix thresholds before the test, and the sample size guide for revenue and continuous metrics shows why revenue per visitor needs even more sample than conversion.

2. Novelty effect: a new badge draws attention because it is new

A review block that suddenly appears on a page returning customers know well is, in its first week, mostly novelty, and the early gain can shrink. To separate novelty from effect, compare new and returning visitors and look at the week by week trajectory, as in novelty effect in A/B testing.

3. Contamination: real social proof changes during the test

All genuine social proof is a live number: the count grows, the average moves, a negative review lands at the top. That creates three problems:

None of this invalidates the test, but it changes the read: log the displayed score and count every day, and know that the result holds for the range of values that actually appeared.

4. Heterogeneity: new vs returning, mobile vs desktop

It is reasonable to expect social proof to help new visitors more, and a long testimonial block to behave differently on a small screen. The danger is discovering this afterwards: with ten independent cuts each tested at 5 percent, the chance of at least one false positive is above 40 percent. Declare two or three segments with a reason up front and read the difference with the right math, covered in heterogeneous treatment effects and Simpson’s paradox.

5. Fake social proof: the test cannot see it, the law can

jurisdiction rule what it says, in short what it changes in a test
United States FTC Rule on the Use of Consumer Reviews and Testimonials (16 CFR Part 465), announced August 14, 2024, in effect since October 21, 2024 bans fake or false consumer reviews and testimonials (including AI-generated ones), compensation conditioned on a particular sentiment, undisclosed insider reviews, suppressing reviews through threats, and buying or selling fake indicators of social media influence such as followers or views; allows civil penalties for knowing violations a variation with cherry-picked reviews, an invented counter or an employee testimonial is not a hypothesis, it is exposure
Brazil Consumer Defense Code, Law 8,078/1990, articles 37, 38 and 67 bans misleading advertising, including partly false or by omission; the burden of proving truthfulness falls on the advertiser; advertising known or presumed to be misleading carries three months to one year of detention plus a fine customer counts, scores and counters must be provable in both arms
Brazil Conar, Brazilian Advertising Self-Regulation Code, Annex Q (testimonials) an identified consumer’s first and last name must be real; the advertiser’s employees must not pose as ordinary consumers; the advertiser must prove a testimonial is genuine when asked an unauthorized or unverifiable testimonial does not go into a variation

In its questions and answers about the rule, the FTC clarifies that organizing reviews is not suppressing them, but organizing them in a way that makes negative reviews hard to find could be deceptive under the FTC Act. For anyone testing review ordering, that is the line. This is context, not legal advice.

Worked example 1: the review badge on a product page

Scenario (illustrative). A fashion store with 60,000 weekly visitors on product pages converts 2.4 percent of visits into orders. Reviews exist today but sit at the bottom of the page. The hypothesis: showing the average score with the review count right under the title, near the price, reduces doubt at the moment it appears and increases orders.

Metrics declared up front. Primary: visitor with a completed order. Guardrail: visitor with an order not returned within 30 days (in this scenario, each buyer places one order). Diagnostic: clicks on the score. Segments: new and returning. All at the template level, across the catalog, because a single product page would never have enough traffic.

Sample size. The team wants to detect an 8 percent relative lift. In the calculator below, enter: current conversion rate 2.4; minimum detectable effect 8, relative; confidence 95; power 80; visitors per week 60000; test two-sided.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The screen shows 103,633 visitors per variation, 207,266 in total and 25 days. The team rounds up to four full weeks, 28 days, which gives 120,000 visitors per arm. To feel the sensitivity, change only the minimum effect:

relative effect you want to detect visitors per variation total days at 60,000 a week
5% 261,572 523,144 62
8% 103,633 207,266 25
10% 66,946 133,892 16
12% 46,922 93,844 11

The peek you should not act on. At the end of week one, with 30,000 visitors per arm, there were 720 orders in the control and 834 in the variation: plus 15.8 percent, p-value 0.0034. Stopping there would have recorded almost double the lift the full test showed. In weeks 2 to 4, with 90,000 per arm, the score was 2,160 versus 2,298, plus 6.4 percent, p-value 0.0364. A stronger first week is consistent with novelty and also with noise; what it does not justify is a decision. The cost of looking early is in the peeking problem.

The primary metric result, after 28 days.

arm visitors orders rate
A, reviews at the bottom of the page 120,000 2,880 2.40%
B, score with count near the price 120,000 3,132 2.61%

The split is clean: 120,000 versus 120,000 raises no flag in the SRM checker. Now paste into the significance calculator: control with 120000 visitors and 2880 conversions; variation with 120000 visitors and 3132 conversions; confidence 95.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows a 2.40% rate for the control and 2.61% for the variation, a relative lift of +8.8%, a p-value of 0.0010, a 95 percent CI of the difference of +0.1% … +0.3% (pp) and the verdict Significant winner · B wins. With more decimals, the lift is 8.75 percent, the p-value is 0.000997 and the interval runs from plus 0.0850 to plus 0.3350 percentage points, or roughly plus 3.5 to plus 14.0 percent in relative terms (dividing the interval by the control rate).

So far, the usual story: B won, comfortably. That is 252 extra orders.

The twist: a guardrail that only matures 30 days later. Orders with a 30-day return window only close almost two months after the test started. When they do:

arm orders returned return rate visitors with a kept order rate per visitor
A 2,880 461 16.01% 2,419 2.02%
B 3,132 689 22.00% 2,443 2.04%

Of the 252 extra orders, 228 came back. To check both reads in the same calculator, first enter orders as “visitors” and returns as “conversions” (2880 and 461 versus 3132 and 689): the screen shows 16.01% versus 22.00%, a lift of +37.4%, a p-value below 0.0001 (the screen prints a less-than sign before the number), a CI of +4.0% … +8.0% (pp) and the verdict “Significant winner · B wins”. Here the calculator only says that B’s rate is higher, and on this metric higher is worse.

Then enter the guardrail: 120000 and 2419 versus 120000 and 2443. The screen shows 2.02% versus 2.04%, a lift of +1.0%, a p-value of 0.7280, a CI of -0.1% … +0.1% (pp) and Not significant yet. With more decimals, the lift is 0.99 percent and the interval runs from minus 0.0927 to plus 0.1327 percentage points, roughly minus 4.6 to plus 6.6 percent relative.

Effect of the review badge on orders and on kept ordersTwo 95 percent confidence intervals on a relative effect axis from minus 10 to plus 20 percent. Orders: estimate of plus 8.75 percent, interval from plus 3.5 to plus 14.0 percent, entirely to the right of zero. Orders kept after 30 days: estimate of plus 0.99 percent, interval from minus 4.6 to plus 6.6 percent, crossing zero.one test, two questions: did it sell more? did it keep more?-10%0%+10%+20%orders28 days+8.75%, p = 0.0010kept orders30 days later+0.99%, p = 0.7280approximate relative interval: interval of the difference divided by the control rate.illustrative scenario calculated with the same engine as the blog calculators
The kept-orders interval crosses zero and excludes the 8.75 percent gain. The test had about 79 percent power to see an 8 percent gain on that metric, so “it did not show up” is informative here, not just a sample shortfall.

What the result says. The badge pushed people into buying who used to scroll to the reviews, read about sizing, and leave. In this scenario, the diagnostics back that reading: clicks on “see reviews” fell in the variation, and the most common return reason in B is fit. The badge answered the wrong doubt (can I trust this store) when the real doubt was size. And it is not just a lack of sample: for an 8 percent gain in kept orders, the 2.02 percent baseline needs 123,628 visitors per variation (29 days), and the test had 120,000.

What to do with it. Do not ship B as is. The next variation pairs the score near the price with a review excerpt filtered by size, which answers the doubt that was driving returns. Same primary, same guardrail, and this time with the final read scheduled after the return window.

Worked example 2: client logos on a SaaS pricing page

Scenario (illustrative). A B2B SaaS gets 12,000 weekly visitors on its pricing page, and 7 percent of them start a trial. About 15 percent of trials become paid accounts within 30 days, which works out to 1.05 percent paid accounts per visitor. The hypothesis: a strip of well known client logos above the plan table reduces the perceived risk of subscribing.

Sample size. In the same sample size calculator above, change the fields: current rate 7; minimum effect 15, relative; confidence 95; power 80; visitors per week 12000; two-sided. The screen shows 9,907 visitors per variation, 19,814 in total and 12 days. The team runs two full weeks: 12,000 visitors per arm.

The trial result. In the significance calculator, enter 12000 and 840 for the control and 12000 and 966 for the variation. The screen shows 7.00% versus 8.05%, a lift of +15.0%, a p-value of 0.0020, a CI of +0.4% … +1.7% (pp) and Significant winner · B wins. That is 126 extra trials.

The paid account result, 30 days later.

metric A B relative lift p-value read
trial starts per visitor 840 of 12,000 (7.00%) 966 of 12,000 (8.05%) +15.0% 0.0020 significant
trials that convert to paid 126 of 840 (15.00%) 130 of 966 (13.46%) -10.3% 0.3486 not significant
paid accounts per visitor 126 of 12,000 (1.05%) 130 of 12,000 (1.08%) +3.2% 0.8015 not significant

To check the last row, enter 12000 and 126 versus 12000 and 130: the screen shows 1.05% versus 1.08%, a lift of +3.2%, a p-value of 0.8015, a CI of -0.2% … +0.3% (pp) and Not significant yet. With more decimals, the interval runs from minus 0.2266 to plus 0.2933 percentage points, something like minus 21.6 to plus 27.9 percent relative. Of the 126 extra trials, 4 became extra paid accounts.

The twist here is different from the ecommerce one. The paid account result does not say logos do not work. It says the test was never able to answer. Go back to the sample size calculator and enter the metric that decides revenue: rate 1.05, effect 15 relative, 95, 80, 12000 a week. The screen shows 70,622 per variation, 141,244 in total and 83 days. The calculator does not show power, so the next numbers are this guide’s own calculation with the same formula: with 12,000 per arm, the power to detect a 15 percent lift in paid accounts was about 21 percent, and the smallest effect that sample could see with 80 percent power is about 38 percent relative.

Days required per metric in the logo testHorizontal duration bars at 12,000 visitors a week to detect a 15 percent relative effect. Trial start, 7 percent baseline: 9,907 visitors per variation and 12 days. Paid account, 1.05 percent baseline: 70,622 visitors per variation and 83 days. A vertical line marks the 14 days the test ran.the test was sized for trials, and the decision is about paid accountstrial start7% baseline, 9,907 per variation12 dayspaid account1.05% baseline, 70,622 per variation83 daysthe test ran 14 days050 days15% relative, 95% confidence, 80% power, 12,000 visitors a week, two-sided
About 21 percent power for the metric that decides value. The “not significant” read on paid accounts is absence of evidence, and the interval of minus 21.6 to plus 27.9 percent relative shows how much is still unknown.

What to do with it. There are three honest paths, and the choice should have been made before the test: accept trial starts as the decision metric, with a declared guardrail on trial to paid conversion and follow-up on the cohort; run for 12 weeks, if the decision is worth that calendar; or accept, in writing, that only large effects on paid accounts will be visible. What you do not do is announce “logos lift conversion by 15 percent”. And one signal becomes a new hypothesis: trial to paid dropped from 15.00 to 13.46 percent, not significant, but in the direction of “big logos attracted trials from people who were not the typical customer”. The next variation tests logos from companies the same size as the typical visitor. How to put a value on this kind of winner is covered in attributing revenue to your A/B test winner.

Checklist before testing social proof

  1. Truth first. Every number, score, logo, testimonial and counter on display is real, provable and authorized.
  2. Doubt mapped. Which visitor uncertainty does the proof answer at that point on the page?
  3. Primary on the transaction, guardrail on the value that stays. Clicks on the proof are only diagnostics.
  4. Sample sized for the metric that decides, with the final read scheduled after maturation (return window, end of trial).
  5. Full weeks, without stopping on the first good week.
  6. Segments declared up front, two or three at most, each with a reason.
  7. Daily log of the displayed proof (score, count, customer number) and of where the control already sees it.
  8. SRM checked before reading any effect, especially when the review block comes from an app that loads late.

Common mistakes

Automate this with Donnu

The specific pain of testing social proof is that the winner shows up fast on the wrong metric and the damage shows up slowly on the right one. No tool fixes that by magic, but a few configuration choices make the mistake less likely.

In Donnu, a conversion goal can be confirmed from your own server (server-to-server), in addition to clicks, form submissions, page visits and custom events. That lets you measure the transaction that actually happened, such as a paid order confirmed by the payment gateway, instead of a button click. On the Pro plan, a campaign can be restricted by device and by visitor type, new or returning, which helps you declare the segment before running instead of hunting for it afterwards. The report is Bayesian and shows the probability that the variation beats the control, with the confidence interval of the lift; on the Pro plan that probability is also charted day by day, so a first-week lead that crosses the 95 percent band and comes back reads as noise instead of turning into a decision.

What stays on you: making sure the proof on display is true, waiting for the guardrail to mature, and sizing for the metric that decides. For that, the sample size calculator, the significance calculator and the revenue per visitor calculator run the math in this guide with your numbers, for free.

References

Read next: Product page A/B testing · SaaS pricing page optimization · B2B SaaS landing page A/B testing · Ecommerce checkout optimization · Novelty effect · Guardrail metrics · Significance calculator · Leia em português

Frequently asked questions

What is social proof on a website?
It is any signal that other people have already chosen, used or approved what the visitor is evaluating: review scores and counts, customer or user numbers, testimonials, client logos, real-time activity, badges and certifications, and user-generated content. It works by reducing uncertainty, so it tends to matter more for expensive purchases, unfamiliar brands and risky decisions, and less when the visitor arrives already decided.
Do reviews increase conversion?
Often, but one of the most quoted numbers does not come from an A/B test. The Spiegel Research Center report at Northwestern, built on PowerReviews data, says the purchase likelihood of a product with five reviews is 270 percent greater than that of a product with none, based on products that accumulated reviews over time. That is strong observational evidence, not a promise for your store. The way to know for your case is to test, and to measure orders that are kept, not just orders.
Which metric should decide a social proof test?
The one closest to money that you can actually size for. Clicks on the stars or the testimonial are diagnostics, never the primary metric. In ecommerce, the primary is completed orders and the guardrail is orders not returned, or revenue net of returns. In SaaS, the primary is usually trial starts and the metric that decides value is paid accounts. In this guide, a review badge lifts orders by 8.75 percent with a p-value of 0.0010 and kept orders by only 0.99 percent, with a p-value of 0.7280.
Can I show a counter of people viewing a product right now?
Only if the number is real and measured. A Princeton study of roughly 53,000 product pages across roughly 11,000 stores found 313 activity notifications, and 29 of them, on 20 sites, were deceptive, most generated by a random number generator or hard-coded so they never changed. In the United States, the FTC rule on fake reviews and testimonials has been in effect since October 21, 2024. In Brazil, the Consumer Defense Code prohibits advertising that is wholly or partly false and puts the burden of proving truthfulness on the advertiser.
How much traffic does a social proof test need?
It depends on the baseline and the effect you want to see. At a 2.4 percent product page conversion rate, 95 percent confidence and 80 percent power, detecting an 8 percent relative lift takes 103,633 visitors per variation, 25 days at 60,000 visitors a week. On a SaaS pricing page with a 7 percent trial start rate, detecting 15 percent relative takes 9,907 per variation. Measuring paid accounts at a 1.05 percent baseline, same relative effect, takes 70,622 per variation, 83 days at 12,000 visitors a week.
Do customer logos work on a SaaS pricing page?
They can lift trial starts without moving paid accounts, which is what happens in the worked example in this guide: plus 15.0 percent in trials with a p-value of 0.0020, and plus 3.2 percent in paid accounts with a p-value of 0.8015, in a test that had only about 21 percent power for that second metric. A logo works as proof when the visitor recognizes the company and identifies with it, so test which logos, for which segment, rather than logos versus no logos.