Social Proof A/B Testing: What to Test and How to Measure
Social proof A/B testing: reviews, logos, testimonials and live counters. What the evidence shows, what to test, and the metric that does not lie.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
Social proof is any signal that other people have already chosen, used or approved what a visitor is evaluating: review scores and counts, customer numbers, testimonials, client logos, real-time activity, badges and user-generated content. It works by reducing uncertainty, which is why it is a common CRO hypothesis and an easy one to measure badly. The usual mistake is not picking the element, it is picking the metric: in the worked example in this guide, a review badge on a product page lifts orders by 8.75 percent with a p-value of 0.0010, while orders that are not sent back as returns rise by only 0.99 percent, with a p-value of 0.7280. This guide is part of our complete guide to conversion rate optimization (CRO) and covers what the research shows (and what it does not), what to test in ecommerce and SaaS, the five measurement mistakes, and two worked examples calculated end to end.
What social proof is and why it works
Robert Cialdini’s company sums up the principle in one sentence: especially when they are uncertain, people look to the actions and behaviors of others to determine their own. The word doing the work is uncertain. Social proof does not persuade anyone on its own; it lends other people’s experience to someone who does not yet have enough information to decide.
For testers, that means the expected effect depends on how much uncertainty exists at that point on the page. A visitor who came from a branded search, already decided, has little doubt left to reduce; a visitor looking at an expensive product from an unfamiliar brand has plenty. The same badge can have zero effect on one page and a meaningful one on another.
Goldstein, Cialdini and Griskevicius added a second condition in a 2008 field experiment: the proof worked better when the group described was in the same situation as the reader, which they called a provincial norm. On a page, “customers your size” or “people who bought this size” tends to be a stronger hypothesis than “thousands of customers”.
The seven types of social proof on a page
| type | what it signals | where it usually appears | main risk |
|---|---|---|---|
| Review score and count | other people bought and rated it | near the title and price on a product page | a perfect score with few reviews looks too good |
| Customer or user count | scale and staying power | home page, landing page, pricing page | an inflated number is deceptive advertising |
| Testimonials | concrete experience of someone similar | landing page, pricing page, checkout | generic or unverifiable testimonials |
| Client logos | known companies trusted you | B2B landing page and SaaS pricing page | logos the visitor does not recognize or relate to |
| Real-time activity | other people are acting right now | product page and cart | a randomly generated or hard-coded counter |
| Badges and certifications | a third party verified something | checkout, footer, security page | a badge with no real verification behind it |
| User-generated content | photos and videos of real use | product page gallery, social media | page weight and curation that hides the negatives |
The last column is almost entirely about credibility. Social proof is a claim about facts, which exposes it to two risks an A/B test dashboard never shows: looking fake and being fake.
What the evidence actually shows
Three sources are commonly cited on social proof, and the quality of evidence behind each is very different. Read each one for its design, not its headline number.
The towel experiment: strong, famous, and a replication that failed
In the first experiment by Goldstein, Cialdini and Griskevicius, the 190 rooms of a hotel belonging to a national US chain were randomly assigned one of two cards. The standard card talked about protecting the environment. The descriptive norm card said almost 75 percent of guests took part in the program. Counting only each guest’s first eligible day, towel reuse was 44.1 percent with the norm versus 35.1 percent with the environmental appeal, with a p-value of .05 in a chi-square test on 433 observations.
In the second experiment, with 1,595 observations at the same hotel and five cards, the four norm messages combined reached 44.5 percent against 37.2 percent for the environmental message, and the “guests who stayed in this room” norm reached 49.3 percent, versus 42.8 percent for the other three norms combined.
Cialdini’s own page presents the result as increases of 26 and 33 percent. The arithmetic matches the paper (44.1 against 35.1 is about 26 percent relative; 49.3 against 37.2 is about 33), but each number comes from a different experiment with a different control. It is a good reminder of how research turns into a headline.
Then there is the part that usually gets left out. In 2014, Bohner and Schlüter repeated the design at two hotels in Germany, published in PLoS ONE. In study 1, with 724 observations, the four norms combined produced 81.9 percent reuse against 83.7 percent for the environmental message, p = .62. In study 2, with 204 observations, the environmental message did best of all (93.3 percent). The authors’ conclusion: descriptive norm messages were not more effective than the standard message, and proximity effects were inconsistent across studies.
There is also a design detail that matters to anyone running experiments. At the US hotel, randomization was by room (190 rooms in experiment 1) and the analysis was by guest. When the randomized unit is larger than the analyzed unit, observations from the same group tend to resemble each other, and a test that treats each one as independent tends to overstate precision. We cover that math in cluster randomization and in choosing the randomization unit.
The 270 percent figure: observational, with a vendor involved
The report How Online Reviews Influence Sales, from Northwestern’s Spiegel Research Center, produced in partnership with PowerReviews (a company that sells ratings and reviews software), is the source of one of the most repeated numbers about social proof. What it says, read in full:
| finding | where it comes from | how to read it |
|---|---|---|
| purchase likelihood with five reviews is 270% greater than with none | high-end gift retailer: about 15.5 million page views, 1,800 products, 7.8 million users over a year, tracking products as they accumulated reviews | observational; a product that accumulates reviews also accumulates catalog time and sales |
| displaying reviews lifted conversion by 190% for lower-priced and 380% for higher-priced products | same retailer, comparing categories | the direction (more effect on risky purchases) transfers, the magnitude does not |
| purchase likelihood typically peaks between 4.0 and 4.7 stars and falls as ratings approach 5.0 | across product categories | the same report also says 4.2 to 4.5 in its closing summary; treat it as a range, not a target |
| nearly all of the gain happens within the first 10 reviews, and the first 5 drive most of it | gift retailer | an argument for a few reviews on many products |
| verified buyer reviews raise purchase likelihood by 15% versus anonymous ones | about 13,500 products, 57,000 anonymous and 65,000 verified buyer reviews | a verified buyer badge is a cheap hypothesis to test |
This is not an A/B test. A product that reaches five reviews is, by construction, a product that has already sold and has been in the catalog longer, so the number does not mean “add stars and sell 3.7 times more”. What transfers is the direction: reviews matter more where uncertainty is higher, a perfect score raises suspicion, and the first reviews are worth more than later ones.
Social proof that is a lie
Mathur and coauthors at Princeton crawled roughly 53,000 product pages across roughly 11,000 stores (CSCW 2019) and found 1,818 instances of dark patterns on 1,254 sites. Among the 313 activity notifications (who just bought, how many people are viewing), 29, on 20 sites, were deceptive: most generated the number with a random number generator or showed hard-coded messages. Among the testimonials of uncertain origin, in one case the same set of testimonials appeared on another store, with different customer names attached.
For testers: a fake counter can “win” an A/B test. The dashboard measures behavior, not truthfulness, so “is it true?” comes before “does it convert?”.
What to test in social proof: the ecommerce vs SaaS matrix
Because social proof works on uncertainty, the useful question is not “should we add testimonials?”, it is “what doubt does the visitor have at this point, and what proof answers it?”. The doubt differs between ecommerce and SaaS, and from one step to the next.
The matrix below organizes the most common hypotheses by the lever they pull. The “measurement trap” column is the one that matters most when you read the result.
| lever | ecommerce: what to test | SaaS: what to test | measurement trap |
|---|---|---|---|
| Placement | score with count right under the title, near the price, versus only at the bottom | logos and a testimonial above the plan table versus below it | clicks on the score rise just because it became visible |
| Specificity | reviews filtered by size or use case versus the general list | testimonial from the same company size or industry versus a generic one | segment defined after looking at the data |
| Score format | average with count versus average alone; showing the star distribution | customer count versus logos | high score with very few reviews |
| Credibility | verified buyer badge; showing negative reviews | testimonial with name, role and company versus anonymous | a conversion gain that turns into returns or churn |
| Activity | “sold in the last 24 hours”, if real | “companies that started this month”, if real | a counter that changes on its own during the test |
| Badges and compliance | security badge next to the card field | GDPR or LGPD information, data location, real certifications | a badge both arms already see through another path |
| User content | customer photo gallery on the product page | case study with a verifiable number | page weight changes along with the proof |
Two hypotheses cut against intuition. Showing negative reviews: the Spiegel report links near-perfect scores to lower purchase likelihood, and our product page A/B testing guide argues that negatives calibrate expectations and tend to reduce returns. Swapping big-brand logos for logos that look like the visitor: the SaaS pricing page guide describes the wall of famous brands that convinces a small company the product is not for them, and the B2B SaaS landing page guide recommends testing which logos, for which segment.
A technical note for ecommerce: Google asks that review content marked up with structured data be readily available to users on the marked-up page. Hiding reviews in the control while keeping the markup conflicts with that. Testing placement, format and emphasis is safer than testing whether reviews exist at all.
How to measure social proof without fooling yourself
Social proof is cheap to implement and expensive to measure. Five problems show up in almost every test of this kind, and each one manufactures a false winner in a different way.
1. The wrong metric: clicks, conversion, revenue or orders that stay
The most tempting metric is the one that moves first: clicks on the stars, opens of the testimonial carousel. It explains why a variation won and does not decide whether it won, because any visual emphasis increases clicks.
The rule of thumb: primary metric on the transaction step, guardrail on the value that stays. In ecommerce, completed orders decide and orders not returned (or revenue net of returns) protect. In SaaS, trial start is usually the feasible primary, and paid accounts decide value, at a much larger sample. The guardrail metrics guide shows how to fix thresholds before the test, and the sample size guide for revenue and continuous metrics shows why revenue per visitor needs even more sample than conversion.
2. Novelty effect: a new badge draws attention because it is new
A review block that suddenly appears on a page returning customers know well is, in its first week, mostly novelty, and the early gain can shrink. To separate novelty from effect, compare new and returning visitors and look at the week by week trajectory, as in novelty effect in A/B testing.
3. Contamination: real social proof changes during the test
All genuine social proof is a live number: the count grows, the average moves, a negative review lands at the top. That creates three problems:
- The treatment changes mid-test. A variation showing 4.8 stars from 12 reviews in week 1 may show 4.5 from 40 in week 4. You are testing an average of versions.
- The control is exposed too. Stars can appear in search results, and reviews travel through marketplaces and social media. Part of the control arrives having already seen the proof, and the measured effect gets diluted.
- The catalog changes. In a template test, new products with no reviews enter the sample over time.
None of this invalidates the test, but it changes the read: log the displayed score and count every day, and know that the result holds for the range of values that actually appeared.
4. Heterogeneity: new vs returning, mobile vs desktop
It is reasonable to expect social proof to help new visitors more, and a long testimonial block to behave differently on a small screen. The danger is discovering this afterwards: with ten independent cuts each tested at 5 percent, the chance of at least one false positive is above 40 percent. Declare two or three segments with a reason up front and read the difference with the right math, covered in heterogeneous treatment effects and Simpson’s paradox.
5. Fake social proof: the test cannot see it, the law can
| jurisdiction | rule | what it says, in short | what it changes in a test |
|---|---|---|---|
| United States | FTC Rule on the Use of Consumer Reviews and Testimonials (16 CFR Part 465), announced August 14, 2024, in effect since October 21, 2024 | bans fake or false consumer reviews and testimonials (including AI-generated ones), compensation conditioned on a particular sentiment, undisclosed insider reviews, suppressing reviews through threats, and buying or selling fake indicators of social media influence such as followers or views; allows civil penalties for knowing violations | a variation with cherry-picked reviews, an invented counter or an employee testimonial is not a hypothesis, it is exposure |
| Brazil | Consumer Defense Code, Law 8,078/1990, articles 37, 38 and 67 | bans misleading advertising, including partly false or by omission; the burden of proving truthfulness falls on the advertiser; advertising known or presumed to be misleading carries three months to one year of detention plus a fine | customer counts, scores and counters must be provable in both arms |
| Brazil | Conar, Brazilian Advertising Self-Regulation Code, Annex Q (testimonials) | an identified consumer’s first and last name must be real; the advertiser’s employees must not pose as ordinary consumers; the advertiser must prove a testimonial is genuine when asked | an unauthorized or unverifiable testimonial does not go into a variation |
In its questions and answers about the rule, the FTC clarifies that organizing reviews is not suppressing them, but organizing them in a way that makes negative reviews hard to find could be deceptive under the FTC Act. For anyone testing review ordering, that is the line. This is context, not legal advice.
Worked example 1: the review badge on a product page
Scenario (illustrative). A fashion store with 60,000 weekly visitors on product pages converts 2.4 percent of visits into orders. Reviews exist today but sit at the bottom of the page. The hypothesis: showing the average score with the review count right under the title, near the price, reduces doubt at the moment it appears and increases orders.
Metrics declared up front. Primary: visitor with a completed order. Guardrail: visitor with an order not returned within 30 days (in this scenario, each buyer places one order). Diagnostic: clicks on the score. Segments: new and returning. All at the template level, across the catalog, because a single product page would never have enough traffic.
Sample size. The team wants to detect an 8 percent relative lift. In the calculator below, enter: current conversion rate 2.4; minimum detectable effect 8, relative; confidence 95; power 80; visitors per week 60000; test two-sided.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The screen shows 103,633 visitors per variation, 207,266 in total and 25 days. The team rounds up to four full weeks, 28 days, which gives 120,000 visitors per arm. To feel the sensitivity, change only the minimum effect:
| relative effect you want to detect | visitors per variation | total | days at 60,000 a week |
|---|---|---|---|
| 5% | 261,572 | 523,144 | 62 |
| 8% | 103,633 | 207,266 | 25 |
| 10% | 66,946 | 133,892 | 16 |
| 12% | 46,922 | 93,844 | 11 |
The peek you should not act on. At the end of week one, with 30,000 visitors per arm, there were 720 orders in the control and 834 in the variation: plus 15.8 percent, p-value 0.0034. Stopping there would have recorded almost double the lift the full test showed. In weeks 2 to 4, with 90,000 per arm, the score was 2,160 versus 2,298, plus 6.4 percent, p-value 0.0364. A stronger first week is consistent with novelty and also with noise; what it does not justify is a decision. The cost of looking early is in the peeking problem.
The primary metric result, after 28 days.
| arm | visitors | orders | rate |
|---|---|---|---|
| A, reviews at the bottom of the page | 120,000 | 2,880 | 2.40% |
| B, score with count near the price | 120,000 | 3,132 | 2.61% |
The split is clean: 120,000 versus 120,000 raises no flag in the SRM checker. Now paste into the significance calculator: control with 120000 visitors and 2880 conversions; variation with 120000 visitors and 3132 conversions; confidence 95.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The screen shows a 2.40% rate for the control and 2.61% for the variation, a relative lift of +8.8%, a p-value of 0.0010, a 95 percent CI of the difference of +0.1% … +0.3% (pp) and the verdict Significant winner · B wins. With more decimals, the lift is 8.75 percent, the p-value is 0.000997 and the interval runs from plus 0.0850 to plus 0.3350 percentage points, or roughly plus 3.5 to plus 14.0 percent in relative terms (dividing the interval by the control rate).
So far, the usual story: B won, comfortably. That is 252 extra orders.
The twist: a guardrail that only matures 30 days later. Orders with a 30-day return window only close almost two months after the test started. When they do:
| arm | orders | returned | return rate | visitors with a kept order | rate per visitor |
|---|---|---|---|---|---|
| A | 2,880 | 461 | 16.01% | 2,419 | 2.02% |
| B | 3,132 | 689 | 22.00% | 2,443 | 2.04% |
Of the 252 extra orders, 228 came back. To check both reads in the same calculator, first enter orders as “visitors” and returns as “conversions” (2880 and 461 versus 3132 and 689): the screen shows 16.01% versus 22.00%, a lift of +37.4%, a p-value below 0.0001 (the screen prints a less-than sign before the number), a CI of +4.0% … +8.0% (pp) and the verdict “Significant winner · B wins”. Here the calculator only says that B’s rate is higher, and on this metric higher is worse.
Then enter the guardrail: 120000 and 2419 versus 120000 and 2443. The screen shows 2.02% versus 2.04%, a lift of +1.0%, a p-value of 0.7280, a CI of -0.1% … +0.1% (pp) and Not significant yet. With more decimals, the lift is 0.99 percent and the interval runs from minus 0.0927 to plus 0.1327 percentage points, roughly minus 4.6 to plus 6.6 percent relative.
What the result says. The badge pushed people into buying who used to scroll to the reviews, read about sizing, and leave. In this scenario, the diagnostics back that reading: clicks on “see reviews” fell in the variation, and the most common return reason in B is fit. The badge answered the wrong doubt (can I trust this store) when the real doubt was size. And it is not just a lack of sample: for an 8 percent gain in kept orders, the 2.02 percent baseline needs 123,628 visitors per variation (29 days), and the test had 120,000.
What to do with it. Do not ship B as is. The next variation pairs the score near the price with a review excerpt filtered by size, which answers the doubt that was driving returns. Same primary, same guardrail, and this time with the final read scheduled after the return window.
Worked example 2: client logos on a SaaS pricing page
Scenario (illustrative). A B2B SaaS gets 12,000 weekly visitors on its pricing page, and 7 percent of them start a trial. About 15 percent of trials become paid accounts within 30 days, which works out to 1.05 percent paid accounts per visitor. The hypothesis: a strip of well known client logos above the plan table reduces the perceived risk of subscribing.
Sample size. In the same sample size calculator above, change the fields: current rate 7; minimum effect 15, relative; confidence 95; power 80; visitors per week 12000; two-sided. The screen shows 9,907 visitors per variation, 19,814 in total and 12 days. The team runs two full weeks: 12,000 visitors per arm.
The trial result. In the significance calculator, enter 12000 and 840 for the control and 12000 and 966 for the variation. The screen shows 7.00% versus 8.05%, a lift of +15.0%, a p-value of 0.0020, a CI of +0.4% … +1.7% (pp) and Significant winner · B wins. That is 126 extra trials.
The paid account result, 30 days later.
| metric | A | B | relative lift | p-value | read |
|---|---|---|---|---|---|
| trial starts per visitor | 840 of 12,000 (7.00%) | 966 of 12,000 (8.05%) | +15.0% | 0.0020 | significant |
| trials that convert to paid | 126 of 840 (15.00%) | 130 of 966 (13.46%) | -10.3% | 0.3486 | not significant |
| paid accounts per visitor | 126 of 12,000 (1.05%) | 130 of 12,000 (1.08%) | +3.2% | 0.8015 | not significant |
To check the last row, enter 12000 and 126 versus 12000 and 130: the screen shows 1.05% versus 1.08%, a lift of +3.2%, a p-value of 0.8015, a CI of -0.2% … +0.3% (pp) and Not significant yet. With more decimals, the interval runs from minus 0.2266 to plus 0.2933 percentage points, something like minus 21.6 to plus 27.9 percent relative. Of the 126 extra trials, 4 became extra paid accounts.
The twist here is different from the ecommerce one. The paid account result does not say logos do not work. It says the test was never able to answer. Go back to the sample size calculator and enter the metric that decides revenue: rate 1.05, effect 15 relative, 95, 80, 12000 a week. The screen shows 70,622 per variation, 141,244 in total and 83 days. The calculator does not show power, so the next numbers are this guide’s own calculation with the same formula: with 12,000 per arm, the power to detect a 15 percent lift in paid accounts was about 21 percent, and the smallest effect that sample could see with 80 percent power is about 38 percent relative.
What to do with it. There are three honest paths, and the choice should have been made before the test: accept trial starts as the decision metric, with a declared guardrail on trial to paid conversion and follow-up on the cohort; run for 12 weeks, if the decision is worth that calendar; or accept, in writing, that only large effects on paid accounts will be visible. What you do not do is announce “logos lift conversion by 15 percent”. And one signal becomes a new hypothesis: trial to paid dropped from 15.00 to 13.46 percent, not significant, but in the direction of “big logos attracted trials from people who were not the typical customer”. The next variation tests logos from companies the same size as the typical visitor. How to put a value on this kind of winner is covered in attributing revenue to your A/B test winner.
Checklist before testing social proof
- Truth first. Every number, score, logo, testimonial and counter on display is real, provable and authorized.
- Doubt mapped. Which visitor uncertainty does the proof answer at that point on the page?
- Primary on the transaction, guardrail on the value that stays. Clicks on the proof are only diagnostics.
- Sample sized for the metric that decides, with the final read scheduled after maturation (return window, end of trial).
- Full weeks, without stopping on the first good week.
- Segments declared up front, two or three at most, each with a reason.
- Daily log of the displayed proof (score, count, customer number) and of where the control already sees it.
- SRM checked before reading any effect, especially when the review block comes from an app that loads late.
Common mistakes
- Quoting 270 percent as a forecast. It is observational, from one retailer, with a reviews vendor as a partner. It gives direction, not a target.
- Copying the towel experiment without reading the replication. The same design did not repeat the effect in Germany.
- Hiding reviews in the control. It tests the least useful question and conflicts with Google’s guideline for marked-up content.
- Showing only positive reviews by default. A perfect score raises suspicion, and making negatives hard to find could be deceptive according to the FTC.
- A wall of big-brand logos for an audience of small companies. The visitor concludes the product is not for them.
- Changing the proof and the layout at the same time. The result becomes impossible to explain or reuse.
Automate this with Donnu
The specific pain of testing social proof is that the winner shows up fast on the wrong metric and the damage shows up slowly on the right one. No tool fixes that by magic, but a few configuration choices make the mistake less likely.
In Donnu, a conversion goal can be confirmed from your own server (server-to-server), in addition to clicks, form submissions, page visits and custom events. That lets you measure the transaction that actually happened, such as a paid order confirmed by the payment gateway, instead of a button click. On the Pro plan, a campaign can be restricted by device and by visitor type, new or returning, which helps you declare the segment before running instead of hunting for it afterwards. The report is Bayesian and shows the probability that the variation beats the control, with the confidence interval of the lift; on the Pro plan that probability is also charted day by day, so a first-week lead that crosses the 95 percent band and comes back reads as noise instead of turning into a decision.
What stays on you: making sure the proof on display is true, waiting for the guardrail to mature, and sizing for the metric that decides. For that, the sample size calculator, the significance calculator and the revenue per visitor calculator run the math in this guide with your numbers, for free.
References
- Goldstein, N. J., Cialdini, R. B. and Griskevicius, V. A Room with a Viewpoint: Using Social Norms to Motivate Environmental Conservation in Hotels. Journal of Consumer Research, 35(3), 2008. Source for both field experiments (randomization by room, rates of 44.1 versus 35.1 and of 49.3, 44.5 and 37.2 percent) and for the provincial norm concept. PDF read in full. sparq.stanford.edu.
- Bohner, G. and Schlüter, L. E. A Room with a Viewpoint Revisited: Descriptive Norms and Hotel Guests’ Towel Reuse Behavior. PLoS ONE, 2014. Source for the replication in Germany (81.9 versus 83.7 percent, p = .62) and the conclusion that descriptive norm messages were not more effective. Article read in full. pmc.ncbi.nlm.nih.gov.
- Spiegel Research Center, Northwestern University, with PowerReviews. How Online Reviews Influence Sales. 2017. Source for the 270, 190 and 380 percent figures, the star rating ranges, the weight of the first reviews, the 15 percent verified buyer effect and the datasets. PDF read in full. spiegel.medill.northwestern.edu.
- Influence at Work (Robert Cialdini’s company). Principles of Persuasion. Source for the definition of the social proof principle (especially when uncertain, people look to others’ behavior) and for presenting the hotel results as increases of 26 and 33 percent. Checked on September 15, 2026. influenceatwork.com.
- Mathur, A., Acar, G., Friedman, M. J., Lucherini, E., Mayer, J., Chetty, M. and Narayanan, A. Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites. Proceedings of the ACM on Human-Computer Interaction, CSCW, 2019. Source for the store crawl, the 313 activity notifications with 29 deceptive ones on 20 sites, and testimonials of uncertain origin. PDF read in full. arxiv.org.
- Federal Trade Commission. Federal Trade Commission Announces Final Rule Banning Fake Reviews and Testimonials (August 14, 2024) and The Consumer Reviews and Testimonials Rule: Questions and Answers. Source for the prohibited practices, civil penalties for knowing violations, the October 21, 2024 effective date, and the distinction between organizing and suppressing reviews. Checked on September 15, 2026. ftc.gov · ftc.gov.
- Brazil. Law No. 8,078 of September 11, 1990 (Consumer Defense Code), consolidated text, in Portuguese. Source for articles 37 (misleading advertising, including partial or by omission), 38 (burden of proof on the advertiser) and 67 (three months to one year of detention plus a fine). Checked on September 15, 2026. planalto.gov.br.
- Conar. Brazilian Advertising Self-Regulation Code, Annex Q: Testimonials, Attestations, Endorsements, in Portuguese. Source for the rules on real names of identified consumers, employees not posing as ordinary consumers, and proving that a testimonial is genuine. Checked on September 15, 2026 in the current edition published by Conar (2025 edition), where Annex Q keeps these items. conar.org.br.
- Google Search Central. Review snippet (Review, AggregateRating) structured data. Source for the guideline that marked-up review content must be readily available to users on the marked-up page. Checked on September 15, 2026. developers.google.com.
Read next: Product page A/B testing · SaaS pricing page optimization · B2B SaaS landing page A/B testing · Ecommerce checkout optimization · Novelty effect · Guardrail metrics · Significance calculator · Leia em português
Frequently asked questions
- What is social proof on a website?
- It is any signal that other people have already chosen, used or approved what the visitor is evaluating: review scores and counts, customer or user numbers, testimonials, client logos, real-time activity, badges and certifications, and user-generated content. It works by reducing uncertainty, so it tends to matter more for expensive purchases, unfamiliar brands and risky decisions, and less when the visitor arrives already decided.
- Do reviews increase conversion?
- Often, but one of the most quoted numbers does not come from an A/B test. The Spiegel Research Center report at Northwestern, built on PowerReviews data, says the purchase likelihood of a product with five reviews is 270 percent greater than that of a product with none, based on products that accumulated reviews over time. That is strong observational evidence, not a promise for your store. The way to know for your case is to test, and to measure orders that are kept, not just orders.
- Which metric should decide a social proof test?
- The one closest to money that you can actually size for. Clicks on the stars or the testimonial are diagnostics, never the primary metric. In ecommerce, the primary is completed orders and the guardrail is orders not returned, or revenue net of returns. In SaaS, the primary is usually trial starts and the metric that decides value is paid accounts. In this guide, a review badge lifts orders by 8.75 percent with a p-value of 0.0010 and kept orders by only 0.99 percent, with a p-value of 0.7280.
- Can I show a counter of people viewing a product right now?
- Only if the number is real and measured. A Princeton study of roughly 53,000 product pages across roughly 11,000 stores found 313 activity notifications, and 29 of them, on 20 sites, were deceptive, most generated by a random number generator or hard-coded so they never changed. In the United States, the FTC rule on fake reviews and testimonials has been in effect since October 21, 2024. In Brazil, the Consumer Defense Code prohibits advertising that is wholly or partly false and puts the burden of proving truthfulness on the advertiser.
- How much traffic does a social proof test need?
- It depends on the baseline and the effect you want to see. At a 2.4 percent product page conversion rate, 95 percent confidence and 80 percent power, detecting an 8 percent relative lift takes 103,633 visitors per variation, 25 days at 60,000 visitors a week. On a SaaS pricing page with a 7 percent trial start rate, detecting 15 percent relative takes 9,907 per variation. Measuring paid accounts at a 1.05 percent baseline, same relative effect, takes 70,622 per variation, 83 days at 12,000 visitors a week.
- Do customer logos work on a SaaS pricing page?
- They can lift trial starts without moving paid accounts, which is what happens in the worked example in this guide: plus 15.0 percent in trials with a p-value of 0.0020, and plus 3.2 percent in paid accounts with a p-value of 0.8015, in a test that had only about 21 percent power for that second metric. A logo works as proof when the visitor recognizes the company and identifies with it, so test which logos, for which segment, rather than logos versus no logos.