A/B Test Hypothesis Generator
Fill in the fields and get a finished hypothesis in the If, Then, Because, Measured by format, plus a quality score that points at exactly what is still missing. Free, live, no signup.
Most A/B tests that teach nothing were already dead at the hypothesis stage. Not because the statistics failed, but because nobody wrote down what was supposed to be learned. A test without a hypothesis produces one answer: it went up or it went down. A test with a hypothesis produces reusable knowledge that outlives the result and points at the next experiment. This tool writes the hypothesis for you from the fields that matter, then does the uncomfortable part: it scores how rigorous the claim is and tells you where it is loose.
Your hypothesis
One line version (for the backlog card)
What holds it up (and what is missing)
- Specific audience and page-
- Concrete, single change-
- Expected effect on behavior-
- Evidence behind the bet-
- Reason explained, not just asserted-
- Primary metric defined-
- Baseline and expected effect as numbers-
- Guardrail metrics declared-
Nothing here leaves your browser: the fields are never sent anywhere. The score is a rigor check, not an oracle: it verifies that the hypothesis is specific, falsifiable and anchored in evidence, which is what separates an experiment from an expensive guess.
How to use it
- Describe the change in concrete terms. Write what the variation does, not the goal behind it. "Make the page feel more trustworthy" is not a change; "add three security badges below the form" is.
- State who and where. Audience and step change the expected effect completely: the same button behaves differently on mobile and on desktop.
- Write the expected effect on behavior, not the financial outcome. Behavior is what the variation can cause directly.
- Record the evidence and pick its source from the list. This is the heaviest item in the score, and it is where honesty pays off.
- Define one primary metric, the current baseline and the relative lift you expect. That pair of numbers is what feeds the sample size math later.
- Declare the guardrail metrics and copy the finished hypothesis into your backlog card.
How it works: the anatomy of the hypothesis and the score
The output follows the canonical experimentation format, with each clause doing a specific job:
The target is computed as baseline × (1 + relative lift / 100), because effect size in A/B testing is declared in relative terms. A 20% lift on a 3.2% rate lands at 3.84%, not at 23.2%: confusing the two is the most common mistake when sizing a test, and it inflates the expected effect until the experiment becomes impossible to detect.
The quality score adds up eight criteria with fixed weights: specific audience (15), concrete change (15), expected effect (10), source of evidence (up to 25), reason explained (10), primary metric (10), baseline and effect numbers (10) and guardrail metrics (5). Evidence outweighs any other single item because it is what separates an experimentation program from a queue of opinions: quantitative data is worth 25 points, qualitative research 20, a UX heuristic 12, a third party benchmark 10, and a hunch is worth zero.
Worked example (reproduces the default output)
The fields ship pre filled with a real product page case. The change is switching the button label from Submit to Get my free quote, the audience is mobile visitors on the product page, the expected effect is that more visitors complete the form, and the reason is session recordings showing 62% of users stalling on the button without clicking. The primary metric is lead conversion rate, currently at 3.2%, with an expected relative lift of 20%.
The target comes out at 3.2 × 1.20 = 3.84%, and that is the number you see in the Measured by clause. The score lands on 95 out of 100: seven of the eight criteria pay in full (15 + 15 + 10 + 10 + 10 + 10 + 5 = 75) and evidence, marked as qualitative research, pays 20 of the 25 available. Exactly 5 points are missing, and they sit in one place: if you also had the funnel data confirming the drop at that step, and not just the recordings, the evidence item would hit 25 and the hypothesis would reach 100.
Now switch the evidence source to a hunch and watch what happens: the score collapses to 75 without a single word of the text changing. That is deliberate. The hypothesis is still well written, specific and falsifiable; it simply stopped having any reason to exist. That is the test that burns three weeks of traffic to answer a question nobody knew why they were asking.
How to read the score, and where it does not help
The score measures rigor of formulation, not probability of winning. A 100 point hypothesis can lose badly, and that is a sign of health: it means you learned something that contradicted the available evidence, which is the most valuable outcome a testing program can produce. What the score guarantees is that, win or lose, you will know what the result means.
It also does not judge the quality of your evidence inside a category. Marking "quantitative data" based on a sample of 30 sessions earns the same 25 points as a funnel with 200,000 visitors, and the tool has no way to tell the difference. The same goes for the baseline: it accepts whatever number you type without checking whether it came from analytics or from memory. The ruler measures structure, and structure is a necessary condition, never a sufficient one.
The most important limit is this: a good hypothesis does not fix an underpowered test. Once yours is written, take the baseline and the expected lift to the sample size calculator and find out how many visitors the experiment demands. If your traffic cannot fund that effect, the problem is not the wording: use the MDE calculator to find the smallest effect you can actually detect and rewrite the expectation around it.
Five mistakes the score flags immediately
- Two primary metrics. When a test has two owners, the winner is whoever argues better after the result. Pick one and demote the other to a guardrail.
- A bundled change. "Redesign the page" contains ten changes. If it wins, you do not know which one won, and you cannot apply the learning anywhere else.
- Effect stated in revenue, not behavior. A button does not cause revenue directly; it causes clicks that cause forms that cause sales. Measure the link the change actually moves.
- A circular Because. "Because it will convert better" restates the Then in other words and is not evidence. The Because has to point at something that existed before the test.
- An expectation with no number. With no baseline and no expected effect you cannot size the test, and an unsized test ends when somebody gets tired of waiting.
Frequently asked questions
- What is an A/B test hypothesis?
- It is the testable claim your experiment will try to knock down. It has to state what changes, for whom, what behavior you expect to see, why you believe it and which number settles the question. If no result could ever prove the sentence wrong, it is not a hypothesis: it is a wish.
- What is the difference between a hypothesis and a test idea?
- An idea is "let us test the green button". A hypothesis is "if we change the button label from Submit to Get my free quote, for mobile visitors on the product page, then more visitors complete the form, because session recordings show people stalling there, measured by lead conversion rate". The idea tells you what to do. The hypothesis tells you what you learn, win or lose.
- Why the If, Then, Because, Measured by format?
- Each clause closes a different gap. The If pins down the change, so you do not alter five things and never learn which one worked. The Then pins down the expected behavior, which is what makes the claim falsifiable. The Because pins down the evidence, which separates an experiment from a guess. The Measured by pins down the primary metric before you see data, which blocks the convenient choice of number after the fact.
- How is the quality score calculated?
- Eight criteria with fixed weights adding up to 100: specific audience and page (15), concrete change (15), expected effect (10), source of evidence (up to 25), reason explained (10), primary metric (10), baseline and effect as numbers (10) and guardrail metrics (5). Evidence carries the heaviest weight on purpose: a beautifully written hypothesis backed by a hunch is still a hunch. Everything runs in your browser and nothing is sent anywhere.
- Do I really need to declare guardrail metrics?
- Yes, and it is the field most people skip. A guardrail is what must not get worse while you chase the primary metric. An aggressive popup can lift email capture and sink revenue per visitor at the same time. With no guardrail declared up front, you celebrate the partial win and take the loss home without noticing.
- Can the hypothesis change after the test starts?
- No. Swapping the primary metric or the audience after seeing data is the shortest path to a false positive, because you start picking the cut that confirms what you wanted. If the hypothesis was wrong, record that and write a new one for the next cycle. An old hypothesis, refuted and documented, is learning; a hypothesis rewritten mid flight is noise.
Keep going
The full reasoning behind each clause is in the guide how to write an A/B test hypothesis. With the hypothesis ready, size the experiment in the sample size calculator, find out how long it will run in the duration calculator and, at the end, read the verdict in the significance calculator.