A/B Testing Your Website: What to Test First
Most A/B tests fail to reach a clear winner. Not because the tool broke, but because the test was too small to matter, the traffic was too thin to prove anything, or the team called a winner three days in on a coin flip. You burn a few weeks, learn nothing, and quietly stop testing.
That waste is avoidable. The hard part of A/B testing is not the software. It is choosing what to test, sizing the test so the result means something, and reading the data without lying to yourself. This guide walks through all three, with a clear answer to the question everyone asks first: what should I test before anything else?
The short version: test the few pages that touch the most revenue, change one big thing at a time, and only trust results that clear a real significance bar. The rest of this article is how to do that on a B2B site where traffic is rarely huge and the sales cycle is long.
What A/B testing actually measures
You split incoming traffic between two versions of a page. Version A is what you have now (the control). Version B is your change (the variant). The tool randomly assigns each visitor to one version and tracks who completes the goal, usually a form submission, a demo request, or a click to the next step.
After enough visitors, you compare conversion rates. If B beats A by a margin that is unlikely to be random noise, B wins and becomes the new control. That is the whole loop.
The trap hides in "unlikely to be random noise." Conversion rates bounce around day to day on their own. A page that converts at 3% one week and 4% the next has not changed; that is just variance. A/B testing exists to separate a real lift from that natural wobble. Skip the math and you will ship changes that did nothing, or kill changes that would have worked.
Before you test anything: do you have the traffic?
This is the question that saves most B2B teams from wasting a quarter. A/B testing needs volume. Not Google-scale volume, but enough conversions to tell signal from noise.
A rough rule: you want at least a few hundred conversions per variant over the test window, and ideally the test wraps within two to four weeks. Run the math backward. If your page gets 1,000 visitors a month and converts at 2%, that is 20 conversions a month, 10 per variant. You would need many months to reach significance on a modest lift, by which point seasonality and other changes have polluted the result.
If that is your situation, do not start with classic A/B testing. Do these instead:
- Fix the obvious problems first. A slow page, a broken mobile layout, or a 12-field form does not need a test. It needs fixing. Our guide on how to improve your website conversion rate covers the changes worth making before you split traffic at all.
- Test bigger, bolder changes. Small tweaks need huge samples to prove. A full redesign of the page produces a larger effect that shows up faster.
- Use qualitative tools. Heatmaps, session recordings, and a few user interviews tell you where people get stuck without needing statistical power.
Low traffic is not a reason to skip optimization. It is a reason to optimize differently.
What to test first: follow the money
When teams ask what to test first, they usually expect a list of clever tactics. Wrong starting point. Start with a map of where revenue enters your site, then test the pages on that path in order of how much traffic and money they touch.
Here is the priority order that works for most B2B sites.
1. Your highest-traffic conversion page
Find the single page that carries the most goal completions: usually a primary landing page, a pricing page, or a demo request page. A 10% lift here is worth more than a 50% lift on a page nobody visits. This is where you have both the stakes and the sample size to get a fast, trustworthy answer.
2. The primary call to action and the form
The CTA and the form are the moments people decide to act or leave. They are small surface areas with outsized impact, which makes them ideal early tests. Try the offer itself ("Get a free audit" versus "Book a demo"), the button copy, and above all the form length. Cutting a B2B form from nine fields to four is one of the most reliable conversion wins there is. For deeper detail, see how to build landing pages for PPC that hold the visitor through to the form.
3. The headline and the offer above the fold
The first screen decides whether people stay. Your headline carries most of that weight. Testing a clearer, more specific value proposition against a vague one often moves the number more than any button color ever will. Most visitors read the headline and almost nothing else, so a stronger one compounds across everything downstream.
4. Page structure and proof
Once the big levers are tested, work on the order of sections, the trust signals (logos, case studies, numbers), and how objections get answered. These produce smaller, steadier gains.
Notice what is not at the top of this list: button colors, font tweaks, microcopy on a secondary page. Those famous "we changed the button to green and got 21%" stories are mostly survivorship bias. Color wins are real occasionally, but they are a rounding error next to a sharper offer or a shorter form. Save them for when you have run out of bigger things to try.
| What you test | Typical effort | Typical impact | When to do it |
|---|---|---|---|
| Offer / value proposition | Medium | High | First |
| Form length and fields | Low | High | First |
| Headline above the fold | Low | High | First |
| CTA copy and placement | Low | Medium | Early |
| Page structure / section order | Medium | Medium | After the basics |
| Button color, fonts, spacing | Low | Low | Last |
How to run a test that you can trust
A clean test follows a simple discipline. Skip any step and the result becomes a guess wearing a number.
Start with a hypothesis, not a tweak. Write it as: "Because [observation], we believe [change] will [outcome] measured by [metric]." For example: "Because session recordings show people abandoning the form at the phone field, we believe making phone optional will lift form completions, measured by submission rate." A hypothesis forces you to predict, which is what makes a result a learning instead of a shrug.
Change one thing per test. If you change the headline, the form, and the image at once and the variant wins, you do not know which change did it. You cannot bank the lesson. Single-variable tests are slower, but each one teaches you something durable about your audience. The exception is a full redesign where you are testing the whole concept on purpose, and accept that you will not learn which element drove the result.
Pick one primary metric. Decide before launch what counts as winning, and make it the closest metric to revenue you can measure. A click is weaker than a form submit; a form submit is weaker than a qualified lead. The further your test metric sits from money, the easier it is to "win" on a number that does not pay.
Calculate the sample size up front. Use any free A/B test calculator. Enter your current conversion rate and the smallest lift worth detecting (say, a relative 15% improvement). It tells you how many visitors per variant you need. Now you know whether the test is feasible and roughly how long it runs. This single step prevents most false results.
Run for full weeks, and do not peek. Buyer behavior differs by weekday. A test that runs Tuesday to Friday misses your weekend pattern and your Monday spike. Run in whole weeks, usually two to four, and resist calling it early. "Peeking" and stopping the moment you see significance inflates your false positive rate badly. Set the end date in advance and honor it.
The chart everyone misreads
Day 3: Variant +18% ← "We won!" (sample tiny, noise huge)
Day 8: Variant +4% ← lift shrinking toward reality
Day 14: Variant +6% ← stabilizing, now near significance
Day 21: Variant +5% ← call it: a real, modest win
Early leads are almost always exaggerated. The lift you see on day three is mostly noise, and it shrinks as data accumulates. Patience is the cheapest way to improve your test accuracy.
Statistical significance without the jargon
You will see two numbers in every testing tool.
Confidence (or significance). Usually expressed as a target of 95%. It answers: if there were truly no difference between A and B, how often would I see a gap this big by chance? At 95% confidence, that false-alarm rate is about 5%. Treat 95% as the floor for acting on a result, not a number to chase by stopping early.
Statistical power. This is the test's ability to detect a real effect that exists. It depends on sample size. A test with too few visitors can miss a genuine win entirely and show "no difference" when there was one. This is why the sample-size calculation matters as much as the confidence threshold.
One honest caveat: with B2B traffic, you will sometimes run a well-designed test that simply ends inconclusive. No clear winner, no clear loser. That is a legitimate outcome, not a failure. It tells you the change did not matter much, which frees you to test something bolder next.
Wiring up the tracking
A test is only as good as the conversion it measures. Before you launch, confirm the goal fires correctly in your analytics. If you are on GA4, make sure the form submission or click is set up as a tracked event, and that it counts once per real conversion, not on every page load. Our walkthrough on setting up events and conversions covers the common mistakes that quietly corrupt test data.
Two things matter most here. First, the goal should reflect real intent: a completed lead form, not a button hover. Second, if your sales cycle is long, connect the test back to lead quality where you can. A variant that doubles form fills but fills them with tire-kickers is a loss, even if the tool calls it a win. The deeper your closed-loop tracking, the more you can judge tests by qualified pipeline rather than raw conversions.
A simple testing cadence
Treat testing as a habit, not a one-off project. A workable monthly rhythm for a small team:
- Pull data and recordings to find the biggest leak in your funnel.
- Write one hypothesis for the highest-impact page.
- Calculate sample size and confirm the test can finish in a few weeks.
- Build the variant, double-check tracking, launch.
- Wait the full window. No peeking.
- Read the result, document what you learned, ship the winner or move on.
Even at one test a month, you compound. Twelve disciplined tests a year, each banking a real lesson about your buyers, will beat fifty rushed ones that prove nothing. If you want the broader framework this sits inside, our CRO step-by-step guide connects testing to the rest of conversion work.
Frequently asked questions
How much traffic do I need for A/B testing?
Enough to collect a few hundred conversions per variant within two to four weeks. Work it backward from your conversion rate: low-traffic pages need either very large effects or a different approach (heatmaps, user testing, fixing obvious problems) rather than classic split testing.
How long should I run an A/B test?
Run in whole weeks, usually two to four, and set the end date before you start. Stopping early the moment you see a lead is the single most common way teams fool themselves, because early results exaggerate the difference and shrink toward reality as data accumulates.
What should I test first on my website?
The highest-traffic page that touches revenue, then its CTA and form, then the headline and offer above the fold. Test the few things that move money before tweaking colors or fonts on minor pages.
Can I test more than one change at a time?
For a standard A/B test, change one variable so you know what caused the result. Testing many elements at once is multivariate testing, which needs far more traffic. The exception is a deliberate full-page redesign, where you accept that you will learn the concept won, not which element did it.
What does statistical significance mean in plain terms?
It is the chance your result is real and not random luck. A 95% confidence level means there is roughly a 5% chance you would see a gap this large even if the two versions were identical. It is the minimum bar before acting, not a number to chase by ending the test early.
What if the test ends with no clear winner?
That is a normal outcome, especially in B2B. It usually means the change was too small to matter. Document it, then test something bolder: a sharper offer or a real structural change rather than a cosmetic one.
The takeaway
A/B testing rewards discipline over cleverness. Pick the page that touches the most revenue, change one meaningful thing, size the test so the answer means something, and wait for the full window before you believe it. Do that and even a low-traffic B2B site can stack up real, compounding wins.
Quick checklist before your next test:
- One clear hypothesis tied to an observation.
- One variable changed, one primary metric chosen.
- Sample size calculated; test can finish in two to four weeks.
- Conversion tracking verified before launch.
- End date set in advance, and no peeking.
- 95% confidence as the floor for acting.
If you would rather not run this alone, that is what we do. We can audit one high-value page on your site and map the two or three tests most likely to move qualified leads, in about 30 minutes and with no commitment. Reach out and tell us which page matters most to your pipeline, and we will show you where to start.