Knowledge Hub
The SEO Split Test You Can Run Without 30,000 Sessions A Month
By John Butterworth · August 25, 2026
You rewrote the title tags on forty collection pages last month, and organic sessions are up eleven per cent. You can't say whether that was your rewrite or the season. Settling that is what SEO AB testing does, written by most of the field as SEO A/B testing, and someone is about to ask you for the answer.
An SEO split test splits a group of pages. Visitors are left alone. Half your collection pages get the new title and half keep the old one. Both halves are watched over the same weeks, so anything happening to the whole site shows up in both.
I'm John Butterworth. I have spent 11 years in SEO and driven more than 3M visits, and the stores I run retainers for sit at exactly the size where a split test gets awkward. Below is which split test your store can run without 30,000 organic sessions a month, what to point it at, and how to read the answer when it comes back.
Start With The Test Design Your Traffic Can Support
Four designs are available to you, and your organic traffic decides which. Work down the list and stop at the first one your store can support.
You Need Same-Template Pages And 30,000 Organic Sessions
SearchPilot runs controlled SEO experiments for large sites and publishes a new one every fortnight. Its own guide states the entry requirement plainly. SearchPilot works with sites that have at least hundreds of pages on the same template, and at least 30,000 organic sessions a month to the group of pages being tested. Traffic to one-off pages like your homepage doesn't count towards that.
Those two numbers do different jobs. Page count gives you enough units to divide into two groups that behave alike. Sessions give each group enough signal for a real change to rise above the week-to-week movement.
Neither number is a cliff edge, though, and SearchPilot says so itself. SearchPilot reports customers testing on sections of their site that only get a couple of thousand sessions a month. What those customers give up is sensitivity, because the change in traffic has to be much larger before it can reach statistical significance.
You Already Have More Product And Collection Pages Than The Bar Needs
To find out how many stores clear the page-count half of that, I read the public XML sitemaps of 36 UK Shopify storefronts on 25th August 2026. We screened 114 UK consumer store candidates ourselves to get there. Thirty-nine confirmed as Shopify from their storefront markup, and 36 of those served a sitemap we could read. On each one I counted product URLs and collection URLs.

Across 36 UK Shopify storefronts whose sitemaps we read on 25th August 2026, the median store carried 681 product pages and 278 collection pages. Of the 36 stores measured, 25 carried at least 100 collection pages and 32 carried at least 100 product pages. The spread was enormous: product counts ran from 6 to 36,919 and collection counts from 0 to 5,832.
Page count, then, is usually already there. A Shopify catalogue generates those URLs whether or not anyone planned to test on them. What stops most stores is the sessions figure, and Search Console is where you check it. Check that your numbers are trustworthy before you build anything on them.
Fall Back To A Matched Page Group, Then A Phased Rollout, Then Before And After
Below a randomised split sit three more designs, and they give up control in that order.
First of those three is a matched page group. You pick the comparison yourself, pairing your changed pages with pages that have behaved like them in the past. Lauren Busby set out the options in Search Engine Land in August 2026. Match on clicks, impressions or rankings. Crawl frequency, page age and seasonality work too.
Get the matching right and the rest of the design follows. Groups don't need to start at the same level, but they do need to have moved in similar ways.
Second is a phased rollout, which keeps a little less control again. You ship the change to one group of markets or categories and leave comparable sections untouched for a few weeks. Those untreated sections act as a temporary control while they last.
Last is before and after, and it is still worth doing. Busby's rule for it is the honest one. Treat the result as lower-confidence evidence, because a change and a result that share a timeline are not proof that one caused the other.
None of the three needs a paid platform. SEOTesting makes the point that if you are running a single time based test, it is totally possible to report on this using a spreadsheet. You look up one URL in Search Console each day and record what it did.

One Googlebot Forces You To Split The Template
Your conversion tool randomises people. It shows one visitor the green button and the next visitor the blue one, then compares what the two groups did. SearchPilot is blunt about how much of what gets called SEO testing is anecdotal evidence or a badly controlled test.
That distinction rests on a mechanical fact. Search has one crawler, and Statsig puts the constraint plainly: it's not possible to show multiple Googlebots random variants since there is only one Googlebot. That leaves the pages as the only thing available to randomise.
Three consequences follow from splitting pages, and you will meet all of them. A variant controlled by a cookie is invisible to the crawler. A change that alters what a visitor sees but not what the crawler is served measures nothing in search. And your unit of change is a template. Your answer describes the whole template.
Because the visitor-splitting version has its own arithmetic, we cover it separately in Your Traffic Decides Which A/B Tests You Can Run. That article deals with conversion rates on a Shopify store. This one deals with clicks from search.
Cloaking Is The Only Thing Here That Breaks Google's Rules
Splitting a template's pages is not cloaking and it creates no duplicate pages. SearchPilot answers the question directly. When doing SEO A/B testing, there is only one version of the page, and nobody is being shown a different version of the same URL. You are changing a subset of pages that share a template.
That distinction is why cloaking appears in Google's own testing documentation, alongside three other requirements.
Google adds one reassurance on top of those four. Depending on what types of content you're testing, it may not even matter much if Google crawls or indexes some of your content variations while the test is live.
The clean-up is handled too. Once the experiment ends, the version you settle on tends to be reindexed fairly quickly.
Which Page Change Is Worth Your One Slot
There's a public library of controlled SEO tests with the results attached, and reading it is the cheapest way to choose a hypothesis. Someone has already spent six weeks finding out whether the idea works somewhere. Our own testable checklist is another place to pull candidates from.
Titles That Mirror The Search Beat Titles That Add Keywords
SearchPilot's title tag testing points one way across several industries. On a retail customer it changed titles from descriptive formats like Dental Implants Options and Pricing to question based phrasing such as How Much Do Dental Implants Cost. Organic sessions rose by more than five per cent.
Adding keywords went the other way. SearchPilot records a page that originally had the title Women's Dresses becoming Women's Dresses Best Women's Dresses, Dresses for Women, Stylish Women's Dresses. That came back inconclusive at the 95 per cent confidence level.
Working from the other direction, seoClarity's split tester found the same shape. An ecommerce liquor site added Buy [Product Name] Online to its title tags and gained 15 per cent more organic clicks. That modifier matched how its audience was already shopping.
On A Product Listing Page, Content Depth And Page Weight Both Moved Traffic
Two SearchPilot tests on listing pages pull in opposite directions, and both of them won.
Nuha Miah, one of the analysts who writes up SearchPilot's controlled tests, ran an April 2026 test adding FAQ content below the existing footer copy on a set of product listing pages. SearchPilot reports that the test delivered a positive result at the 95% credible interval, with an estimated +9.7% increase in organic clicks. Those FAQs answered category-level questions, which gave the page more long-tail language.
Veneta Mihaylova's test the following month went the other way, reducing the number of products displayed per category page from 48 down to 36. Layout, navigation and filters stayed as they were. SearchPilot reports that one as positive at the 85 per cent credible interval.
Both of those results are reconcilable. A listing page gives search engines little text to read and a lot to load, so adding query-aligned words and removing weight are both acting on real limits of that template. Which one your pages need depends on where yours are weakest. If that is weight, what really moves it on a Shopify theme is rarely the score in the tool.
You Learn More From The Title Tests That Lost
A losing result takes an idea off your list for good, where a winning one still has to be proved again on your own pages. That makes the failures the more efficient read, and SearchPilot publishes plenty of them.
Clearest of those failures is the video label test. A media customer added labels like (video) and (with video) to article titles, so that Football Highlights became Football Highlights (with Video). Both versions performed worse. Organic traffic dropped for the variant pages.
Prices in titles split on implementation. Static prices such as Apartments in Paris from EUR59 led to a negative seven per cent result on a travel rental customer's location pages, because an outdated price misleads the person reading it. Dynamic prices pulled from a live feed gained ten per cent for the same customer. Identical idea in both tests, and the outcome turned on whether the number was current.
SearchPilot's own advice on all of it is to test in your own context. The right answer varies with your speed baseline and your category structure. Reading its six title tests together, one line does the summarising. Adding more keywords doesn't always help and can make titles worse.
Four Decisions To Settle Before The Test Goes Live
Whether an SEO AB testing run produces a readable number comes down to four decisions. Settle them before anything ships, because none of them can be repaired once the test is running.
A Control Group Absorbs Everything Except Your Own Second Change
Freeze everything outside the change, and keep the template frozen for the whole run. Rewriting the page content or updating the wider template mid-test makes the effect of your change impossible to separate from the rest of the work.
A control group absorbs whatever happens to everybody. SearchPilot's methodology note states what that buys you. By detecting changes in performance of the variant pages compared to the control, you know the effect was not caused by seasonality, sitewide changes or a Google algorithm update. What it can't absorb is a second change you made yourself to the same pages.
One case is worth naming, because a control group can't help you if you're watching the wrong pages. When the change is an internal linking module, the pages carrying the new links and the pages that benefit from them aren't the same pages, so the destinations are what you track.
Your Two Groups Have To Have Moved Together Before You Start
A random 50/50 split doesn't by itself make two groups comparable. If one side happens to hold your stronger categories, the split has done nothing for you.
Match on trajectory. Today's traffic is the wrong thing to pair on. Two groups sitting at similar traffic are a poor pairing if one has been growing for months while the other has been sliding. Two groups at different levels are fine if they have historically moved in the same shape.
Getting that pairing right is also what protects the test from Google itself. Google's own Search Status Dashboard, the public record of confirmed ranking updates, records the May 2026 core update starting on 21st May and running 11 days and 21 hours. March's core update ran 12 days and 4 hours.
A rollout that long sits inside a normal test window. A control group moving alongside your variant group absorbs it, and a before and after comparison has nothing to absorb it with. That is the whole reason SEO AB testing bothers with a control at all.
A Tag Manager Change Never Reaches Googlebot
SearchPilot's answer on Google Tag Manager is blunt. While you can make changes to a page using Google Tag Manager, Google may not see the changes, because they're being made with JavaScript. The crawler renders the page without your injection and gets served the control version of every page in the test.
Cookies fail the same way, and the crawler is served the wrong page again. Google's documentation is explicit: if you're using cookies to control the test, keep in mind that Googlebot generally doesn't support cookies. A cookie-controlled variant only ever shows the crawler the version a browser refusing cookies would get. If a tag manager is your only route to the change, the test isn't ready to run.
Match The Window To What You Changed
SEOTesting sets the windows by what you changed. Page title and meta description changes get two weeks, content improvements four, and new links six. Where there is no deadline it suggests six weeks every time, since there is no harm in capturing data over a longer period.
Those windows assume you can wait. SearchPilot puts the general case this way. Generally, positive or negative SEO experiments take 2-4 weeks to reach statistical significance, but a trend can often start to appear within less than a week. A trend isn't a result, so set the end date before you start.
Reading A Result That Says Nothing
Plenty of tests come back without a clear winner, and two of the results above did exactly that. Reading a null takes different rules from reading a win, and getting them wrong is how a change gets shipped on noise alone.
Define Failure And Inconclusive As Well As Success
Write down what success, failure and an inconclusive result each look like before you have any data. Lauren Busby makes the same point in Search Engine Land about setting those three states in advance.
That order matters more than it sounds, because once a chart is on screen any threshold you pick is chosen partly by what the chart already shows and partly by what you were hoping for. Writing the decision rule takes a minute. It is the only thing that makes the answer binding on you.
The Alt Text Test Whose Credible Interval Crossed Zero
Noe Servin, who has written up several of SearchPilot's title and listing experiments, ran the alt text test from June 2026 that is worth keeping in mind. A classifieds customer rewrote the alt text on its category page listing images, combining each listing's name with the page category heading.
SearchPilot reported the numbers in full. The results of this test were inconclusive. The best estimate indicated a 7.3% increase in organic sessions; however, the 95% credible interval ranged from -4.2% to +19.4%, which crosses zero.
Crossing zero is what settles it, because one of the plausible values is no effect at all. SearchPilot reports the test as inconclusive. Reporting the 7.3 as a win would have been wrong.
Convention here is a 95 per cent confidence level, and SEOTesting documents the p-value below 0.05 that sits behind it. Report the interval and not the headline number.
A Stable Null Will Not Narrow With More Data
One more line in that write-up decides what you do next. SearchPilot notes the outcome was highly stable, suggesting that it is unlikely to become significant with more data.
An underpowered test still has an interval that narrows as data comes in. A stable one has stopped narrowing, so waiting won't change the answer. Recognise that and the slot goes to your next hypothesis. Another month of the same test buys you nothing.
There is a second half to this, and seoClarity's AI-generated title testing supplies it. Its results split on the quality of the titles already in place: poor baselines improved, and strong ones saw no lift or came back worse. Where your pages start decides what a change can do for them. Treat a published result as a hypothesis for your site, and never as a setting to copy.
Where An SEO Consultant Fits Into Your Testing Programme
A store gets a handful of these a year. Spending one on a hypothesis nobody sharpened costs you the six weeks and the answer, and so does spending it on two page groups that were never comparable.
Choosing the hypothesis and pairing the groups is what an outside consultant is for, rather than the execution. We work with teams who have the resource to ship a test and want the design settled first. Which template can carry it, which change is worth the slot, and how the two groups get paired. I run these engagements myself.
If your sessions sit below the threshold, an SEO AB testing programme is not your first job and we will tell you so. Getting organic traffic up to where a test can read at all is the work that comes first, and it's the same eight disciplines we run on any retainer.
Book a call through our SEO consulting service and bring your Search Console. Half an hour is enough to say which of the four designs your store can support, and which change is worth spending an SEO AB testing slot on.

