What a before and after actually measures
Eight pages were rewritten on April 15. Over a 90-day window that began two weeks later, they drew 13.7% fewer clicks a day than over the 90 days before the change. The rewrite worked, and by a lot.
The pages are not real. They belong to a sample property built to test the method behind our free Did it work? tool, and the right answer is known because it was planted: the rewrite multiplied those eight pages' traffic by 1.9. Over the same months the site as a whole lost 52% of its daily clicks, the kind of slide a seasonal business sees every year. The rewritten pages fell far less than everything around them. Measured against untouched pages that had matching traffic before April, they ended up 88% above where they would otherwise have been.
Both figures, minus 13.7% and plus 88%, are simple arithmetic on the same two Search Console files. Only the second is the effect of the change. Plenty of SEO reporting stops at the first kind of number, and a client who is shown a falling line after paying for a rewrite has every reason to leave. The rest of this piece is about producing the second number: what it takes, why each step is there, and when the data cannot support any number at all.
A before and after fails because a search property is never still. Four things in particular move traffic between any two windows without anyone touching a page.
The first is seasonal demand. Google's own guide to debugging traffic drops lists seasonality and changing interests among the main causes, alongside algorithm updates and technical faults. A tax preparer measuring a May change against an April baseline will find a disaster. A ski shop measuring an October change in December will find a triumph.
The second is Google itself. It says that several times a year it makes significant, broad changes to its ranking systems, and that improvements made after one can take several months to show. An update that lands inside your window moves the changed pages and the unchanged ones alike. Google's advice for diagnosing an update is itself a before and after: wait a full week after it finishes rolling out, then compare with a week before it started. That is the right tool for the question "what moved?". It cannot tell you why.
The third is anything sitewide. A domain gaining links, a new competitor, a migration, a change in how the results page is laid out: all of them lift or sink every page at once, and a before and after on a subset of pages hands the whole movement to whatever you changed.
The fourth is statistical, and the easiest to miss: regression to the mean. Pages are usually chosen for work because they just had a bad stretch, and anything picked at a low point tends to drift back toward its usual level on its own. Say a page averaged 40 clicks a day for a year, slipped to 25 in the month before its rewrite, and settled back at 38 afterwards. A before and after reports a 52% gain for a page that simply returned to normal.
A control group builds the counterfactual
The question a client is really asking is counterfactual: what would these pages have done if we had left them alone? Nobody can observe that directly. The next best thing is other pages on the same site, over the same two windows, that looked like the changed pages beforehand and were not touched. Whatever happened to them, the season, the update, the tide, happened to the changed pages too.
The details decide whether a comparison holds up or quietly flatters the result, so here is the matching exactly as the Did it work? tool does it.
- The control pool is every page that was not changed, appears in the before export with at least one click, and is not missing from a truncated export. With fewer than 10 such pages, the tool refuses to form a control set at all.
- Changed pages are matched largest first, because high-traffic pages have the fewest similar neighbors and would otherwise lose them to smaller pages.
- Each changed page gets up to five controls: the untouched pages whose before-window clicks are closest to its own. Distance is measured on a log scale, so the gap between 10 and 20 clicks counts about the same as the gap between 1,000 and 2,000.
- A caliper caps how far a match can stray: roughly a factor of two in before-window clicks. A page with fewer than two controls inside that band is reported as unmatched rather than paired with something unlike it.
- Matching is without replacement. Each control backs one changed page and no other.
That last rule costs precision on paper and buys honesty in practice. If one quiet page is allowed to serve as the control for five changed pages, any quirk in that page, a stray spike or a bad week, is counted five times, and the result looks more certain than the data warrants.
The controls then produce an expectation. If a changed page's five controls between them went from 10 clicks a day to 8, the page was expected to fall by the same ratio, to 80% of its own before rate. The measured effect is the page's actual after rate minus that expectation, in clicks per day. The effects of all matched pages are added up, and everything is kept as a daily rate so that windows of slightly different lengths still compare.
In the sample, the eight rewritten pages were matched to 40 controls drawn from a pool of 132. The rewritten pages ran at 78.6 clicks a day after the change. Their controls implied 41.8 without it. The difference, 36.8 clicks a day, is the 88%.
Two cautions. A control is only a control if the change could not reach it. A template edit that also touched other pages, or new internal links pointing into neighbors, contaminates those neighbors, and a contaminated control pulls the measured effect toward zero. And a page that went from no clicks to some has no baseline, so it is reported as a count of clicks, never as a percentage of zero.
Matching also has a limit worth stating plainly. It matches on the level of traffic in the before window, not on its trend, so it does not remove regression to the mean when pages were picked because they had just dipped. A long before window, 90 days rather than 14, dilutes a recent dip. Choosing pages for reasons other than a fresh drop, or saying in the report that they were chosen that way, is the rest of the fix.
A placebo test tells you how noisy the site is
Even with no change at all, pages rise and fall week to week. A measured effect of plus four clicks a day might be your rewrite or might be what this site does every quarter. The control group cannot answer that on its own. A placebo test can.
The tool picks a random set of untouched pages with the same traffic profile as the changed set, each one inside the same factor-of-two band as the changed page it stands in for. It pretends those pages were changed, gives them their own matched controls from what remains, and runs the identical calculation. Then it does it again, 1,000 times, with a fixed random seed so that the same files always produce the same answer.
Those 1,000 results are a portrait of ordinary churn on this property. The middle 95% of them is the range that doing nothing produces. The real result is only called a gain if it sits above that range, and a loss if it sits below. The same test that confirms a win catches a change that did harm.
In the sample, the placebo range ran from minus 2.06 to plus 1.47 clicks a day. The real measurement was plus 36.8, and none of the 1,000 placebo sets came close. Synthetic sites are tidier than real ones, and on a real property the band is usually much wider, which is exactly why it has to be measured rather than assumed.
The band also states the resolution of the test. Half its width, divided by the expected clicks, is roughly the smallest change this comparison could have detected. The sample could resolve a change of about 4%. A small site might only resolve 60%, and a client deserves to know that before being told a change "did nothing".
Four verdicts, and one of them is not yet
Before computing anything, the tool checks for hard stops, any one of which ends the analysis with an instruction to re-export. The two windows overlap. The before window ends on or after the change date. The after window starts before it. The two exports were taken with different filters, or from different properties. None of these has a soft mode, because each one means the comparison would measure the wrong thing.
Past the stops, there are four possible answers.
- Measured gain, or measured loss: the effect sits outside the placebo range. Report the clicks per day, the window and the number of controls.
- No effect this comparison can detect: the effect sits inside the range. Say so, and give the resolution, so "no effect" is read as "nothing larger than about X%".
- Not enough traffic yet to tell: either window is shorter than 14 days, the changed pages drew fewer than 30 clicks before the change, or the pool is too small to run a placebo test.
- No control set could be formed: fewer than 10 untouched pages carry traffic, or no changed page has comparable peers. The raw before and after can be described, and it cannot be attributed.
"Not yet" is a legitimate deliverable, and the hardest one to say to a client. An agency that says it when it is true earns the right to be believed when it says a change worked.
Small pages cannot prove small effects
Clicks are counts, and counts are noisy in a predictable way: the random wobble in a count is roughly its square root. A page with 100 clicks in a window could easily have had 90 or 110 for no reason at all, about 10% either way. At 400 clicks the wobble shrinks to about 5%. That is why big changes on small pages, and small changes on big pages, can both be measured, while small changes on small pages usually cannot.
The rule of thumb the tool uses: to detect a relative change d, you need roughly 2 × (2.8 / ln(1 + d))² clicks in each window. The 2.8 bundles two standard choices: a 5% chance of calling an effect that is not there, and an 80% chance of catching one that is.
| Change you want to detect | Clicks needed in each window |
|---|---|
| 10% | about 1,730 |
| 20% | about 472 |
| 30% | about 228 |
| 50% | about 96 |
Say the changed pages drew 19 clicks in a 90-day before window and 33 in the 90 days after. That is a 74% jump, and it is inside the noise: at these counts, a 95% confidence interval around the change runs past zero. At 33 clicks a quarter, reaching 472 would take more than three years. The tool reports that kind of projection as an estimate, with its assumption stated (traffic holds at the current rate). Often the useful conclusion is not "wait" but "measure a larger group of pages together".
Search Console has traps that look like results
The method needs two Performance exports from the same property with the same filters, one ending before the change went live and one starting on or after it, plus the list of pages you changed. Getting those files right matters as much as the statistics, because several documented Search Console behaviors produce movements that look like effects.
Start with the dates. A Search Console zip includes a Filters.csv that records the date range as you picked it, often a relative label. This site's own export for the six months to September 23 says only "Last 6 months". The absolute dates are in the daily chart file, and that is where the tool reads them. A Query or Country filter makes the whole export a slice of the site, and every figure from it describes that slice.
Next, the row cap. Google's post on Performance data filtering and limits says the most you can export through the interface is 1,000 rows. When a page appears in one export and not in a capped other, its traffic there is unknown, not zero. The tool sets such pages aside and counts them, rather than scoring a collapse that never happened.
Query tables have a bigger gap. The same Google post says anonymized queries, those not issued by more than a few dozen users over a two to three month period, are omitted from tables but included in chart totals. The effect can be large. In this site's six-month export, the 1,000 named queries accounted for 92 of the 312 clicks in the chart. The Pages table summed to 313. That is why the tool measures pages: a query-level proof here would describe under a third of the traffic.
Impressions have the opposite problem. Google's Performance report help explains that the chart aggregates by property, so two of your results on one results page count as a single impression, while the Pages table counts each page. In the same export, page-level impressions came to 470,248 against 429,408 in the chart. Never add up page impressions and call it the site total.
Average position is an average of a moving mix. Google's definition is the topmost position your site held in a result, averaged across impressions. When a rewritten page starts appearing for a batch of new long-tail searches at position 40, its average position gets worse even if every ranking it already had improved. Position is useful for diagnosing what happened. Clicks decide whether it worked.
Finally, timing. Google says collected data is usually available in two to three days, and the newest points are preliminary, so an after window that ends yesterday is short a few days. Recrawling a changed page can take anywhere from a few days to a few weeks, which is why the tool warns when the after window starts less than 14 days after the change. At the other end, Search Console keeps 16 months of data, so a baseline from well over a year ago may simply be gone by the time someone asks for it.
Write the result so anyone can check it
A result belongs in a single sentence that carries its own working: the rate, the window, the comparison and the test. From the sample, it reads like this.
The eight rewritten pages ran at 78.6 clicks a day from April 29 to July 27, against 41.8 implied by 40 untouched pages matched on their January to April traffic: plus 36.8 clicks a day, or 88%. None of 1,000 placebo sets of untouched pages produced a difference that large.
Every figure in that sentence can be checked against the files. What it leaves out is deliberate. It does not annualize, because multiplying a 90-day rate by four assumes the effect and the season both hold for a year, and neither is measured. It does not quote average position. And it does not claim a revenue figure, because nothing in a Search Console export measures revenue.
The free Did it work? tool runs this whole method on two Search Console exports, in your browser, with the sample above preloaded so you can see each step. AIO Copilot runs the same measurement on every change for every client and puts the result, with its range, in the report the client reads.
The part no tool can do for you happens before the change ships. Write down the date it goes live and the exact pages it touches, and export the before window that day. Six months later, when someone asks whether the work paid off, those three things are the difference between an answer and a chart.
Frequently Asked Questions
Why is a before and after comparison not enough to prove SEO worked?
Because it credits the change with everything else that happened in the same window: seasonal demand, Google core updates, sitewide trends and pages drifting back from a bad month. A control group of similar pages you did not touch shows what would have happened anyway, and the effect is the gap between that and what the changed pages did.
What is a control group in SEO?
A set of pages on the same site, measured over the same two date windows, that had similar traffic to the changed pages before the change and were not touched by it. Their movement between the windows is the expectation. Whatever the changed pages did beyond that expectation is the measured effect.
How many clicks do I need to measure an SEO change?
Click counts are noisy in a predictable way, so the answer depends on the size of the effect. To detect a 20% change at 95% confidence with 80% power you need roughly 472 clicks in each window; a 50% change needs about 96. With fewer, the honest verdict is that there is not enough traffic yet to tell.
Should I use average position to judge an SEO change?
No. Google calculates average position from the topmost position your site held in each result, averaged across impressions, so it moves whenever the mix of queries changes, even if no ranking did. Use it to diagnose what happened and let clicks decide whether the change worked.
How long should I wait after an SEO change before measuring it?
Google says recrawling can take from a few days to a few weeks, and Search Console data usually arrives two to three days after it is collected. Start the after window at least two weeks after the change went live and give each window at least 14 days; longer, equal windows are better.
Can I measure an SEO change with query data instead of page data?
Only for a minority of the traffic. Search Console leaves anonymized queries, those not issued by more than a few dozen users over two to three months, out of its query tables, and the interface exports at most 1,000 rows. Page tables carry far more of the clicks, which is why page-level measurement is the default.
Never miss an update
Get the latest AI and SEO strategies delivered to your inbox.