Ranking Changes Are Not an Experiment

A before-and-after on one keyword has no control group, so it cannot support a causal claim. You changed a title tag, the position improved two places, and there is no version of that week in which you didn’t change it. The comparison you want — this page with the change against the same page without it, over the same period, on the same result page — is not available to anyone.

This is not pedantry. Most of the folklore in SEO is a single before-and-after that somebody generalised.

What is missing is the counterfactual

For a keyword, there is one page, one query, and one timeline. Everything else that could explain the movement stayed in the picture:

  • Background volatility. The keyword has a normal wandering range and your change landed inside a week of it — you can’t read a position without a baseline.
  • Other people’s changes. Competitors publish, update, and consolidate on their own schedule, and any of it reorders the list.
  • Seasonal and weekly shape. Demand and intent shift over the calendar, so a same-day comparison across a week boundary compares two different weeks — weekday and seasonal effects in position data.
  • Your other deploys. Almost nobody ships one change in isolation, and template edits touch pages you weren’t thinking about.
  • Layout changes on the result page. Gaining or losing a feature block changes the number your tracker reports while the ordering stands still — position zero is a reporting convention.
  • Processing lag. The effect of a change appears whenever it is recrawled and reprocessed, not when it shipped, so the timing evidence is blurred.

Any of these is a sufficient alternative explanation for a two-place move. Several are usually present at once.

What you can randomise is pages

The unit you can assign at random is not the keyword — it is the page. If you have a large set of comparable pages generated from one template, you can split them into buckets, apply the change to one bucket, leave the other alone, and compare how the two groups’ outcomes trend over the same period.

That design buys you the thing a single before-and-after lacks: a control that lived through the same weeks, the same volatility, and the same competitor activity.

It needs conditions that most sites cannot meet:

Enough comparable pages. Two buckets of a handful of pages will produce a difference by chance. This is a large-catalogue technique — product, location, or category templates — not something you run on a ten-page site.

Random assignment, not convenient assignment. Splitting by category, by age, or by “the ones we got to first” builds the difference into the buckets before you start.

One change. If the treated bucket also got new internal links, you have tested two things and learned about neither.

A pre-committed read date. Decide when you will look and what result would count, before you look. Otherwise you will stop at the first favourable week.

Position is a poor outcome measure even then

Position is a bounded ordinal, so differences in it do not add up the way the bucket comparison wants them to — moving 3 → 2 and 43 → 42 are not the same event and averaging them pretends they are (position is an ordinal, not a quantity). A bucket mean position is therefore hard to interpret and easy to swing with a few long-tail terms.

Impressions and clicks per bucket are better outcome measures. They are counts, they aggregate honestly, and they come from your own property rather than from a sampled retrieval. Position is still worth recording as a supporting signal — it explains how an outcome changed — but it should not be the thing you declare significance on.

And even a well-run bucket test is not a randomised trial. The buckets share a host, a crawl budget, and an internal link graph, so they are not independent of each other. Treat the result as evidence about your site, not a discovered ranking factor.

When you can’t run a split at all

Most of the time you can’t, and the honest fallback is a documented observation rather than a test:

  1. Freeze a control set of untouched keywords or pages with similar volatility, and read the treated set relative to it.
  2. Write down the prediction and the read date before shipping.
  3. Annotate the change on the series at the moment it ships.
  4. Report the result as consistent with, or inconsistent with, the prediction — not as proof.

That is weaker, and it is still far better than the alternative, which is attributing whatever the chart did next to whatever you did last. The same discipline in reverse is how a drop should be handled — how to diagnose a ranking drop.

What to actually do

  1. Stop calling single-keyword before-and-afters tests. Call them observations, and keep them.
  2. Only claim causality from a design with a control that experienced the same period.
  3. Randomise pages, never keywords, and only where the catalogue is big enough.
  4. Measure the outcome in clicks and impressions, with position as the explanation rather than the verdict.
  5. Pre-commit the read date and the losing condition, so a null result is a result and not a reason to keep looking.