You Can't Read a Position Without a Baseline
“We’re at position seven” is not a finding. Seven could be the best this keyword has ever done or a two-week low. Without knowing where it normally sits, the number supports no conclusion, no alert threshold, and no report.
The baseline isn’t a benchmark someone else publishes. It’s this keyword’s own history, on your settings, and there’s no shortcut to having it.
What a baseline actually is
For any tracked keyword, three things:
A central tendency. Where it usually sits. Prefer the median over the mean, because a single week spent at 60 during an outage drags a mean permanently and leaves a median untouched.
A spread. How far it wanders in normal operation. This is the part people skip and it’s the part that makes the baseline usable, because “outside its normal spread” is the only defensible definition of an event.
A duration. How long you’ve been observing. Two weeks of daily checks is enough to spot gross problems and not enough to characterise spread. A quarter is comfortable.
Once you have all three, most of the hard questions in rank tracking get easy: a drop is a drop when it’s outside the spread, and everything inside is weather.
Why spread varies so much per keyword
Some queries are settled: the intent is unambiguous, the result set has been the same for months, and positions barely move. Others are permanently unsettled, because the query means several things and the ordering keeps being reconsidered.
A tracked set almost always contains both, and a global rule (“investigate any five-place move”) over-fires on the unsettled ones and misses real events on the settled ones. A two-place move on a keyword that’s held position four for four months is a bigger deal than a fifteen-place swing on a keyword that swings fifteen places every week.
This per-keyword spread is what your alert thresholds should be built from — rank alerts that don’t cry wolf.
Things that reset a baseline
A baseline is only valid while the measurement is unchanged. These invalidate it, and each needs an annotation on the chart:
- Changing tracker. Different tools report different numbers, so the series before and after aren’t comparable — why two rank trackers disagree.
- Changing location, device or language settings. Same tool, different question.
- Changing check frequency. Daily and weekly series have different spread by construction; weekly data is smoother because it samples less.
- Changing the target URL. A new page competing for the term is a new measurement.
- A layout change on the result page. Position may be comparable, but its value isn’t.
- Moving the sampling weekday. A small offset, consistently applied, and it looks like a step in the series — weekday and seasonal effects in position data.
The habit that makes this survivable is annotating the series at the moment you change something, not reconstructing it later from memory.
Baselines for keywords you just added
New keywords have no baseline, which is fine as long as you don’t pretend otherwise. Two ways to cope:
Borrow from the group. If a keyword sits in a group of similar terms with established spreads, the group’s typical spread is a reasonable provisional assumption. Mark it as provisional.
Watch, don’t alert. Keep new keywords in a no-alert state for their first few weeks. They’ll produce apparent dramatic movement as they settle, and none of it means anything.
The temptation is to react to a new keyword’s first big move. Don’t. You have one observation and a change, which is two numbers and no context.
Baselines at aggregate level
The same logic applies to whatever aggregate you report — band counts, a visibility index, share of voice. Each has a normal range, and month-on-month movement inside that range is not a result.
Aggregates are noisier than people expect at small keyword counts, because a handful of keywords crossing a band boundary moves the count visibly. If your tracked set is thirty keywords, a band count of eleven moving to nine is well within what chance produces. Small-sample behaviour is the reason to report direction over several months rather than deltas month to month, and the same reason to keep the keyword set frozen — visibility scores and share of voice.
What to actually do
- Record the median and the observed range for every tracked keyword, and refresh them quarterly.
- Define a noise floor per keyword or per group, not globally, and don’t investigate inside it.
- Annotate every settings change on the series at the time you make it.
- Give new keywords a few weeks of no-alert observation before they count.
- Never quote a position without its range in a report. “Position 7, normally 6 to 9” is a sentence nobody can misread.