Maintenance

Part of Ranking signals: a clear guide with practical examples

Ranking signals comparison: what to know and why

Ranking signals comparison of what each kind of evidence can carry, why a correlation study cannot support advice, and the comparison worth running yourself.

Most published claims about ranking factors come from one kind of study: take a few thousand items that currently rank, measure some attributes, and report which attributes track position. The results get written up as advice. The method cannot support advice, and the gap between what it measures and what it is used to claim is the single biggest source of confident nonsense in this field. This page compares the kinds of evidence you can actually get, and what each one will and will not carry.

What to take away

  • A correlation study describes what already ranks. It does not show what would happen if you changed something.
  • The sample in these studies is selected on the outcome, which is the one sampling choice that makes causal reading impossible.
  • The only evidence that supports a decision about your own site is a change you made on your own site, measured against something that did not change.

What a ranking factor study actually measures

The procedure is roughly: collect items in ranked positions, measure attributes the tool can see, compute a correlation between each attribute and position. Everything that follows depends on that sample.

The sample contains items that ranked. It contains nothing about the items that did not, so you are looking at survivors and reasoning about the population. Attributes also arrive in clusters, so the study cannot separate them: an operation that publishes long pages also has an editor, a design budget, an established audience and a domain people already trust. Correlate any one of those with position and you will find something, because you are really measuring the cluster. That is a textbook confounder, and no amount of sample size fixes it.

Direction is the third problem. Several of the attributes people measure are consequences of ranking rather than causes of it. Something visible collects links, mentions and engagement because it is visible. The correlation is real and the arrow points the other way, which is the ordinary meaning of the warning that correlation does not imply causation.

Finally, the number is an average over wildly different queries and audiences. A relationship that is positive in one segment and negative in another can average to nothing, and a relationship that holds only in one segment can average to something. A single figure across a whole corpus hides both.

Comparing what each kind of evidence can carry

Evidence What it can honestly support What it cannot support Main weakness
One person's before and after A hypothesis worth testing Anything general No control, and only the wins get published
Correlation study across ranked items A description of what currently ranks Any claim about the effect of a change Sample is selected on the outcome
Watching a competitor gain A prompt to look at what else changed that week Attribution to any one thing they did You see their output, never their inputs
A platform's own published guidance What the platform says it wants and will penalise The weighting, or how the words map to code Written for a wide audience, deliberately general
A natural experiment on your property A reasonable inference, if the timing is clean A precise effect size The world changes at the same time you do
A controlled test on your own property A specific effect, on your site, for that period Generalisation to other sites Slow, and needs enough traffic to see a signal

The order matters more than the rows. Evidence near the bottom costs more and answers a narrower question, which is exactly why it is worth more. A study covering ten thousand domains tells you about the average of ten thousand situations, none of which is yours.

How a correlation study earns a place

It has one honest use: generating candidates. If an attribute tracks position across a large sample, that is a reason to put it on a list of things to test, and no more than that. Treat the output as a queue of hypotheses and the study becomes useful. Treat it as a set of instructions and you will spend a quarter on the strongest confounder in the data.

There is a second honest use, which is describing a moment. A well-run study is a snapshot of what the visible results look like now. That is genuinely informative about the current shape of a results page, and it does not require any causal claim at all. Reading one that way is closer to reading a census than reading an experiment. The record of how these claims have changed over time is a useful check on how quickly the snapshots go stale.

The comparison you should actually run

Against the studies, set the one thing you control: a change you make deliberately, on a slice of your own property, with a comparable slice left alone. It answers a narrow question and it answers it about you. The structural background sits in the overview of ranking signals, and the stage map of a ranking pipeline tells you which stage a change could plausibly touch, which is worth knowing before you spend a month on one that could not have mattered.

Two habits make the comparison honest. Write down what you expect before you look, so you cannot fit the story to the result afterwards. And decide in advance what result would make you abandon the idea. A test you would explain away either way is not a test.

Common questions

Are ranking factor studies dishonest?

Usually not. Most are careful about their method in the methodology section and get misread in the summary. The failure is in the reading far more often than in the arithmetic.

Why do two studies disagree about the same attribute?

Different samples, different measurement tools, different corrections, different periods. Any of those can flip a weak correlation. Persistent disagreement across studies is a signal that the underlying effect is small or conditional, not that one team is wrong.

Can I compare my site against a competitor as a control?

Not really. A control has to be similar in every respect except the change, and another company differs in every respect at once. The nearest usable version is comparing two sections of your own site. The failure modes worth knowing before you try are worth reading first.

What about studies run by the platform itself?

They have access nobody else does and a strong interest in the conclusion. Read them for the mechanism they describe rather than the effect size they report, and note what the comparison group was. Work on how recommendation research is designed shows how much the framing of a question drives the answer.

More in Maintenance

Reviews

Ranking signals: a clear guide with practical examples

Ranking signals sorted by the five producers of evidence, why a signal is not a factor, four tests for any claimed input, and examples read the right way.

Features

Ranking signals platforms explained with examples

Ranking signals platforms as a stage map: why pipelines have stages at all, how to read a symptom back to one, and what the map deliberately leaves out.

Rules

Ranking signals research: practical details and examples

Ranking signals research you can run: setting the question narrowly, building a split that cannot flatter you, and reading the result without fitting a story.

Reviews

Ranking signals timeline: what to know and why

Ranking signals timeline of a claim, from one account's odd week to folklore: seven stages, why aggregation is the dangerous one, and how to date a claim.