Card on method changes that break year-over-year trend report comparisons. Before you act on annual trend reports comparison
Image: Algorithm Trend Updates

Maintenance

Part of How to read an annual trend report in about twenty minutes

Before you act on annual trend reports comparison

Annual trend reports comparison: the six method changes that quietly break a year-over-year series, and what to publish when the series is broken.

The most natural move with this year's report is to set it beside last year's. That is almost always wrong. An annual series looks like a time series, but usually is not one.

The measured thing, the people measured, and the question wording all shift between editions. This page asks whether two editions are comparable, and what to do when they are not.

What to take away

  • Check whether the two editions measured the same thing before you compare any number in them. Most of the time they did not.
  • A change in the sample is as disruptive as a change in the definition and much harder to see.
  • Where a series is broken, say so and report the two points separately. A joined line across a break is worse than no line.

Six things that break a series

How to spot it

The definition of the measured thing
Compare the wording of the question or metric
The population sampled
Read the method note in both editions
The recruitment method
Often only in an appendix, if at all
The response rate
Rarely published; look for it
The categories offered
Compare the answer options
The analysis cut
Compare the table headings

What it does to a comparison

The definition of the measured thing
The two numbers are about different things
The population sampled
The difference is the sample, not the world
The recruitment method
Changes who answers, in a direction you cannot sign
The response rate
A large drop can move a figure on its own
The categories offered
Adding a category takes share from the others
The analysis cut
A different denominator, presented as the same number

Any one of these is enough to make a year-over-year difference uninterpretable. In practice several move at once, because a publisher improving their method changes three of these in the same edition.

What Breaks a YoY Series

What changed

Definition
Compare question wording
Population
Read method note
Recruitment
Often only in appendix
Response rate
Rarely published
Categories
Compare answer options
Analysis cut
Compare table headings

How to spot it

Definition
Different things measured
Population
Difference is the sample
Recruitment
Changes who answers
Response rate
Composition shift moves figure
Categories
New category takes share
Analysis cut
Different denominator, same label

Effect on comparison

Definition
Population
Recruitment
Response rate
Categories
Analysis cut

Improving the method is good, but it breaks the series anyway.

Publishers differ in which rows they touch and how plainly they say so, a spread mapped in how the sources diverge.

The response rate row is the quiet one. If the same instrument goes to the same population and fewer people answer, the respondent mix has changed. Any movement in the results is partly that.

Official statistical practice treats such a discontinuity as a break to document, so the rules a defensible survey has to disclose put method changes among required disclosures.

Why an improvement breaks the series

This is the part people find counterintuitive. A publisher who fixes a badly worded question, widens a sample, or corrects a weighting has made this year's number better and made the comparison worse. Both are true at once.

The underlying issue is that a series requires the measurement to be held constant even when it is imperfect, and holding an imperfect measurement constant is a real cost that publishers reasonably decline to pay. What a careful publisher does is run both versions for one edition and report the overlap, which is expensive and rare.

When you meet an unexplained jump between editions, a method change is a more likely explanation than a change in the world, and it should be your first hypothesis rather than your last. Ranking the dull explanation ahead of the dramatic one is the same habit as checking causes in the right order after a service goes down.

Same name, different thing

The subtlest break is the one where nothing in the method changed and the concept moved anyway. A term in wide use drifts in meaning, respondents answer according to the current sense, and the question wording never changed to reveal it.

That is a construct validity failure: the measurement no longer captures its intended concept. It cannot be detected from the numbers, only from knowing the subject well enough to notice that the words moved.

The defense is the same one that protects any measurement over time: writing down what you meant when you defined it. That is the discipline of operationalization done properly.

Panels, and what following the same people buys

Some reports track the same respondents across editions rather than drawing a new sample each time. That is panel data, and it is genuinely stronger for measuring change, because the composition is held constant by construction.

It brings its own problems, and they are worth knowing. Panels attrit, and the people who drop out are not a random subset. Panels also learn: people who have answered the same question three times answer it differently the fourth. Neither ruins the design, and both mean a panel result is not automatically the more trustworthy one.

The related design where a defined group is followed forward, a cohort study, is worth recognizing for the same reason. When a report describes its design in these terms, it is telling you something real about what its year-over-year comparisons can support. Copying that description into your own file is part of the maintenance a report deserves after publication.

What to do with a broken series

Report the two points separately, with the break named between them. That is the honest presentation and it is less satisfying than a line, which is why it is rare.

If you need a direction rather than a magnitude, look for a claim that survives the break: something both editions asked in the same words, or a ranking rather than a level. Rankings hold up better than levels under a method change, though they are not immune.

And record what you found for next year, because you will do this again. That note is part of the same upkeep, and it is the difference between doing the work once and doing it every year.

Common questions

How do I compare two reports from different publishers?

Usually you cannot, and the reasons are the six rows above with nothing held constant. What you can compare is where they agree on direction, which is weak evidence and better than a false precision.

What if the method note is missing?

Then the comparison cannot be made, and saying so is the correct output. The alternative is producing a number whose meaning nobody can state, which is the standard failure described in why most figures cannot be had.

Are longer series more reliable?

Not automatically. A long series has had more opportunities to break, and old breaks tend to be forgotten. Reading the method notes of the earliest editions is often more revealing than reading the latest.

Does a big change ever just mean a big change?

Sometimes, and it should be the hypothesis you reach after eliminating the six, not before. The cheap explanations are also the likely ones.

More in Maintenance

Latest from Records Desk