Hubble social media lead Erin Kisliuk (right) takes a photo with Dr. Nancy Grace Roman (center), NASA's first Chief of Astronomy and "the Mother of Hubble," at NASA's Goddard Space. Feature updates research explained with examples
Photo by NASA Hubble Space Telescope on Wikimedia Commons, Public domain

Guides

Part of Feature updates: methods, tools and useful context

Feature updates research explained with examples

Feature updates research done properly: why early data is the worst data, the sequence to follow, the shapes worth recognizing, and when to act on one.

An update lands, your numbers move, and the pressure to explain it arrives the same afternoon. Almost everything that goes wrong from here comes from investigating in the wrong order: forming the conclusion in week one, then collecting the evidence that fits it. This page sets out a sequence that survives contact with a real week, and the shapes in your own data that tell you which question to ask next.

What to take away

  • The first two weeks after a change produce your worst data, for three separate reasons. Treat that period as observation, not analysis.
  • Before and after on your own numbers is not a comparison. You need something the change did not touch.
  • The shape of the movement narrows the explanation faster than the size of it.

Why early data is the worst data

A staged rollout means part of your audience is on the new behavior and part is not, so any figure covering the window is a blend of two systems in unknown proportion. The blend shifts every day the rollout advances, which produces a smooth curve that looks like a trend and is an artifact of coverage.

New things also draw attention because they are new. Interest that fades on its own gets recorded as an effect of the change, which is the ordinary novelty effect, and it runs in the same direction as the hope that the change is good. The mirror image is real too: people avoid an unfamiliar interface at first and come back later.

Third, you changed. Once you are watching a dashboard daily you post differently, respond faster, and check things you never checked. That is a genuine intervention and it is confounded with theirs.

An open technical notebook with handwritten columns of readings and dated entries
Photo: Technical Notebook, U.S. National Archives via Wikimedia Commons, public domain.

The dated, appended, never overwritten notebook is the whole method. Everything below is a description of what belongs in it.

The sequence

  • Before anything: capture the baseline. Export the numbers you will want to compare against, at the granularity you will want them, and store them somewhere that is not a screenshot.
  • Write the prediction. One sentence saying what you expect to see and by when, dated. This is the only defense against explaining any outcome afterwards.
  • Freeze your own releases for the observation window if you can. If you cannot, log every one of them with a timestamp.
  • Observe for a period long enough to cover the rollout and at least two full weekly cycles.
  • Only then compare, and compare against something external to the change.
  • Decide, and record the decision with the evidence that supported it, including what you could not rule out.

The step people delete is the freeze. It is also the one that decides whether the analysis is possible, because a site that shipped four things during the window has four candidate explanations and no way to separate them.

Shapes worth recognizing

What the series does Most consistent with Check next
A step down that holds flat at the new level Something eligibility-shaped: a filter, a policy, a template no longer qualifying Whether the drop is uniform across pages or concentrated
A spike, then decay back toward the old level Attention rather than a durable change Whether the tail is above or below where it started
A slow drift over weeks with no clear edge A staged rollout, a seasonal move, or a slow competitor gain Whether a comparable series drifted too
A hole for hours or a day, then full recovery Availability or collection, not ranking The provider status record and your own logging
A change in one segment only Something scoped by region, device or account type Whether the segment boundary matches a known scope

Shapes are hypotheses. The value is that each row sends you to a cheap check rather than to a rebuild. A uniform step across every page and a step concentrated in one template are different problems with different fixes, and telling them apart takes an hour.

The comparison that actually works

You need a series that the change could not have touched, moving in the same period, for the same audience. Candidates: a surface the change did not cover, a country outside the rollout, a section of your site you did not modify, or a different channel entirely. Compare the change in your affected series to the change in the untouched one. If both moved together, you are looking at the season or the news, not the update.

Where no clean comparison exists, the fallback is to model the series before the change and ask whether the observed values after it depart from the extrapolation. That is interrupted time series analysis, and it is weaker than a real control because it assumes nothing else changed. It is still much better than eyeballing two averages. The same reasoning about controls, units and power that governs a deliberate experiment applies here, and the protocol for testing a ranking change sets it out in full.

When to act

Act when the movement is larger than your normal variation, has persisted past the rollout window, is present in the comparison, and points at something you can change. That is four conditions and most alarms fail at least one of them.

Acting early is expensive in a way that is easy to miss. You spend the effort, you lose the clean baseline, and if the movement was going to revert you will now credit the reversion to your fix and repeat the same work forever. The knowledge that a movement was noise is worth having, and it is only available to people who waited. What to keep in the record so the wait pays off is covered in the statistics worth keeping on feature changes, and where to look for the account of what actually shipped is in the comparison of sources.

Common questions

How large does a movement need to be before it is worth investigating?

Larger than the range your metric covers in an ordinary month. If you do not know that range, that is the first thing to work out, and it costs one afternoon with historical data.

Everyone in my industry is reporting the same drop. Does that confirm it?

It raises the chance that something external happened, which is useful. It does not confirm the cause, and it does not tell you the size, because the people posting are the ones who dropped. Treat it as a prompt to check a control series.

Can I just ask support what changed?

Ask, and expect a general answer. Support sees the public documentation and a script. A specific confirmation is a gift; its absence is not evidence.

What if the update genuinely was catastrophic for us?

Then the sequence still holds, only compressed. Capture the baseline, identify the shape, find a comparison, and fix the thing the evidence points at. Panic-rebuilding a site during a rollout has ended badly often enough that the pattern is well documented in the record of how these changes have played out, and the general background sits in the overview of feature updates.

More in Guides

Latest from Field Desk