Features

Part of Ranking signals: a clear guide with practical examples

Ranking signals updates 2027: guide and criteria

Ranking signals updates 2027 and sourcing: why an operator's own page is the only citable source, what it supports, and how to write about a closed system.

There is a rule about sourcing that sounds arbitrary until you try to break it: about a specific ranking system, the operator's own documentation is the only citable source. Everything else, however careful, is inference from the outside about a mechanism nobody outside can see. This page sets out why that holds, what the operator's own page can and cannot be made to say, and how to write about a system you cannot inspect without either fabricating or going silent.

What to take away

  • The operator is the only party with access to the code, the weights, and the change history. Everyone else is reasoning from outputs.
  • That does not make the operator's page true in every respect. It makes it the only statement with standing, which is a different property.
  • Written this way, the honest form of most claims is conditional: here is what the published document commits to, and here is what would follow if it holds.

Why outside evidence cannot reach the claim

A ranking system is observable only through what it returns. Given outputs alone, many different mechanisms produce the same visible behavior, and no amount of sampling separates them. This is the ordinary situation of reasoning about a black box: you can build a model that predicts the outputs well and still be wrong about every internal detail.

Three specific gaps make the outside view weaker than it looks.

You cannot see the counterfactual. The one comparison that would settle a causal question, what this same item would have done under a different configuration, is unavailable to anyone without the ability to run both.

You cannot see the population. Outside samples are drawn from what is visible, which is what already ranked. Reasoning about the system from that sample is reasoning from survivors, and the problem with a sample selected on the outcome does not shrink with more data.

You cannot see the change history. When behavior shifts, the outside observer sees the shift and nothing about its cause. It could be a configuration change, a policy change, a shift in who is using the product this month, or a difference in your own material. All four look identical from here.

What the operator's own page is, and is not

An operator's published description is a primary statement about their own product. It is citable because it is the operator speaking, and because it is the thing they can be held to. That is the whole of its authority, and it is worth being precise about the limits.

The document does The document does not
Commit the operator to a description in public Reveal weights, thresholds, or ordering
Name the categories of input the system uses Say how much any one input matters
Describe policy: what is demoted, what is removed Describe the code that implements the policy
Establish a date: this was the description then Guarantee that the description is current
Give you the operator's own vocabulary Guarantee that the vocabulary maps onto internals

Read that way, a published description supports a narrow class of sentences well. It supports "the operator states that X is used". It does not support "X is the third most important factor", and no honest reading gets you there.

The same discipline applies to any organization writing about its own methods. Research bodies publish method statements for exactly this reason, and a page like Pew Research Center's account of how it describes its own methods is citable about Pew in a way that no outside description of Pew would be.

Writing about a system you cannot inspect

The practical question is what to do when you need to explain something and the only strong source is a general document. Four moves cover almost every case.

Move the claim up a level. Instead of a claim about one system, make the claim about the class of systems, where the constraint argument does the work. The reasons a feed must select and order at all hold whoever built it.

State the source's own words as the source's own words. "The published description names watch time among its inputs" is defensible. The paraphrase that drops "the published description names" is not.

Convert a mechanism claim into a test. If you cannot support "X causes Y here", you can usually support "here is how you would find out on your own property", which is more useful anyway and is set out in designing a change you can read.

Drop the claim. This is the move people skip. If a sentence needs a specific weight or ordering to be worth saying, and no such source exists, the sentence should not be written. A missing paragraph is cheaper than a fabricated one.

Keeping the citation honest over time

An operator's page can change without notice, and often does. Anything you build on one should record what the page said and when you read it, so that a later reader can tell whether your sentence has come loose from its source. The same failure that makes undated advice unusable applies to your own writing, and the life cycle of a ranking claim is the fastest way to see where your page will end up if you do not date it.

Common questions

Is a platform's own page a biased source?

It is an interested source, which is not quite the same problem. Read it as a statement of what the operator wants to be held to in public. That has real value and clear edges: it tells you what they will defend, not what the system does.

What about an operator's engineering blog or research paper?

Same category, with an extra caution. A paper often describes a system that was in production some time before publication, sometimes a simplified version of it, and the gap is rarely stated. It is citable about the described method rather than about today's product.

Can an independent audit substitute?

Only within the access it was granted. An audit run against outputs has the same limits as any outside study; an audit with real access can say much more, but then the question is what access was given, and that is usually the least documented part.

How do I write a comparison across systems if I cannot cite anything specific?

Compare on structural properties instead of on internals: what surface it is, what the unit is, what the reader arrived wanting. That is what the stage map of a ranking pipeline and the wider set of producers of evidence are for, and neither requires a claim about anyone's code.

More in Features

Reviews

Ranking signals: a clear guide with practical examples

Ranking signals sorted by the five producers of evidence, why a signal is not a factor, four tests for any claimed input, and examples read the right way.

Features

Ranking signals platforms explained with examples

Ranking signals platforms as a stage map: why pipelines have stages at all, how to read a symptom back to one, and what the map deliberately leaves out.

Maintenance

Ranking signals comparison: what to know and why

Ranking signals comparison of what each kind of evidence can carry, why a correlation study cannot support advice, and the comparison worth running yourself.

Rules

Ranking signals research: practical details and examples

Ranking signals research you can run: setting the question narrowly, building a split that cannot flatter you, and reading the result without fitting a story.