GuideResearch-backed

Base Rates vs Stories: Why the Outside View Beats Intuition

Start forecasts with a relevant reference class, then update for case-specific evidence instead of letting a coherent story erase historical frequencies.

The reference-class ladder. A worksheet for testing broad, narrow, recent, and mechanism-matched reference classes before updating a forecast with local evidence. Download the SVG asset.
Direct answer

Select a defensible reference class before reading the case story in detail. Estimate the class’s outcome distribution, record why the case belongs, and update only for evidence that changes the likelihood of the outcome. If several reference classes are plausible, show the forecast under each instead of hiding the choice.

Stories suppress the denominator

A founder is unusually determined. A student uses an elegant study plan. A project has an admired leader. These facts may matter, but narrative coherence makes them feel more diagnostic than they are. Prediction requires another question: what happened to comparable cases?

The inside view simulates the focal case from its plans and details. The outside view begins with a distribution of outcomes among a reference class. The outside view often disciplines optimism, but “use the base rate” is incomplete. Which base rate? Choosing the class is itself a model decision.

Evidence inside the case boundary: the reference-class ladder

Evidence snapshotHigh confidence

Classic research on intuitive prediction found that people can let descriptive similarity dominate prior probabilities. Kahneman and Lovallo argued that forecasting from the focal plan can produce overly bold predictions relative to comparable ventures. Large forecasting tournaments found substantial and persistent differences in probabilistic accuracy, supporting trainable practices such as decomposition, updating, and comparison.

kahneman-tversky-base-rate, kahneman-lovallo, mellers-superforecasting

Claim sources: kahneman-tversky-base-rate, kahneman-lovallo, mellers-superforecasting

Build a reference-class ladder

List at least four candidate classes:

  1. Broad: all projects of this general kind.
  2. Narrow: projects with similar scale, maturity, and constraints.
  3. Recent: comparable cases under the current environment.
  4. Mechanism-matched: cases sharing the causal bottleneck most relevant to the outcome.

For each, record inclusion criteria, sample size, time period, measurement, survivorship, and outcome distribution. Do not choose only the class whose rate supports your preferred action.

The broad class offers stability but may ignore crucial differences. The narrow class offers relevance but may become tiny or selected. A range across defensible classes is often more honest than one precise prior.

Run the OUTSIDE update

O — Outcome: Define what will happen, by when, and how it will be measured.

U — Universe: Specify the population from which comparable cases could have been observed.

T — Taxonomy: Build the reference-class ladder and justify inclusion.

S — Starting rate: Record the outcome distribution before local detail.

I — Individual evidence: Identify case facts that are causally or statistically diagnostic.

D — Direction and size: Move the estimate explicitly, not merely verbally.

E — Exceptions: State structural change and evidence that would invalidate the reference class.

The initial probability need not be perfect. Making it visible prevents the story from becoming an unrecorded prior of nearly zero or one.

A worked forecast

A team estimates that a new data migration will finish in eight weeks because the architecture is clear and the engineers are experienced.

The project plan is useful for identifying tasks. It is not yet a calibrated duration forecast. The team assembles three classes: all prior migrations, those of similar data volume, and those involving the same legacy dependency. Completion times are recorded, including abandoned and delayed projects rather than only successes.

The mechanism-matched class has the widest delays because external data owners must approve mappings. The current project has an experienced team, which supports an upward update, but it also depends on the same approval bottleneck. The resulting forecast becomes a range with milestones and a probability of finishing by week eight—not a confident date attached to a story.

The rival model: this time is different

Sometimes it is. A new regulation, technology, epidemic, incentive, or measurement definition can break historical comparability. Local evidence should dominate when it identifies a genuine structural change linked to the outcome.

But novelty must be specific. “Our people are exceptional” is weak unless exceptional performance predicts the bottleneck. “The process now eliminates the approval stage that caused most historical delay” is diagnostic and testable. The outside view should yield to mechanism, not enthusiasm.

Adversarial reference classes

Ask two people to choose the reference class independently: one who benefits from action and one responsible for downside. If their classes differ, expose the selection rule before debating the rate.

Stress-test:

  • excluding survivors that conceal failed cases;
  • shifting the historical window;
  • measuring the outcome consistently;
  • separating projects that stopped for strategic reasons;
  • choosing the class before revealing the focal case’s desired forecast;
  • comparing a causal bottleneck, not superficial resemblance.

This prevents “base rate” from becoming an authoritative costume for motivated selection.

When the reference class behind the reference-class ladder breaks

Give the outside view less weight when the data-generating process has changed, the reference class mixes mechanisms with opposing effects, measurement is incompatible, or strong local evidence has a validated relationship to the outcome.

Give it more weight when the case story is vivid but the proposed difference has no demonstrated predictive value, when incentives reward optimistic forecasts, or when similar plans have repeatedly missed in the same direction.

If no useful class exists, do not fabricate one. Decompose the event, use multiple models, make a wider forecast, and design information-gathering steps.

Base-rate failures

  • Citing a population rate that excludes the focal case.
  • Choosing the class after seeing the desired conclusion.
  • Treating a tiny narrow class as certain.
  • Ignoring abandonment and survivorship.
  • Refusing any update because “base rates always win.”
  • Updating for prestige, confidence, or detail with no predictive link.
  • Replacing a probability with “likely” and avoiding accountability.

The bounded verdict from the reference-class ladder

Limits and counterevidence

Reference classes depend on classification, measurement, and historical stability. They can encode discrimination or past constraints that should not govern an individual decision. Rare events and structural breaks may offer little comparable data. High-stakes forecasts should show sensitivity to multiple priors and avoid using group rates as automatic judgments about a person.

Test the local causal mechanism, use the probability in an expected-value comparison, and assess repeated forecasts through calibration.

Named sources

Evidence and further reading

  1. On the Psychology of Predictionresearch · accessed 2026-07-28
  2. Timid Choices and Bold Forecastsresearch · accessed 2026-07-28
  3. Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictionsresearch · accessed 2026-07-28
Publication record

Published July 29, 2026. No substantive revision has been recorded. Evidence last verified July 28, 2026.