Chengwan Repair Station receives inquiry R17. One person hears a tone and says 70 percent; another points to eight successes on the display wall; a third assigns fifty-fifty because nobody knows. As new messages arrive, everyone quietly rewrites what they originally meant, so all appear correct in hindsight.
If everyone can rewrite the original forecast after the outcome, what are we judging?
A reliable probability judgment is not a guessed number but a process in which the object, evidence, version, and adjudication can return to the original question
"This commission will probably close" sounds like a judgment but says neither what counts as closing, when it will be adjudicated, which materials support it, nor whether later messages changed the version. If it cannot be checked or proven wrong, is it probability, hope, or an impulse to act?
⏱ 8 min
01 · Stop One: Give Judgment an Object
Stop assigning numbers. Fix the event, outcome, time, observation window, sample entry, and type of unknown in an event card that can be revisited.
Define the event
Prediction is not hope or action
R17 needs an identity, outcome condition, confirming material, and version boundary. Hoping to close, preparing materials, and receiving formal confirmation within the window are different events.
Fix the timeline
Preserve what was visible then
Issue time, message arrival, discovery, and adjudication remain distinct. Later evidence may form a new version but cannot flow backward; late outcomes cannot silently change a closed window.
Trace the sample
The display wall is not the population
Eight success cards must be traced to sixteen historical inquiries, including units, duplicates, missing states, and entry paths. Selected visible outcomes cannot directly become event frequencies.
Classify the unknown
Unknown is not automatically fifty-fifty
Missing past data, unexamined present conditions, omitted representation, and genuine future variation call for different responses. Investigate what can be checked and preserve boundaries around the rest.
Common Misreadings
✗ Give a rough probability first and define it precisely later.
✓ A change in object, window, or adjudication creates a different question. Number-first reasoning encourages definitions that merely rationalize intuition.
02 · Stop Two: Let Evidence Actually Update Judgment
When conditions change, do not nudge the old number by feel. Recheck the comparison class, direction of evidence, dependence among sources, and material quality.
Rebuild the comparison class
Recount after conditions change
A base rate applies only to a defined starting state, window, and set. If R17 gains budget confirmation or decision-maker involvement, seek records with those conditions rather than reasoning backward from all outcomes.
Identify what a message changes
One reply has limited reach
A message may exclude one failure cause, change timing, or repeat old information. State where it is more expected and what it fails to distinguish instead of awarding points for tone.
Check source relations
Three people may repeat one message
Sources sharing an upstream message, citing one another, or passing through one filter are not independent evidence. Map the chain before deciding how much information was added.
Resolve conflict through quality
Quantity cannot replace traceability
Check acquisition time, object match, completeness, and bias. High-quality counterevidence can contract a judgment; a low-quality majority need not prevail.
Common Misreadings
✗ Every new piece of evidence should move the probability up or down.
✓ It may be nondiscriminating, already contained in earlier evidence, or relevant only to another version. First state which condition it changes.
03 · Stop Three: Recognize Noise Disguised as Common Sense
Survivorship, correlation stories, causal explanations, and consensus wrap limited materials in apparent common sense. This part removes the wrapping.
See the selection mechanism
What is displayed is not all that happened
Success stories, respondents, and complete records are selected. Unobserved outcomes cannot be filled in arbitrarily, but their absence must limit inference.
Separate co-movement
Correlation is not causation
Variables may move together because of a common condition, time trend, or filter. Mechanisms need comparisons, temporal order, and alternatives; a fluent story is not enough.
Stop stories from overwriting
Outcomes make the past feel inevitable
Hindsight reorganizes earlier evidence and turns unrecorded hesitation into foresight. Original versions and adoption times keep ex ante judgment separate from later explanation.
Examine consensus entry points
What everyone says changes what becomes visible
Consensus affects reporting, preservation, and responses, thereby shaping later data. Majority opinion can be evidence and part of the evidence-generating mechanism at once.
Common Misreadings
✗ A large enough sample automatically removes selection bias and consensus effects.
✓ If entry continuously favors one outcome, more observations reproduce the bias more precisely. Quantity cannot replace an account of provenance and missingness.
04 · Stop Four: Preserve Judgment as a Cycle
Record probability judgments, calibrate groups, interpret scores, preserve counterevidence, and adopt provisionally so each judgment can feed the next observation.
Preserve the forecast record
Object, version, evidence, and time together
Keep identity, issue time, probability, grounds, adoption status, window, and adjudication material. Append the outcome; never overwrite the original.
Use calibration correctly
Compare a group of like judgments
Calibration asks how often outcomes occur across forecasts assigned similar probabilities. Scores compare fuller predictions but depend on difficulty, grouping, and incentives; they are not single-event rankings.
Let judgments contract
Counterevidence is not an enemy
A reliable system preserves counterexamples and identifies which inference they weaken. Provisional adoption authorizes limited use, not permanent correctness.
Define reopening and handover
Close the old event; version the new question
After R17's original window closes, a later opportunity becomes a new event. A successor should see the old judgment, actual use, adjudication, and unresolved questions without rewritten history.
Common Misreadings
✗ The outcome is zero or one, so a prior 0.6 forecast was meaningless.
✓ One result cannot assess probability quality alone. Check event definition, evidence, and version preservation, then examine calibration across comparable forecasts.
Key Concepts
Event card
A judgment unit preserving the object, outcome conditions, confirming material, time window, and version boundary.
It lets the forecast return to its original question after the outcome.
Comparison class
Historical objects sufficiently comparable in starting state, conditions, and observation window.
A base rate has meaning only relative to a comparison class.
Source chain
A map of provenance, acquisition time, shared upstream sources, and selection processes.
It prevents repeated transmission from masquerading as independent evidence.
Calibration
A long-run comparison between groups of like probability judgments and outcome frequencies.
It shifts evaluation from one lucky hit to cumulative judgment quality.
Provisional adoption
Using a version within a stated object, evidence base, and purpose while preserving counterevidence, boundaries, and reopening conditions.
It permits action without sealing limited evidence into permanent truth.
Map of the Book
Part One: Where Judgment Begins Where judgment begins — events, time, samples, and unknowns
Part Two: How Evidence Changes Judgment How evidence changes judgment — comparison classes, updating, sources, and quality
- Chapter 7: Recounting After Conditions Change
- Chapter 8: The Pitfall of Inferring Conditions from Outcomes
- Chapter 9: What One Message Distinguishes
- Chapter 10: Why Updating Is Not Merely Enlarging an Old Number
- Chapter 11: Are Multiple Sources Independent
- Chapter 12: When Materials Conflict, Check Quality First
Part Three: How Noise Disguises Itself as Common Sense How noise becomes common sense — selection, survivorship, causation, and consensus
- Chapter 13: The Selected Sample
- Chapter 14: Survivorship and Unobserved Outcomes
- Chapter 15: Changing Together Is Not Causing Each Other
- Chapter 16: What Further Comparisons Does Explanation Require
- Chapter 17: How the Story Overwrites the Original Prediction
- Chapter 18: How Consensus Changes Visible Material
Part Four: How Judgment Keeps Generating How judgment continues — records, calibration, contraction, and handover
After reading, you will understand
- Define an adjudicable event before discussing a probability
- New messages cannot rewrite an old judgment; they create a new version
- Base rates depend on comparison classes; displayed successes are not populations
- Multiple sources may share an upstream origin, so counts do not equal independent information
- Causal stories and consensus can help manufacture the evidence later observed
- Good judgments are reviewable, contractible, reopenable, and transferable