One Held Outcome Cannot Yet Answer a Set of Judgments
At 19:00 on Thursday, Gu Ning turned to the next page of the continuous record. Tang Ke asked: if a judgment that adopted three quarters later comes to pass, does that show it was well calibrated? Gu Ning said we can record this outcome, but we still cannot evaluate how a set of three-quarters judgments performed from a single holding.
Three quarters does not hold every time, nor is it a schedule guaranteeing three holdings in every four. To check the correspondence between probability and outcome, one needs to find the corresponding judgments actually issued, state the scope and version, and then look at the frequency of holding among them.
This chapter remains within the fictional Chengwan. R17 currently has no final probability adoption and no complete confirmation outcome. What appeared at 19:00 is a discussion of evaluation method, not R17 having entered a set of scorable results. The batches in what follows are all set up as independent examples and do not join Chengwan's original registration.
The previous chapter gave the ledger entrance; this chapter uses the entrance to examine calibration. If the objects are unclear, the values not genuinely adopted, and the outcomes not adjudicable, one cannot simply pick a probability out of a narrative and pair it with a later ending.
The Correspondence of Probability and Conditional Outcome
For a binary event, write "held" as Y=1 and "not held" as Y=0, and the ex-ante issued affirmative probability as p. Ranjan and Gneiting's Combining Probability Forecasts (2010) formulate calibration in an explicit probabilistic framework as P(Y=1 | p) = p.
This means that, within that framework, the probability of holding corresponding to a judgment that reports a given probability should agree with the reported value. It is a relation between forecast and outcome, not merely whether the reported value lies between zero and one, nor whether the recorder sounds confident.
An actual finite ledger does not directly let us read the conditional probability itself. We can group identical or similar reported values and observe the corresponding frequencies as diagnostic material. The frequency agreement found by diagnosis and the property above, which has been fully proven, must be kept apart.
This chapter adopts the binary-event formulation and does not extend into the different calibration concepts for predictive distributions of continuous variables. The formal framework makes our question explicit; it cannot obtain for Chengwan probabilities not issued, outcomes not known, or infinitely many repetitions.
Overall Mean Agreement Is Not Enough
Take an independent twenty-object example whose task definitions and relative windows correspond, all values issued while the corresponding outcomes were unknown, and every item finally fully adjudicated. Ten belong to an ex-ante identifiable lower-opportunity group and ten to a higher-opportunity group.
Given outcomes: in the lower group two hold and eight do not; in the higher group eight hold and two do not. Ten of the twenty hold in total, a frequency of one half. This setting merely supplies an example ledger; it does not mean we already know such stable conditional groups exist in reality.
The first set of reports gives one half for all twenty items. Its reporting average is one half, and the outcome frequency is also one half. The second set reports one fifth for the lower group and four fifths for the higher group, with a reporting average that is also one half.
The two sets agree in overall mean, yet their judgment structures differ. Keeping only the overall mean, the reader cannot see who reported different degrees for different objects. Overall agreement is one relation; it cannot complete the whole examination of probability grouping.
How the Second Set of Reports Enters Probability Groups
In the second set, the one-fifth group has ten items with two holdings, an observed frequency of one fifth; the four-fifths group also has ten items with eight holdings, an observed frequency of four fifths. In this finite sample the two reporting groups agree separately.
We can accurately say that the two observed group frequencies of the example agree with the reported values. We cannot take one further step and say that each object's true probability of holding has been verified, or that the next batch will necessarily agree in the same way.
The frequency check preserves the number of items per group. Ten items and a thousand items, even at the same proportion, provide different diagnostic material; but quantifying that uncertainty further requires explicit sampling, dependence, and other assumptions—the count alone does not directly generate precision.
The first set of reports has only a single one-half group, with ten of twenty holding, and is also in agreement on that reporting group. It has not automatically become miscalibrated by failing to distinguish objects. Calibration and the ability to express differences between objects are related but different questions.
Calibration Alone Does Not Accomplish Discrimination
If the current material can identify the lower and higher groups before the outcome, the second set of reports expresses that distinction. The first set remains in agreement on the overall frequency, yet has not put this available distinction into the probabilities.
So rather than asking only which is better calibrated, one may also ask: when a distinction has support, can the judgment assign appropriate degrees to different objects? This question involves discrimination, and it cannot be solved by arbitrarily pushing numbers toward zero or one.
Without the corresponding material, more extreme reporting may only sound more confident in tone. Calibration diagnostics can constrain such degrees, and later scoring will also help compare, but one lucky strike by an extreme report does not mean discrimination has been obtained.
This chapter does not equate discrimination with all usefulness in use. Even if a set of probabilities expresses differences between objects, whether to act and how to allocate resources still require separate judgment. Probability evaluation describes how a representation performs; it cannot accomplish on its own the action allocation of the second volume.
The Same Average Can Hide Opposite Groupings
In the same twenty-object example, take a third set of reports that reverses the pattern, giving four fifths to the lower group and one fifth to the higher group. Its reporting average is still one half, and the overall outcome frequency is still one half.
But the four-fifths reporting group actually has only two holdings, a frequency of one fifth; the one-fifth reporting group actually has eight holdings, a frequency of four fifths. The two deviations leave no trace in the overall average, because their directions are opposite and their magnitudes equal.
This comparison does not need to declare every object of the third set misjudged. It points out precisely that the correspondence between the finite reporting groups and the outcome frequencies disagrees. A single summary number erases the very conditional relation we are examining.
The ledger therefore should not preserve only the reporting mean and the total of holdings. At least when a calibration diagnostic is planned, it must be possible to return to the corresponding probability groups and outcomes. The overall average may be retained, but it cannot replace the omitted grouping.
Reporting Groups and Conditional Subgroups Differ
Return to the first set of reports, which gave one half across the board. At the reporting-value level there is only one group, with observed frequency one half. If it is further split by the ex-ante identifiable low and high conditions, the two subgroup frequencies are one fifth and four fifths respectively.
This does not require retracting the claim that the original reporting group agrees overall. It shows that within certain conditional subgroups the same reported value still carries unexpressed differences. The passing of an overall calibration diagnostic has not checked every finer condition.
Subgroups should be specified by the question and the material obtained. If groups with many or few holdings are assembled arbitrarily after seeing the outcomes, differences are of course easy to find; such exploration can generate new questions, but cannot pose as a distinction that was available ex ante.
We can first ask which conditions there was originally reason to check, and then look at the corresponding subgroups. The calibration question can thereby be refined step by step, instead of reading one overall agreement as reliability under all conditions, or using an arbitrary post-hoc subgroup to declare all judgments invalid.
Grouping Boundaries Are Also Part of the Method
Actual forecasting rarely reports the same value repeatedly; instead, similar numbers such as 60 percent, 70 percent, and 80 percent appear. For convenient summarization they may be placed into a wide group, but the representative value within the group needs to be stated.
If the group contains both 60 percent and 80 percent, one cannot assume every original report was 70 percent merely because the heading says "the 70 percent group." One may compute the within-group reporting average and compare it with the corresponding frequency; the average is a summary, and the relation of the original values should still be preserved.
Which inclusion rule a boundary adopts also affects which group an object enters. When 70 percent falls exactly on the boundary between adjacent intervals, a consistent rule is needed; one cannot, after seeing the outcomes, move the boundary objects that held into one group and those that did not into another.
Finer grouping may reveal more of the differences in reported degrees, but may leave each group with sparse material. Wider grouping increases counts, but opposite deviations may cancel. No set of boundaries obtains full support merely by looking tidy.
One Wide Group Can Hide Two Deviations
Set up another ten-item example, not joined with the previous twenty. Five report 60 percent ex ante, and all five finally hold; another five report 80 percent ex ante, and two finally hold. The within-group reporting average is 70 percent; seven of ten hold in total, and the observed frequency is also 70 percent.
If 60 percent through 80 percent are placed in one wide group, the average and the frequency agree exactly. But viewed separately, the 60 percent reporting group has frequency one, and the 80 percent reporting group has frequency two fifths; the two finite-sample deviations cancel in the wide-group average.
Here again, one cannot infer from five outcomes that the corresponding true probabilities are one or two fifths. It only shows what a wide-group summary omits. Agreement and disagreement within the sample should both be expressed by the actual counts and boundaries.
So a handsome grouped chart can be a useful entrance, but it is not the whole diagnosis. The reader should know the grouping rule, the within-group counts, and the reporting average, and be able, when necessary, to inspect the finer structure. The smoothness of a chart and the support of a statistical relation do not share all their meaning.
Sparse Material Is Not Filled into a Stable Frequency
If a 90 percent report appears only twice and both times hold, the observed frequency is one. This proportion can be recorded, but two holdings do not prove that future 90 percent reports must hold, nor should the mere inequality of one and 90 percent be taken as a thorough refutation of the model.
Finite outcomes fluctuate, and how the fluctuation is explained requires its corresponding conditions. We can first report the counts and the deviation, await more material, and check whether there is an obvious problem of objects or of recording. There is no need to fabricate precision just to obtain a definite conclusion.
Adding material also cannot mean collecting only the objects that are easy to adjudicate or that easily agree with the reports. If a sparse group is later replenished selectively, the scope of judgment may have changed. Growing counts and preserved scope are two relations that must be checked together.
This chapter does not directly give confidence intervals or formal test thresholds. Those require explicit statistical models and conditions of applicability. Not adopting those tools does not prevent us from accurately stating how many items were observed, how they were grouped, and what the current material still cannot support.
Repeated Forecasts Cannot Disguise Themselves as More Final Outcomes
In the previous chapter, L02 had two ex-ante forecasts on the same event. If both enter the diagnostic chart, they correspond to the same outcome. They can be used to study judgments at different time points, but the counts and dependence relations must be stated.
Suppose one event is forecast ten times and finally holds; a "held" label is then added at ten probability positions. This is not ten distinct events all holding. Placing it alongside the material weight of ten independent final outcomes would exaggerate the support available for inference.
This does not forbid evaluating the whole update record. One must first state the question: is it the calibration of first forecasts, the calibration near the deadline, or the diagnosis of each time point in the update process? The version rule and the event relation determine how the outcomes are read.
Therefore the summary should preferably retain both the number of forecast records and the number of distinct events. Giving only one total line count easily leads readers to believe all lines provide repeated material of the same kind. Once the task relation is explicit, the appropriate statistical treatment can then be discussed.
The Adjudicated Subset Still Has a Limited Scope
There is a batch of issued probabilities, some windows not yet closed, some outcome material still awaiting verification. One can currently diagnose the adjudicated part, but should state that it does not automatically equal the calibration of the whole batch.
If the easier-to-confirm objects receive outcomes first, or a certain class of ambiguous messages is harder to adjudicate, the currently adjudicated subset may have a particular structure. We cannot omit how the subset was formed merely because each adjudication is genuine.
Recording all pending items as "not held" changes the target; deleting all pending items without reporting hides the scope. The ledger can simultaneously preserve the total entries, the number adjudicated, the window-not-yet-reached and material-pending-verification categories, giving the diagnosis a clear denominator.
When outcomes are later recovered, a new diagnostic version can be generated. A change in the curve or frequencies may come from the widening scope, not necessarily from the forecasting model being re-adopted. Changes in evaluation also need provenance; they cannot each time be attributed to a change in the recorder's ability.
A Change of Time Position Changes the Available Material
A probability formed Thursday morning under greater unknowns, and a probability formed Friday near the deadline, may face different remaining time and different messages. If these are mixed into one group without explanation, within-group agreement may not answer the ability we truly want to compare.
One may diagnose separately by similar forecast lead times, or explicitly evaluate a multi-time-point reporting scheme as a whole. But when the same values come from different information positions, the reader should know of the mixture, rather than assume the tasks' conditions are naturally identical.
In particular, affirmative statements made after sufficient material for holding has been obtained must be kept separate. They may correctly express a state of knowledge, yet they do not show how probabilities were reported under unknown outcomes. That a sentence uses "will" does not turn a known outcome back into ex-ante uncertainty.
In Chengwan, R17 still has no final probability adopted; one cannot, to fit into some probability group, treat the historical candidate calculation as a formal report of Thursday morning or of 19:00. However clear the methods page, it cannot obtain the facts for missing inputs.
Agreement Across Batches Does Not Guarantee the Rules Have Not Changed
If an earlier batch used the old model and a later batch changed the material entrance and the grouping rules, they should be preserved separately. Mixing the two batches and finding frequency agreement may conceal each batch's deviations in different directions, or may hold for a new mixed scope.
We may evaluate the whole, provided we state what rules and objects the whole comprises. If we want to judge the continuation ability of the new model, we should return to the judgments actually issued after the new rules were formed, not merely look at the same batch of outcomes used to modify the old model.
Discovering that the 70 percent group ran high or low in the old batch can motivate future adjustment, but the new values still need subsequent material. Reassigning the old batch so that the grouping agrees describes a completed fit or restatement, not a new rule validated by new outcomes.
The old records therefore still have value. They help explain why the model changed, while the new records help check how the change performs. The relation between the two batches continues; a single back-calculation cannot compress the rule that formed the model and the rule that tests it into one and the same proof.
The Hit Ratio Does Not Preserve Degrees of Probability
Take another example of one hundred fully adjudicated items, ten held and ninety not held. One set of ex-ante reports gives 10 percent for every item; another gives 1 percent. Under a rule that guesses "held" whenever the report exceeds one half, both sets would uniformly guess "not held."
Under this binary guessing rule, both sets hit ninety items, a hit ratio of 90 percent. But the first set's reporting group has an observed frequency of 10 percent, agreeing with the report; the second reports 1 percent while the observed frequency is 10 percent, a clear deviation in the finite grouping. The same hit ratio has not preserved the difference in degrees between the two sets.
This proves neither that the first set is ideal under all conditions, nor does the hundred-item setting obtain future guarantees on its behalf. It only shows that compressing probabilities into affirmative or negative first removes part of the very information calibration is examining.
Therefore "often guessing right" cannot be used directly as a synonym for "probability calibration." The original reported values should be preserved; if the binary guess has a working use it can be noted separately, but at evaluation time the latter must not be allowed to overwrite the former. The next chapter's degree-based scoring also depends on this distinction.
The Quality of Outcome Labels Still Needs Checking
The "held" labels in a calibration table come from adjudication. If a label counts partial acceptance as complete confirmation, or counts a modified version as the first version, then even if the grouped frequencies agree with the reported values, this is agreement only with respect to that set of ambiguous labels.
Diagnosis therefore cannot replace outcome quality checking. One should first confirm that the labels correspond to the original events, and only then discuss the relation between reports and labels. When an adjudication is found wrong, it can be corrected and a new diagnostic version generated, rather than keeping a misclassification to preserve a pretty chart.
Multiple persons adopting the same labeling rule helps consistent checking, but does not by itself guarantee the rule is adequately supported. Whether the original text and the conditions correspond still requires returning to the material. Neatness within a statistical table and accuracy of material usage are two tasks that can be completed separately and can also go wrong separately.
Chengwan's existing event cards serve precisely here. Receipt of a checklist cannot be upgraded into complete acceptance, and the station being able to schedule cannot be upgraded into acceptance by both parties. Calibration needs subsequent actual forecasts and adjudications, but these existing boundaries should still hold as the counts grow.
Calibration Diagnostics Point to Specific Checks
If some probability group's observed frequency has long run below the reported value, one can check whether a class of conditions was overestimated, whether material was double-counted, whether different windows were mixed in, or whether there is an outcome adjudication problem. The deviation proposes directions of checking; it does not by itself select which of these causes is at work.
If the frequency runs above the report, the same questions must be asked. Reliability means more than avoiding optimism; underestimation can equally make the reports fail to correspond to the material. Direction should be determined by examining the record, not by the prevailing mood of caution or of boldness.
Sometimes the grouping agrees overall, yet a clearly identifiable ex-ante conditional subgroup deviates markedly; that conditional scope must then be reopened. It may also be that there are not enough counts to support finer division; that should be honestly preserved, and one should not force new rules into being just to produce an improvement suggestion.
The calibration page and the competing-explanations page are thereby connected: one page describes how reports and outcomes correspond; the other asks why. The former need not have explained all causes, and the latter cannot claim, on the strength of a single causal story, that the frequency relation has been repaired.
Stating the Boundaries of This Examination Accurately
Gu Ning drafted a diagnostic statement: the batch used, the actual issuance rule, the outcome definition, the grouping boundaries, each group's count, mean report and corresponding frequency, and the pending and excluded scopes. It can be compressed, but the key relations cannot be omitted.
If there is no formal test, the observed deviation is not written as having reached some significance conclusion; if there are no statistical intervals, the width of lines on a chart is not written as a precision range. The state of the method, like the state of forecast adoption, needs to be clear.
Where this observation agrees, we may write "agrees" directly; there is no need to deny it in a show of modesty. What must be avoided is expanding finite agreement into permanent calibration, or expanding the agreement of one group into reliability under every condition.
Different diagnostic purposes can stand side by side; the reader should know which question each chart answers. An overall table, a probability-group table, and a conditional-subgroup table are not the same result repeated three times, but the same batch of records examined from different relations.
Calibration Does Not Mean the Process Has Been Exhausted
In RC's internal stance, probability is a limited representation organized for a current question. Calibration checks connect the reported values with outcomes and can constrain the representation; they do not prove the model already contains all the conditions of reality, nor verify that the foundational philosophy has been empirically established.
Even if some scope continues to agree, it may still need reopening once new conditions enter. Openness does not mean any adverse outcome can be ignored; material already sufficient to show a group's deviation should still be preserved and should motivate the corresponding check.
Feedback can promote calibration when the reports, objects, and adjudications are genuinely returnable. If every time the numbers are retroactively changed, the groups swapped, or the target moved, the curve can look better while losing its constraint on the original judgments, and the old explanation may grow more rigid.
So the work of calibration is not to confer a permanently trustworthy identity, but to keep examining the correspondence within an explicit scope. The recorder may provisionally adopt a supported representation while preserving the counts, the scope, and the conditions of change.
At 19:20, Saving the Diagnostic Questions
At 19:20 on Thursday, Gu Ning saved the structure of the calibration page. Actual reporting groups, overall means, and conditional subgroups were kept separate; counts, pending items, versions, and batches had returnable positions; the agreement of the independent examples served only as methodological illustration, not written as Chengwan performance.
R17's formal probability is still not formed, and the complete confirmation is still not obtained. The historical five confirmations and the rough B condition table retain their original use and are not padded out into a batch of issued and adjudicated probability forecasts. The 19:00 discussion also does not change the Friday 17:00 deadline.
The next chapter will adopt an explicit scoring rule to compare the relation between degrees of probability and single outcomes. Scoring can supplement diagnosis, but it can also be misread when the scope is confused or the incentives change; it needs the objects preserved by this chapter and the previous one, and cannot let a single average sign off on all the judgments.