FORM NOT VOID, MIND NO CORE

Chapter 21: How Scoring Helps and How It Misleads

2026.09.13

A Score First Needs a Convention

At 19:40 on Thursday, after the calibration page, Gu Ning wrote down the scoring question. Tang Ke wanted one short number that could compare reports of different degrees, instead of looking at each probability group separately. Gu Ning said a score can be adopted, but first the number's computation, and whether larger or smaller is better, must be made clear.

The previous chapter's calibration page checks probability groups and their corresponding frequencies. A score can instead give a number to each actual report and its outcome, then summarize over an explicit scope. The two have different uses; one average cannot replace the diagnosis of every condition.

This chapter continues the fictional Chengwan. At 19:40 there is no new reply from the commissioning party, no complete confirmation, and no final probability adoption; R17 still has no numerical forecast paired with an outcome that can be formally scored. The scoring computations use the already-specified independent examples and are not padded into mainline results.

We choose one simple rule so the reader can verify item by item. The purpose is not to hand the recorder a permanently reliable ranking, but to show how degrees of probability accept the constraint of outcomes, and how the rule loses its original question when used improperly.

This Book Uses the Squared Error of the Affirmative Probability

For a binary event, write the affirmative probability as p, "held" as Y=1, and "not held" as Y=0. This book adopts the per-item Brier score BS = (p - Y)^2, with the convention that smaller is better. If there are n reports fit for the corresponding evaluation, the average score is the sum of these per-item squared errors divided by n.

Gneiting and Raftery, in Strictly Proper Scoring Rules, Prediction, and Estimation (2007), discuss the quadratic score and strictly proper rules. Their paper uses a larger-is-better formulation; we explicitly switch here to the loss direction and do not mix numbers of different directions.

There is also a convention that sums the squared errors over all categories. For a binary event it computes both the affirmative and the negative category, yielding twice this book's per-item value. A uniform factor does not change the ordering within the same scope, but it changes the reported scale of the score.

Therefore, upon seeing a Brier score elsewhere, first ask about the convention and the direction. Throughout the rest of this book we use the squared error of the single affirmative probability, with values from zero to one, smaller being better. An identical name does not mean two numbers on different scales can be compared directly.

The Same Probability Faces Two Outcomes

If p = three quarters and Y = 1, the error is negative one quarter, and its square is one sixteenth, i.e., 0.0625. If the same report faces Y = 0, the square is nine sixteenths, i.e., 0.5625. A report giving a higher affirmative degree bears a larger loss when the event does not hold.

If p = one half, then whether the outcome holds or not, the squared error is one quarter. It has not obtained the best score by virtue of a neutral tone; its per-item loss is merely the same under both outcomes.

If the report is one and the outcome holds, the score is zero; if the report is one and the outcome does not hold, the score is one. A zero shows that this report-outcome pair fits perfectly under the rule; it does not show the recorder had sufficient ex-ante grounds for adopting one.

Per-item scoring links degree and outcome, yet it cannot infer the whole of the ex-ante basis backward from the actual outcome. An unsupported extreme report may also obtain zero once; a supported non-extreme report may encounter a less favorable realized outcome.

One Poor Result and Honest Reporting Can Coexist

Take an independent thought experiment: a recorder, judging carefully under existing information, takes the probability of holding to be three quarters. She reports three quarters; the event later does not hold, and the actual loss is nine sixteenths.

Had she instead reported one half at the time, this non-holding would have cost one quarter, a better actual result. But this is a single realization compared after seeing the outcome; it does not prove that reporting one half ex ante would have better matched her judgment at the time.

Scoring does not promise that honest reporting wins every time. It must be discussed at the position where the outcome is not yet known: under the probability the recorder has adopted, which way of reporting has the smaller expected loss. The expectation is computed over both outcomes and their probabilities, not over the one that happened, alone, after the fact.

This distinction protects not anyone from evaluation, but the question of evaluation itself. A poor actual score should be preserved and may prompt checking; but every poor outcome cannot automatically be read as dishonesty or as proof that the original probability should not have been adopted.

How the Expected Loss Attains Its Minimum

Let q be the affirmative probability the recorder has carefully adopted at the time, and p the probability actually reported. If the two possible outcomes are evaluated according to q, the expected squared loss is q(1 - p)^2 + (1 - q)p^2.

Expanding, this equals (p - q)^2 + q(1 - q). For fixed q, the second term does not change with the report p, and the first term is uniquely zero when p = q. Therefore, under this condition, reporting according to q has the uniquely minimal expected loss.

For example, with q = three quarters, reporting p = q gives an expected loss of three sixteenths, i.e., 0.1875; reporting one or one half instead gives an expected loss of one quarter either way. Reporting one fits the holding outcome better, but still has to face a one-quarter chance of non-holding.

Here q is the careful judgment given in the thought experiment, not a claim that we already know R17's true probability. The formula states an incentive property of reporting; it neither obtains sufficient material for anyone's q nor proves that the carefully believed probability is calibrated.

Strict Propriety Is a Conditional Property

Ranjan and Gneiting's Combining Probability Forecasts (2010) state strict propriety in the smaller-is-better loss formulation: reporting the probability being assessed should have strictly smaller expected loss than reporting any other probability. It is not a ranking guarantee holding for every actual outcome.

In the computation above, the premises include that outcomes are adjudicated according to the corresponding binary target, that the reporter takes expected loss as the standard of comparison, and that the rule has not quietly changed elsewhere. If the target of evaluation changes, the original property cannot sign off on the new target directly.

Honest reporting and judgmental accuracy are likewise distinct. A person can honestly report an estimate based on insufficient material, while subsequent frequencies show deviation; a person can also deliberately alter a number and happen to guess right. The scoring property removes one reason to alter the report; it does not automatically supply observation and modeling.

We therefore check both the reporting rule and the raw material and sustained performance. Strict propriety is not a moral certificate for a person; it is a property of a rule that can be stated under explicit mathematical conditions.

How the Small Ledger's First Reports Summarize

Return to the previous chapter's independent ledger, L01 and L02. L01 first reported one quarter, outcome not held, loss one sixteenth; L02 first reported three quarters, outcome held, loss also one sixteenth. The average of the two first reports is still one sixteenth.

This average accurately describes the two first reports; it does not show that two recorders, or one recorder, have become reliably consistent. We have only two example events, and no assumption of outcome independence or sampling from some population has been given.

L02 later updated to three fifths; under the same holding outcome the loss is four twenty-fifths, i.e., 0.16. This update's per-item score is larger than the first report's. It can be recorded as it stands, but one outcome alone does not refute the whole basis of the update.

If the update process is studied by bringing all three reports in, the average is 0.095, i.e., nineteen two-hundredths. The computation has three records but still only two event outcomes. This scope differs from the average of the two first reports and cannot be treated as the same result without explanation.

Comparing Sets of Reports Within the Same Event Scope

In the previous chapter's twenty-object example, the lower group of ten had two holdings and the higher group of ten had eight holdings. The first set reported one half throughout; each item loses one quarter, so the average is 0.25.

The second set reported one fifth for the lower group and four fifths for the higher group. In the lower group, the two holdings each lose 0.64 and the eight non-holdings each lose 0.04, totaling 1.6; the higher group's losses also total 1.6. The overall average is 3.2 divided by twenty, i.e., 0.16.

The third set reversed the two groups' reports. In the lower group, the two holdings each lose 0.04 and the eight non-holdings each lose 0.64, totaling 5.2; the higher group likewise totals 5.2, and the overall average is 0.52. The three sets are compared within the same event and outcome scope, with identical scale and denominator.

That the second set has the smaller average this time both preserves degrees and corresponds to the example's conditional distinction. But the ordering of one finite set is not a permanent ordering of the future. The calibration page still has its use: it shows how each probability group corresponds, while the average gives only one summary.

A Good-Looking Average May Omit Local Problems

In one batch, a large number of easy objects score very low, a few difficult objects score high, and the overall average may still look decent. If the work most needs judgments on the difficult objects, giving only the overall average may not answer the user's question.

One may keep the overall result while also examining the subgroups of tasks there was ex-ante reason to distinguish. When a subgroup is small, state its count too; do not immediately declare a stable defect upon seeing a poor score, and still less cherry-pick the most favorable scope by cutting groups after the fact.

Scoring and calibration can therefore be used together. The average score helps compare reports within the same scope; calibration and conditional diagnosis help reveal the structure of correspondence; both need the ledger's objects and versions, and neither replaces all the other's checks.

If some subgroup does not belong to the original purpose of evaluation, excluding it can be reasonable, but the reason must be written down and applied consistently. Exclusion should not be decided by the known scores; otherwise the scope of evaluation has changed after the fact.

Omitting Pending Items and Errors Changes the Denominator

L03 is not yet adjudicated and should not be given a squared error now; but it remains on the list of issued reports. The summary can say that two first reports are currently scorable and one ex-ante report is pending; it cannot merely say the batch has two items and let the third vanish.

If outcome material is missing, a recording error, or an unclear event definition, scoring can also be deferred, with the reasons stated separately. Unscorable lines are retained as material for process checking; there is no need to force an outcome in to make the average complete.

If the objects with larger losses are more easily classified as anomalies and deleted, the change in the denominator directly affects the result. One should first see whether an anomaly has a genuine target or material reason, rather than finding an exit for high values under the heading "atypical this time."

Later, when pending items close or outcomes are corrected, the average can change. The change may come from new adjudications, not from a new run of the model. The summary version marks the provenance so the reader sees when a given batch had which inputs.

Reports of Known Outcomes Cannot Be Mixed into Ex-Ante Results

L04 wrote down the affirmative status one after sufficient material for holding had been obtained; the squared error is zero. The computation is not wrong, but the task position differs from forecasting under an unknown outcome.

If many such statements are added, the average will shrink; but this mainly shows that known material has been correctly restated, not that probability judgment at unknown time points has improved. Formally identical p and Y do not mean the tasks are the same.

The evaluation page should distinguish according to the task scope stated in advance. If the purpose was in fact to express a state of knowledge, L04 can be checked separately; only its zero must not be allowed to obtain, on behalf of L01 through L03's ex-ante forecasts, a certificate of ability.

Likewise, forecasts made very near the deadline with more material in hand should state their lead time and information scope when mixed with early forecasts. Only when the version rule is clear do the scores have a comparable position.

Rewarding Only Affirmative Outcomes Changes the Reason for Reporting

Consider an independent erroneous rule: when the event holds, the reported affirmative probability p is rewarded; when it does not hold, there is neither reward nor deduction. If the recorder judges the probability of holding to be q, the expected reward is q times p.

As long as q is greater than zero, pursuing the expected reward under this rule favors reporting one rather than reporting q. Even if q is only one quarter, reporting one still yields a higher expected reward than reporting one quarter.

This requires no presumption that anyone deliberately exaggerates. The rule gives, mathematically, a reason to alter the report; how actual persons behave requires further conditions. What we see first is that a number called a score does not necessarily encourage reporting true degrees.

If the work only commends bold affirmations when events hold while forgetting the records of non-holdings, its recording structure may also approach this one-sided reward. Whether this is so should be checked against the evaluation rules and the complete ledger, not judged merely from someone's confident tone.

Doubling by Outcome Is Not an Ordinary Scale Adjustment

Take another thought experiment: given q = one half, multiply the squared loss for the holding outcome by two while leaving the non-holding outcome at its original value. The expected loss is then (1 - p)^2 + one half times p^2.

Reporting one half gives an expected loss of three eighths; reporting two thirds gives one third, which is smaller. This change no longer lets q itself be the optimal report, so one cannot simply say that weighting some loss more heavily still preserves the original honesty property.

It differs from multiplying the whole score uniformly by two. A uniform factor changes only the scale; changing the factor by outcome changes the comparison between the two possible outcomes. When different weights are used, one must re-examine what is actually being evaluated.

Some error in action may indeed be more costly, but action loss and the scoring of probability reports are not the same object. Different costs can be discussed in earnest, but cost weights cannot be slipped into the score without explanation, with the claim that the new score still measures the same kind of probability degree.

The Purpose of Weights and Rankings Needs Stating

Different tasks may differ in importance to the work. If each item is given a fixed weight before the outcome is known, and what is being evaluated is explicitly the weighted task scope, the summary can proceed over that scope. The provenance of the weights should be preserved and not decided by the results.

It also affects the overall performance the user sees. If one class of tasks carries a large weight, the whole represents that class more; this need not be wrong, but one can no longer say the average represents every item equally in some overall population.

If the weights change with the reports, the outcomes, or post-hoc preference, the rule's properties need further checking. One cannot assume that because the per-item squared error is strictly proper, arbitrary aggregation and reward procedures inherit the same property.

Especially when only the ranking matters and expected loss does not, the participants' objectives may again differ. This chapter derives no contest strategy; it only preserves the boundary: the expectation property of a standard rule does not unconditionally guarantee reporting behavior under every ranking institution.

Choosing Tasks for the Score Changes the Claimed Ability

A recorder who accepts only easy-to-judge objects may obtain a smaller average. Such a result can describe the actual tasks; it cannot silently expand into judgmental ability over all contacts.

If the task scope changes between periods, the batches and entry rules should be preserved. Improvement in the average may come from clearer material or from model improvements; a falling score alone cannot decide which.

A reasonable work scope may explicitly contract, without having to accept every question to prove ability. Only, after contracting, the conclusion should contract with it: how things stand on these tasks—not writing the excluded difficult tasks off as resolved.

The completeness of the evaluation scope matters as much as the scoring formula. A correct formula proves only that the given inputs were computed correctly, not that the inputs were free of favorable selection. The continuous ledger gives this check a returnable position.

A Comparison Baseline Also Needs a Time of Issue

Seeing an average of 0.16, a reader may ask whether it is already good. A single number does not choose the comparison object for us. One may compare with another set of ex-ante reports in the same task scope, or with an explicit simple baseline, but the baseline's provenance needs to be stated.

For example, the previous chapter's set of all-one-half reports is a set of reports given in advance within a thought experiment, and can therefore be computed on the same outcomes. In real work, if one sets a constant from a batch's holding frequency only after seeing those outcomes and then calls it a pre-existing forecast baseline, later material has been borrowed.

A back-computed baseline may still have diagnostic use. It helps observe how a complex set of reports compares with a post-hoc constant; it just cannot be written as a then-available, genuinely issued competing forecast. If a constant is to be adopted in the future, its formation batch and its subsequent issuance positions should be preserved.

The comparison must also preserve the corresponding objects. If another model did not answer certain difficult tasks, its smaller average cannot be ranked directly against a model covering all objects; one may first compare on the common scope while reporting each side's uncovered part. A conclusion on the common scope does not represent the whole scope.

Returning from a Score to an Original Record

Gu Ning envisioned a reader who, seeing some 0.5625, could return to the pair "reported three quarters, outcome not held," and then find the then-current material and adjudication basis. The score is then not the only remaining memory, but an entrance back into the judgment.

If only the score can be found but not the original probability, multiple corresponding pairs may also exist. The binary squared error gives the same value for different report-outcome pairs; the score alone cannot recover the full content. Preserving the original inputs remains a necessary relation.

The summary page may be concise, but the detail page should connect the original reports and outcomes. When a computational error is found, the computation can be corrected; when an outcome adjudication is found wrong, an adjudication correction is recorded separately. Both changes affect the score, but neither should be conflated as an improvement of the forecasting model.

A Score Does Not Say Who Caused the Outcome

Tang Ke asked: if the attached-sentence process later brings lower scores, does that show it increased confirmations? Gu Ning reminded her that probability performance and process effect remain two questions. A report fitting the outcomes better does not mean some action changed the outcomes.

A new clue may improve forecasting without participating in producing the target; an action may also influence the target and thereby change the task structure. One must return to the competing explanations and comparison conditions of Chapters Fifteen and Sixteen; scoring cannot substitute for causal judgment.

Nor does scoring decide whether an undertaking is worth accepting. A higher probability may face consequences one cannot bear; a lower probability may correspond to an acceptable small trial. The score evaluates the representation; action still requires targets, constraints, and paths of failure.

This volume therefore places the score on the judgment page. It helps compare degrees of reporting; it does not decide for the station whether to invest equipment, accept the commission, or choose the moment of exit. A clear division of labor lets the second volume take over the practical questions, rather than receive an average treated as an all-purpose conclusion.

After a Poor Score, Ask Which Gap

If some loss is large, first verify the event, the report version, and the outcome adjudication for accuracy, then examine the basis of adoption. An original probability that had support but met an unfavorable realization differs, as a finding, from raw material that was double-counted.

If losses are repeatedly large under some class of conditions, one can consult the calibration page for direction and magnitude and keep the competing explanations. The average raises a question; it does not by itself say which input to change, still less guarantee that every modification lowering the old batch's score will improve the new batch.

Old outcomes can help form candidate rules; new rules still need subsequent checking. Rewriting the original reports to hug the outcomes can lower the back-computed score, but it does not improve the performance of the forecasts actually issued.

The improvement page should therefore write how the model or process will change, not merely that the next score should be lower. The score is one representation of feedback; concrete modification requires returning to the material and the conditions of applicability.

Giving Scoring a Position Where It Can Be Constrained

In RC's internal stance, a score is likewise a representation excerpted for an explicit task. It has exact mathematical relations and a finite scope of use; internal consistency does not obtain truthfulness for external objects, materials, or adjudications.

We may explicitly adopt the squared error, accurately preserve poor results, and still keep a place for what explanation has not yet sufficed. Limitedness permits neither deleting known adverse outcomes nor requiring a complete account of all causes immediately after every score.

If feedback connects the original reports with the actual outcomes, it can promote calibration; if it only rewards good-looking results, rewrites the original numbers, or selects the inputs, it may make the existing explanation more rigid. This distinction is checked against the concrete record; the name "score" does not confer it.

A usable scoring page should state the scale, the direction, the objects, the versions, the denominator, the weights, and the pending scope. When the reader sees a number, the corresponding task can be found, rather than the number standing in for the whole of the recorder's ability.

At 20:00, Saving the Scoring Convention

At 20:00 on Thursday, Gu Ning saved this book's scoring convention: the squared error of the binary affirmative probability, smaller is better, averages computed over an explicit record scope. First reports and update evaluations are kept separate; statements of known outcomes are not mixed into the corresponding ex-ante results; pending items and corrections have continuing positions.

The L ledger and the twenty-object computations are retained as independent methodological examples and do not count as Chengwan statistics. R17 still has no final probability paired with a complete confirmation; the candidate three quarters does not enter the formal results just because the formula is written; the attached-sentence process is still not implemented.

The next chapter continues by examining how adverse material makes judgment contract. One poor outcome, repeated group deviations, and a clear material error are not the same signal; revision must state which layer it targets, not merely display stronger conviction in the face of failure.