Chapter 3: Where Samples Enter From
2026.09.13Eight Cards That Look Very Confident
Having saved the event and time notes for R17, Gu Ning watched Tang Ke take eight display cards from the cabinet. Each had photographs of tables and chairs before and after repair, a task name, and a completion note. She said these commissions had all been finished, and now the other party was asking in a lot of detail, so Gu Ning could use them for reference.
These eight cards and the numbers below are settings newly added to the Chengwan thought experiment in this chapter, not real business statistics. Gu Ning is still at the Thursday-morning stage of organizing information, and R17 has not received complete confirmation. The appearance of historical material does not preemptively generate an outcome for the current request.
The cards do show that the repair station completed work, but they do not directly answer whether these requests received first-version confirmation within the specified window. A piece of work may have started only after several revisions, and may have been confirmed at a later time. Final completion and confirmation received by the next day are not the same outcome.
What Gu Ning seeks is material that can help the current judgment, not a denial of past achievements on the repair station's behalf. She first asks Tang Ke why there are only these eight cards, where the unfinished requests are, and whether the cards cover all registrations. The number eight needs a provenance; it cannot become a complete denominator merely by being neatly displayed.
Ye Cheng, who has not appeared before, joins the organizing. He too is a fictional member of this book, tasked with helping locate old request records and not written as possessing all past facts. Ye Cheng finds the registry: before R17, there were requests R01 through R16, sixteen in all. The display cards belong to eight of these, not to eight other new objects.
First, Look at How Material Reaches the Eye
The display cards were made after repairs were completed. They serve to show what has been done, not to preserve how all inquiries ended. Requests before completion do not vanish from the registry for lack of a card, nor is the card obliged to bear the entire use of prediction material alone.
Gu Ning therefore faces a material entry point that already has a purpose. She sees completed results first because the files in the cabinet are stored for displaying results. This entry point provides real information while letting only certain objects appear easily. Purpose and coverage must be identified together.
If the display cards were taken directly as all requests, Gu Ning would infer entry conditions backward from existing results. Every one she sees is complete, perhaps only because unfinished objects never enter this stack. That the card contents are true does not prove that once a request appears it is usually completed.
This chapter will not hurry to compute a magnitude of influence for this selection. To state the influence, one needs to know how the entry point actually works, which objects did not enter, and what relation they bear to the target event. Asking how something came to be seen is more useful than first arguing over whether eight is enough.
An entry point can also be a set of examples recommended by someone, a table of search results, or experiences one easily recalls. We do not here invoke these names to assert claims about real platforms or memory research; we compare the same logic: objects already in hand are not necessarily all the objects we wish to observe.
The Registry Range Is Not the Whole World
The registry lets Gu Ning find sixteen historical requests, but sixteen is not all the needs Chengwan might have. People who never made an inquiry are not in it, nor perhaps are exchanges that merely asked about prices without being registered as requests. Gu Ning must state which entry point this material covers.
She provisionally takes the registry as the search range for this historical material: requests that already obtained identifiers. This range makes the search executable, yet it does not prove that all table-and-chair needs have the same record, nor that future entrants into the registry are the same as non-entrants.
Whether the range suits R17 still needs later comparison. R17 is already registered and shares at least one entry point with these sixteen; whether its task, communication conditions, and confirmation window are similar has not been checked. Sharing a registry is not sharing all generating conditions.
If what Gu Ning truly wants to judge is whether any stranger's inquiry will be confirmed, the registry may omit a great many not-yet-registered inquiries. This is a mismatch between the target object and the material range, and it cannot be solved by writing sixteen as a larger number.
A sample provenance note should first be able to state clearly: where this set of objects was obtained, what class of objects had a chance to enter, and what class was not in the search range at all. Without this statement, so-called past experience easily expands a limited entry point into the whole world.
What Exactly Does One Row of a Record Represent?
Ye Cheng finds three sheets for one request: the initial inquiry, a scope modification, and the final confirmation. If each sheet counted as a request, one past stretch of communication would become three outcomes. They are material of the same object at different stages, not three independent commissions.
Conversely, the same client may also propose two distinct tasks. If these are merged only by personal name, two requests are again compressed into one row. The statistical unit must be chosen according to the question; it cannot be decided by sheet count or name count alone.
Here we start from the request identifier and first return the material to R01 through R16. Under each request, multiple versions, multiple time points, and different outcomes may be kept, but request count does not increase merely because material grows. This connects with the previous chapter's point that many predictions are not many outcomes.
The current event concerns the first version receiving confirmation. If a historical request did not preserve a so-called first version, it may still retain other material, yet cannot for now be filled into the same outcome column. Missing a corresponding version is a material state, not something one more completion photo can supply.
The statistical unit is therefore not a natural cell pre-existing in the table. It is identified jointly by the question, the object, and the relations among the material. A classification may be provisionally usable and still must state how it merges and how it splits, so the organizer does not silently change the number of objects for the judgment.
Why Duplicated Material Looks Like More Evidence
Tang Ke's eight cards and the original communication records Ye Cheng later finds may cover the same requests. If the two stacks are simply added together, one obtains more entries than there are actual requests. Different file names do not mean different events behind them.
As a setting for this chapter, of the sixteen registered requests, twelve have corresponding original communication material, and four have their raw material temporarily missing. The eight display cards overlap with these twelve raw materials; they do not add up to twenty historical requests.
Original communication records and display cards can complement each other. The original records say when what was discussed; the cards say what was later completed. They provide observations at different stages, and do not thereby become two independent verifications of the same confirmation outcome.
Gu Ning can annotate the provenance of the material under each request, keeping photographs, summaries, and original messages separately. This neither discards useful content nor lets a request with more files obtain more outcome counts. Material richness and object count are different and must both be preserved.
Independent sources and shared errors will be addressed further in Part II. This chapter first settles an earlier step: duplicated files must be returned to the same object. If this step is not completed, then no matter how one compares probabilities afterward, one may keep reasoning on an already-duplicated quantity.
Twelve Can Be Found; the Other Four Must Not Be Forgotten
After the search for raw material, Gu Ning has twelve requests that can continue to be checked, while the other four retain only registration information. She may study the twelve first, but she cannot write the four as nonexistent, nor say that all past material is now complete.
The cause of the four's missing status is temporarily unknown. Perhaps nothing was preserved at the time, perhaps they are stored elsewhere, perhaps the search range was insufficient. Different causes may have different effects on comparison, but Gu Ning has not obtained the material and cannot invent an explanation to make the table complete.
If the missing material occurs only among a certain class of unfavorable requests, the existing material will lean toward the other class; if the missing status bears no such relation to outcomes, the effect may differ. This chapter only raises the conditions to be examined; it does not declare which case Chengwan's four belong to.
A clear state can be preserved: registered, raw material not obtained, outcome not yet checkable under the current definition. This differs from having observed non-confirmation. Filling the missing directly as "no" changes the outcome count; deleting it also changes the range of observable objects.
Preserving a gap is not a demand for infinite waiting. Gu Ning may, after stating the range, temporarily use the checkable material while listing the objects not yet covered and the means of later searching. Whether a limited judgment may be adopted needs further analysis by use and by the importance of the gap; a single missing item does not automatically cancel all material.
Same Field, Possibly Different Meaning
The old registry has a column called "confirmed." Ye Cheng finds that some records say full acceptance was received, some say the other party was willing to keep talking, and some merely mark that the contact person verified the table-and-chair inventory. One field preserved different actions and cannot be directly converted into one kind of outcome.
Gu Ning has no reason to first assume the old organizer was careless. The registry may originally have served to arrange communication, and "confirmed" merely meant that a certain step could continue. That the prediction question later needs a narrower meaning is a change of use, not necessarily proof that the original table never had value.
She should return to the raw material and identify what each entry actually confirmed. The parts corresponding to this chapter's conditions are preserved; the parts that do not correspond have their difference stated; for those without raw material, no judgment is forced from the field name. Conceptual correspondence requires an actual bridge.
Likewise, the old table's "next day" may count one day from receipt of the inquiry, or one day from the proposal of a final scope. The previous chapter's temporal distinctions enter the sample here: if the window's starting point does not correspond, outcomes all called "confirmation the next day" may not treat the same question.
The role of sample cleaning is not to uniformly replace old words with new ones, but to make object, field, and time point genuinely connect. Text substitution alone may make the surface format consistent while the material meaning remains inconsistent, making hidden differences harder to discover.
Objects Easy to Locate May Fill the Table First
Gu Ning finds that requests on the display cards are the easiest to locate material for, because the tasks were later completed and the relevant files were stored together. Other requests may require searching one by one. If she organizes only the easiest batch, the work finishes quickly, yet the coverage range is decided by search convenience.
Convenience has real value. With limited resources, it is impossible to trace all material indefinitely. The question is whether she knows that convenience is selecting objects and states the range when using the conclusion, rather than converting organizational ease into proof that these objects are most representative.
One can list all registration identifiers first, then record each one's acquisition state. The search order can then be flexible, and unobtained objects will not vanish from the table. Forming the range first and then recording progress against it maintains reviewability better than continually adding easily found success examples.
If certain requests need to be searched first, the reason can be stated: closer to the current task, time material involving the boundary, or missing status that might change the interpretation. Priority reasons should connect to the question, not merely to whether some outcome is reassuring.
Gu Ning does not settle all organizing costs in this chapter. What she obtains is a visible relation: how material arrives in sequence may change her belief about how the past occurred. Keeping this relation lets later judgments know which parts she has not yet observed.
More Material Does Not Necessarily Increase Relevant Range
Tang Ke suggests finding more completion photos; Gu Ning says that is fine, but whether the photos increase material for the current event needs specific judgment. Ten photos of the same request may make the completion state clearer without adding another request or supplying a confirmation arrival time.
Likewise, more different requests can expand object count, yet if all lack the first-version scope and window, they still cannot be used to check this outcome directly. Quantity and degree of correspondence are different dimensions; a large stack of material should not substitute for a gap in a specific field.
For prediction, useful material and plentiful material are not interchangeable. A single original message may address a key time question, while a whole stack of later results addresses only production completion. Each has its use, and each must be separately stated as to which step it supports.
This also does not require collecting only one strict format. Different material can enter the judgment through explicit interpretation. If old notes and original messages jointly identify the acceptance object and arrival range, they may suffice to support a limited adjudication. Format uniformity is not the sole source of reliability.
This chapter does not propose a formula by which more samples necessarily improve accuracy. How material should be increased depends on the target object, the acquisition mechanism, and the error structure. Without first identifying these conditions, persistently increasing quantity may only enlarge one kind of mismatch.
Seeing All Registrations Is Still Not Seeing All Conditions
Even if the raw material for all sixteen is eventually obtained, Gu Ning still does not know whether they suit comparison with R17. Table-and-chair scope differs, communication stage differs, and receipt arrangements may differ. Complete material and transferable judgment must be separated.
Complete acquisition can protect objects from omission, but it cannot thereby guarantee that the same probability applies to every request. Chapter 4 will ask which class of past objects should serve as the current reference, how to choose the denominator, and why certain conditions need to be grouped.
This chapter therefore does not compute the eight completion cards, twelve locatable materials, or sixteen registrations into a confirmation probability for R17. They respectively describe display, acquisition, and registration ranges, and have not provided a complete outcome count under a single event definition.
This is not to delay calculation. Calculation requires knowing what the numerator and denominator respectively represent. If a ratio is given now and only later explained to be the share of completion cards among registrations, the value Gu Ning obtained has answered another question.
Reasoning can proceed in order: first settle the question, then locate the material entry point, identify the unit and corresponding fields, preserve the gaps, and then compare which set of material can support the target judgment. Order gives each value a provenance, rather than using the outcome to search backward for a suitable explanation.
Does Organizing Material Change the Original History?
Ye Cheng's merging of three sheets under one request is an act of organizing; it does not mean the earlier three exchanges never happened. He is merely distinguishing the event unit from material entries. Keeping the original sheets lets one later check why this merge was adopted.
If another organizer proposes that these two sheets actually belong to different tasks, this should be checked against object and version. Classification can be revised; raw material should not be discarded merely because a first organizing has been completed. Processual completeness also requires a return path at the material level.
Revision also cannot change outcomes at will. Filling missing times with a default "next day" so that all twelve become usable, or rewriting "kept talking" as "explicit acceptance," is not obtaining new material. It turns the unknown into unproven certainty through a table operation.
Theory reduction in RC means that limited representation supports limited observation. The registry range, request unit, and material state in this chapter are all organizations of a complex process. They have effect, but completion of the organizing cannot substitute for all the conditions of the original history.
Hence Gu Ning must preserve the reasons for classification and the important revisions, rather than pursuing a single indisputable perfect table. Allowing material to return, and allowing others to point out an inappropriate unit, is what makes later statistics rest on revisable objects.
How New Records Reduce the Next Gap
Gu Ning can start from R17 and preserve material according to the already-formed event card and time note. When people look back in the future, they will at least know what was asked then, what the confirmation condition was, and when the message was obtained. New records can reduce certain gaps, but cannot backfill what was not preserved for old requests at the time.
Future preservation rules also need an entry range. If predictions are established only when Gu Ning is very confident, what is left later will be the objects she chose to predict, not all registered requests. The chapters on selection and calibration will continue this; here we first write down from which step the record begins.
At registration one can mark whether a prediction was established, why it was not, and whether outcome material continues to be preserved. Absence of a prediction does not mean absence of a request, nor does it mean failure. A statement of range helps later comparisons between predicted objects and other objects.
Gu Ning cannot, in order to obtain a more complete history, require the client to provide all private information unconditionally. Details irrelevant to the event conditions need not enter; if necessary material requires limited preservation, its use and checkable range can be stated. Completeness is not unlimited collection.
After the new rules run, gaps, duplications, and misreadings may still occur. They are not a guarantee of future material quality, but an arrangement that makes problems surface earlier. Whether gaps are actually reduced requires material from subsequent records; it cannot be proven at once by a format that looks clear.
Can Others Check This Organizing?
Gu Ning hands the range table to Tang Ke, not to demand she search all files again, but to ask her to check several relations that would change units and outcomes. Whether all eight display cards correspond to registration identifiers, whether original communication material was assigned to the correct request, and whether the four missing items were merely not found this time can each be raised as a question.
First set a comparison: a certain modification note concerns the same batch of tables and chairs; Ye Cheng treats it as a separate request, while Tang Ke holds it to be only a new version of an old task. The two need not first argue over who knows the past better; they can return to the object in the note, the original identifier, and the communication provenance. After the dispute is handled, the reason for merging or preserving should also be recorded.
If the existing material still cannot distinguish the two, this item can be marked as a unit pending check, rather than entering the certain object count without explanation. Provisionally keeping two interpretations differs from counting two requests at once. The former admits classification is incomplete; the latter has already turned uncertainty into an added count.
Checking may also reveal that some original message is only a copy of a later summary. Different file dates do not necessarily mean a new observation was produced. One should preserve which raw material it points to, so the provenance table does not mistakenly write the copying process as a newly added fact. Clear provenance also protects reasonable reliance on the original record from expanding without limit.
Such review does not require everyone to redo the entire organizing each time. One can check around key merges, gaps, and field meanings, and only then extend the relevant range once a rule problem affecting other entries is found. Check costs should connect to the judgment they might change, so that minor problems are not crushed by infinite proof.
Finally, preserve the organizing version. If a missing piece of material is later found, update the acquisition state and state the new provenance; if a merge error is later found, revise the correspondence. Why the old count changed then becomes visible, without the previous table having never existed. The provenance note becomes a process open to continued checking, rather than another final label one can only accept.
Attaching a Provenance Note to the Sample
At the end of organizing, Gu Ning saved a provenance note, not a confirmation probability. The historical search range is the sixteen registered requests R01 through R16; the eight completion display cards are objects within it; twelve have locatable original communication material, four are temporarily missing; duplicated files are returned to requests; old field meanings continue to be checked.
She separately lists the judgments still incomplete: among the twelve, which preserved a corresponding task version, which time ranges are checkable, and which outcomes satisfy the conditions connecting to R17. The counts are now identifiable; the comparative use still needs the next step.
The value of the provenance note lies in letting someone who later sees a ratio know from which objects it was produced. It need not disclose all raw material, but should be able to point out the search range, unit, overlap, and gaps. The number no longer seems like an answer that naturally surfaced from the past.
Readers can check a set of material they plan to use for a judgment: why this set is seen first, which objects did not enter, what each row represents, how duplicates are handled, whether gaps are still preserved, and whether old fields correspond to the current question. If these relations are not yet clear, do not let material quantity serve as proof of the judgment's reliability.
The judgment obtained in this chapter is that a sample is not a heap of the examples at hand. It requires an entry point, a unit, provenance, and acquisition state to jointly form an observable range. Material can be true while its selection is limited; records can be many while objects repeat; material can be complete while the use has not yet corresponded.
The Chengwan main line is still organizing the judgment conditions for R17, without preemptively confirming the current outcome through historical display. The next chapter continues from this range note and checks comparison classes and base rates: which past requests can truly enter the denominator, and by what conditions outcomes should be counted. Once the provenance is clear, the ratio begins to have a question it can return to.