FORM NOT VOID, MIND NO CORE

Chapter 5: Local Failure and Systemic Failure

2026.09.07

On the twenty-fourth day, the deviation begins to move beyond the site where it first occurred. The clock cannot declare for us that "collapse has arrived," yet it forces whoever judges to say how far the propagation has traveled. One service interruption does not necessarily show that the whole system has collapsed; many locations reporting normal operation at the same moment does not necessarily show that the system can still honor its promises. The distinction between the local and the systemic is not decided automatically by the number of faults; it depends on how the various locations are connected, what substitutes remain, how quickly losses cross boundaries, and whether information can return to the decider before the problem grows.

Imagine a fictional public library-lending network. Three branch libraries share a catalog, transfer vehicles, and one central repository that holds part of the collection, while each branch also keeps books on site. One day, a reader at the east branch cannot pick up a reserved book; the staff see the system displaying "available for pickup," yet the book is nowhere to be found in the stacks. The west and south branches are still lending normally at that moment. This may be nothing more than a misplaced book; it may be a transfer status that was not updated; or it may expose a systematic rupture between the shared catalog and the actual inventory. The lending network is only a map for propagating pressure; it is not to be used to judge real libraries, nor does it offer an information-system or logistics solution. All nodes, persons, and faults are a thought experiment, meant to establish criteria for propagation rather than to simulate a destructive path anyone could copy.

The previous chapter showed that divergence can supply a corrective signal, but a report's inconsistency with the model is not yet systemic failure. This chapter pursues the question further: under what conditions does a local difference stay local, and under what conditions does it pass through multiple nodes; and how are the two claims "just an isolated case" and "a network-wide crisis" used by positions of power.

Failure Is First Relative to a Promise

A reader at the east branch not receiving a book is one adverse outcome. To judge where the failure lies, we must first state what the network has promised. If the page says only that the central record shows a lendable copy, then an on-site absence means the record needs correction; if the page explicitly tells the reader to pick the book up at the east branch, then the service promise also includes the transfer and handover having been completed.

The system's boundary changes with the question. When judging "is this book on the shelf," the east branch's stacks are the primary object; when judging "can the reservation service deliver a lendable copy to a designated location," the catalog, the central repository, transport, and notification are all inside the boundary. We cannot choose the smallest boundary first and then use causes outside it to exempt the whole from responsibility. Nor should the boundary expand without limit. Vehicles being affected by weather does not mean the weather itself is a controllable component of the lending system; but how the system represents delay, whether it retains substitutes, and how it notifies readers still belong to design and to promise. An external cause can explain the triggering event; it cannot automatically answer for internal preparation.

Severity of Outcome and Range of Propagation Are Two Dimensions

A crucial reference material not obtained on time may gravely affect a particular reader, and still be a local event. Ten ordinary books briefly delayed across different branches involve more locations, and may still have no common cause. Severity, number, and systematicity must be recorded separately.

Calling a severe case a system failure is sometimes a way to get the problem resourced; but if the concept is used only by emotional intensity, it becomes hard to judge at which level the repair should land. Acknowledging individual loss does not require exaggerating the range of propagation. Local does not mean minor, still less does it mean negligible.

Conversely, a systemic problem may at first produce only one visible consequence. If all three branches depend on the same status field, and only the east branch happened to trigger the relevant workflow that day, the east branch is the first site where the problem becomes visible, not the only affected structure. Waiting for more readers to fail before acknowledging a common cause turns a numerical threshold into a causal judgment. We therefore need at least two sentences: what consequence occurred, and which shared relations could allow it to occur again. The former protects those already harmed from being diluted by statistics; the latter prevents a problem that has not yet spread from being treated as already over.

A Node Is Not a Box on the Organization Chart

A branch can be treated as a node; so can the catalog service, the transfer stage, and the notification stage. A node is not a natural unit but a boundary drawn temporarily in order to analyze certain inputs, outputs, and responsibilities. If the division is too coarse, internal failures become invisible; if it is too fine, every step can push the problem onto the adjacent position.

The staff at the east branch say "the central system displays it wrong"; the central staff say "the librarians never confirmed receipt into the repository" — both sides may be pointing at a real link. If the analysis searches only for the single bad node, a consequence formed across many steps is dismantled into local pieces, none of which is responsible.

The division of nodes must preserve relations: who supplies the status, who acts on it, who can correct it, who makes the promise to the reader. Organizational affiliation is not the whole of it. Outsourced transport may sit outside the institution yet inside the service chain; a person inside the library, though belonging to the same organization, may not control the relevant decision. This also constrains this chapter's systemic language. Calling something "systemic" cannot be used to cancel the responsibility of specific acts; calling something "operator error" cannot cancel the conditions that let the same error propagate repeatedly. The two explanations can sit at different levels and hold simultaneously.

Nodes can also be divided from the user's side. Inside the network, the catalog, transport, and the branches count as three segments; the reader faces only one service, "reserve, then pick up." That some internal segments still function normally cannot directly entail that the external promise is normal; likewise, one external failure does not prove that every internal segment is broken. The analysis must build a correspondence between the two boundaries rather than choose whichever side favors it. If the organization keeps statistics only by department, cross-departmental gaps easily end up with no owner. Every node meets its own local metric, yet no one confirms the end-to-end result. Systemic failure is sometimes not a box ceasing to work, but a promise between boxes that no position sees in full.

Shared Dependency Determines How Far a Fault Can Travel

The west and south branches have different buildings, staff, and on-site collections, and look like three independent paths. But the reservation status comes from the same catalog, and part of the transfers rely on the same vehicles. Spatial dispersion is not dependency dispersion. Counting only the number of branches mistakes a shared entry point for real redundancy. If the east branch's problem originates in one misplaced book at that branch, the other branches need not be affected; if it originates in the shared catalog writing "loaded onto the vehicle" prematurely as "available for pickup," the other branches may repeat the failure under the same conditions. A common cause does not require all nodes to be interrupted simultaneously; it only requires that a single error of judgment be inheritable at many places.

RC's stability continuation lists path diversity and continuous correction as conditions of continuance. This chapter borrows that to pose an analytical question; it does not treat a philosophical concept as a law of reliability engineering. Real propagation still has to be accounted for by actual dependencies, timings, and records.

Apparent Substitution May Still Pass Through the Same Entry Point

The east branch can suggest that the reader go to the west branch to pick up the book, apparently offering a substitute. If the west branch must also rely on the same erroneous status to confirm the book, the change of location has not bypassed the fault. Whether a substitute path is real depends on whether it can complete the objective independently under the current failure conditions. Full independence is usually expensive. Each branch building its own catalog and keeping its own vehicles is not necessarily more reliable, and data conflicts and maintenance burdens also increase. A system can share resources, but it must know which class of common failure the sharing introduces, and retain for important promises commensurate verification or a manual fallback path.

Substitution may also be reachable only for some people. The west branch truly has the book, but the reader needs longer travel time; for those who can move flexibly it is a usable path, for those constrained by time, body, or caregiving arrangements it may not be. Writing nominal substitution down as service restored moves the switching cost out of the system's picture. Substitute capacity therefore includes resources, information, and the conditions of the subject. It is not "there is another button," but whether the objective can be completed, within an acceptable time limit and cost, by another chain of relations. What counts as acceptable still contains a value judgment that must be stated openly, not unilaterally announced by the system.

One possible dark corollary would rewrite path diversity as dispersed responsibility: since readers can in theory go elsewhere, each node need only offer options and need not answer for whether the options are reachable. The system's normative wish concerning available margin is inverted into institutional exoneration, while the switching cost is borne by those with the least margin.

Recovery Time Changes the Nature of the Same Fault

If the east branch discovers the status error within minutes, contacts the reader, and reschedules, the fault may remain a repairable deviation. If the erroneous status persists for days while new reservations keep entering, a record problem originally local accumulates into promise distortion across many nodes. Time does not merely measure inconvenience; it also determines propagation. When recovery is slower than the entry of new tasks, unfinished items increase; when feedback is slower than the updating of decisions, old judgments continue to be copied. This chapter establishes only the relative relations; specific speeds are left for the next chapter.

The same recovery time limit means different things for different promises. Leisure borrowing can wait; a material with a hard deadline may not. A system cannot publish only an average recovery time; it must also identify which tasks undergo irreversible loss while waiting. Averages are useful for the aggregate; they cannot speak for every category of consequence.

Fast recovery also does not prove the cause has been corrected. A staff member finds another copy ad hoc, the reader receives service, and the erroneous status may remain. Recovery outcomes and repair mechanisms should be recorded separately; otherwise successful remediation keeps the common dependency invisible.

Backlog also forms during recovery. Old tasks are not yet closed, new reservations keep entering, and staff spend their time explaining and re-entering records, so existing service capacity declines. The initial fault may be small, yet the attendant coordination labor occupies other nodes. Propagation travels not only through technical dependency but through shared personnel, attention, and decision time.

Backlog does not necessarily mean the system is about to collapse. The organization can temporarily lower other activities, add explanation, or adjust promises until tasks restabilize. The point of judgment is whether the recovery arrangement genuinely reduces unfinished items, or merely renames them, defers them, or moves them outside the statistical boundary. If the center computes only current success rates, readers who have already given up, gone elsewhere, or stopped asking may disappear from the denominator. The metric looks recovered; the loss is completed by exit. A counterfactual question must be kept here: absent additional individual cost, was the original promise actually honored?

A Fault Tree Is a Set of Questions, Not a Determined Answer

Faced with the east branch incident, we can draw several possible paths: the book was never entered into the repository; the status was updated prematurely; the transfer was missed; the notification was wrong; the on-site search failed. The so-called fault tree is here only a checklist of questions, helping to distinguish which conditions are singly sufficient to produce the consequence and which must occur together. This chapter offers no real technical design.

Laying out the paths prevents the first explanation from prematurely becoming the only explanation. A system displaying wrong information need not be caused by software; a missing book on site need not be staff error. Each path requires its own material, and the unknown should remain unknown. Tree diagrams also mislead. Real feedback may form loops: the erroneous status leads staff to stop searching; the stopping of search removes correcting information; the catalog then continues to display the original status. If we draw only one-way causes, the reverse influence of behavior on information disappears.

Simultaneous Occurrence Is Not a Common Cause

On the same day, the east branch's reserved book is missing, the west branch's entrance gate briefly fails, and the south branch closes part of its area for an on-site event. Three faults occur simultaneously, yet they may be mutually independent. Treating temporal proximity as systemic collapse creates a common narrative that exceeds the material.

A common cause may also lack synchronized consequences. A shared catalog error first affects the east branch with heavy reservation volume, and only days later affects the south branch. Demanding same-moment occurrence would instead miss propagation by dependency. Correlation has to be explained through concrete shared relations, not by intuition about counts or timing.

Intentional action still less can be inferred directly from synchrony. Multiple nodes failing at once may come from a common update, the same external condition, or similar but independent pressures. To conclude that someone coordinated the manufacture, additional evidence is required. What is structural and what is a judgment of intent must be kept apart. Conversely, "there was no unified scheme" also cannot terminate systemic responsibility. If a shared rule lets multiple positions stably repeat the same error, then even if every individual executes it in good faith, the shared condition still needs to be changed. Institutional consequence does not depend for its existence on finding a central intention.

Feedback Coupling Determines Whether an Error Can Reinforce Itself

The east branch cannot find the book, yet because the system displays it as available it is required to keep searching; each failure to find it is logged as a local execution anomaly, while what the center sees remains an overall normal catalog. Local reports do not modify the shared status, so the error maintains itself through the feedback structure. If evaluation in turn rewards "on-time completion rates," librarians may temporarily close anomaly records to avoid their branch's numbers falling. The short-term metric improves, and the center receives fewer problems. Evaluation is not an external add-on; it changes whether feedback is willing to travel.

At this point an initial deviation expands through two loops: status guides action, and the record of action in turn validates the status; evaluation decides which records are worth reporting, and the volume of reports is then used to judge stability. Only when common dependency combines with feedback coupling does a local event acquire systemic propagating power.

Coupling is not the weaker the better. A shared status can let other branches quickly avoid the same problem, and centralized feedback can also speed repair. What matters is whether information can contain counterexamples, not that all nodes each run in silence.

RC's cycle of expectation, reality, and re-expectation provides the conceptual entry. If deviations of expectation can modify the shared judgment, coupling supports learning; if deviations are only reinterpreted as local execution problems, coupling enlarges rigidity. This remains a conditional inference; it cannot substitute for real material in proving which kind a given system belongs to.

Feedback may also be fast in only one direction. A new rule from the center reaches the branches immediately, while a branch's anomalies must pass through layers of aggregation before they can return. On the surface all nodes are tightly connected; in fact the coupling has a directional disparity. When the speed of command propagation and the speed of error return are asymmetric, the system copies expectations quickly and absorbs counterexamples slowly. This structure can be exploited deliberately, or it may simply be the accumulation of the division of labor. If the center compresses reporting categories in the name of uniformity, it can maintain a tidy view; if branches hide problems to protect themselves, they also distort the common judgment. We cannot presuppose a sole manipulator; we need to examine who can change the channel, who benefits from the delay, and who bears the late correction.

Repairing feedback does not mean letting all raw information flood the center. Excessive reporting consumes processing capacity, and the important signals become harder to make out. Reasonable aggregation should preserve the object, scope, and source of anomalies, and allow compressed information to escalate when consequences change.

How "Merely Local" Misdraws the System Boundary

The central staff say the east branch incident belongs to local management, on the grounds that the other branches are still normal. This may be a correct first assessment, or it may substitute the distribution of outcomes for dependency analysis. If the east branch has no authority to modify the shared status, leaving responsibility local requires the position with the least authority to repair a common condition.

The local label also has resource consequences. A problem classified as an isolated case receives less verification and can only be resubmitted repeatedly by those affected; the paucity of reports then proves the problem rare. Classification, resources, and visibility reinforce one another, keeping early systemic signals in a local appearance. This does not mean every case must be escalated to a network-wide incident. A comprehensive response has costs, and may interrupt services that still work. Graded handling is reasonable, provided the escalation criteria look at shared dependency, substitution, and recovery — not only at how many nodes have already failed.

Localization Can Conceal Common Benefit and Common Design

In normal times the network saves cost and extends transfers thanks to the unified catalog, and the benefit is described as overall capability; when something goes wrong, the consequences are handed back to a single branch or to the reader. The shared structure exists only in success and vanishes in failure. This asymmetric narrative lets the center retain the gains of scale while pushing the burden of repair downward. The dark usage need not deliberately manufacture faults. It can exploit the power of classification, slicing the losses produced by the same mechanism into many isolated cases, so that each affected person lacks the material to prove systematicity; the center holds the aggregated data, yet responds only case by case. When harmed subjects cannot compare horizontally, the system declines common responsibility for lack of common evidence.

A stronger inference may also appear: since local shocks can sometimes strengthen overall resilience, edge nodes can be made to bear trial and error repeatedly, so that the center learns from their failures. The original theory's acknowledgment of error and fault tolerance is rewritten as a license to distribute loss. Whether a shock is acceptable must be checked against who chooses, who is harmed, whether refusal is possible, and whether the benefit returns — not only against what the whole later learned.

Some experiments genuinely must be conducted locally, because the reversible range is smaller. Critique cannot write every limited trial as sacrifice. The difference lies in whether the risk is stated in advance, whether the bearer has a real choice, whether stopping conditions exist, and whether, after failure, repair is provided rather than data alone. "The system must endure local attrition" may also refer to ordinary wear, not to letting a particular group repeatedly lose basic services. When systemic language magnifies the object, individual consequences easily become abstract costs. A responsible analysis must keep aggregate learning and concrete loss in the record together.

Crisis-Declaration Can Also Expand Central Power

Another dark path declares any local anomaly a system crisis, and on that ground suspends normal appeals, concentrates information, and expands provisional authority. Before propagation is proven, the crisis narrative has already changed who can decide. Critique of localization does not automatically support total centralization.

Some risks do call for limited protection before the evidence is complete, because waiting may cause losses that cannot be withdrawn. The key is to separate action from conviction: one can say "a common cause is not yet confirmed, but we are suspending the relevant promise and verifying," without writing the preventive measure as proof that systemic collapse has been established.

Provisional centralization also needs end conditions. If every new anomaly extends the state of crisis, the center can derive authority from uncertainty indefinitely. The record of action should preserve the material as it stood, so that afterwards one can check whether the response was proportionate, not justify every measure merely by the absence of enlarged loss. Opponents may likewise exaggerate systematicity to gain rhetorical advantage. If every independent fault is spliced into the same story of collapse, the explanation absorbs all events without undertaking to change conditions. Whoever criticizes the system must likewise state the connecting paths and the counterexamples.

Judge Where Failure Spreads with a Multi-Scale Record

The east branch incident can carry three kinds of record at once: the reader level records the unfulfilled reservation and its actual consequences; the node level records the searching, the handover, and the local status; the network level records whether the shared catalog, the transfers, and the other branches share the same conditions. The three kinds of record answer different questions of responsibility.

Multi-scale does not mean uploading data without limit. Readers need not disclose information irrelevant to the purpose of borrowing, and the center need not take over every on-site operation. The minimal necessary sharing should center on the judgment of propagation: which statuses are jointly used, whether the anomaly can be replicated, whether existing substitutes are real. If verification finds that the book was simply misplaced at the east branch, local repair plus redress to the reader suffices; the shared system need not be wholly rebuilt over one error. If it finds that the status-update rule could let many branches notify prematurely, the repair must return to the common link, while the concrete consequences at the east branch are retained. Systemic responsibility does not absorb individual redress.

The record should also tolerate what cannot yet be classified. Early in the investigation one can write "currently found only at the east branch; not yet confirmed whether the cause is shared," along with the basis for the next update. The unknown is not an evasion of responsibility, so long as it has an object of verification, a responsible position, and a time boundary. Writing "local" too soon suppresses the investigation; writing "systemic" too soon lets the conclusion precede the material.

Near-failures ultimately corrected by hand also carry informational value. Such an event should not be counted as a loss equivalent to one that occurred, yet it can show which promise depends on improvised backfill. If every act of backfill rewrites the record as success, the system mistakes the maintainers' extra labor for the reliability of the design itself. The one who backfills may also exaggerate their own indispensability, demanding more authority on the grounds of a near-failure. Seeing invisible labor does not exempt review. We need to distinguish their actual repair, the upstream conditions that made it fragile, and whether they control the knowledge and the handover.

Four Criteria of Propagation

The first is shared dependency: whether different nodes use the same resource, information, or rule, and whether a local error can be inherited through it. Sharing itself has value; the judgment asks only that the common failure surface be seen.

The second is substitute capacity: whether other paths bypass the current fault and are genuinely reachable in resources, time, and the conditions of the subject. Nominal options cannot automatically count as recovery. The third is recovery time: whether repair speed suffices relative to the entry of new tasks and the accumulation of consequences. A deviation that can be corrected quickly and an error that keeps copying belong to different risks.

The fourth is feedback coupling: whether local counterexamples can change the shared judgment, or are only assessed as local execution problems. The more centralized the feedback, the faster correction may be — and the stronger the spread when it closes. The four are not a system score that can be summed. They help pose the next questions; real judgment still requires domain material. Trying to compress them into a single index would recreate the very metric blind spot the previous chapter criticized.

The Level of Intervention Should Follow What Can Be Changed

Whoever can change the shared catalog should receive the relevant counterexamples; whoever is responsible for local handover should account for the on-site process; whoever promises the reader should ensure that redress has an entrance. Responsibility need not be concentrated in one subject, nor can it evaporate through the division of labor. Repair must also verify consequences. Modifying a status field does not mean the reader has obtained a substitute; apologizing to the reader does not mean the common error will not recur. Mechanism repair and concrete redress are each completed separately, so that neither concludes in place of the other.

The system may in the end also decide to accept some class of low-consequence local error, because complete elimination costs too much. This choice must state frequency, bearers, substitutes, and avenues of appeal; it cannot replace the distributive judgment with "every system has attrition." Inevitability still requires evidence, and does not automatically equal fairness.

The central judgment this chapter reaches is this: systemic failure is not the simple sum of local failures, but the loss, through shared dependency, limited substitution, recovery time limits, and feedback coupling, of the capacity to keep honoring a promise sustainably. A consequence can be severe without propagating, and an early isolated case can also expose a common structure. Judging between the local and the systemic is therefore not choosing a louder name for the problem, but deciding to which level the material must travel, who has the capacity to correct it, and whether the loss is still spreading. Refusing the exaggeration of crisis matters as much as refusing systemic responsibility; both must let the connecting relations submit to verification.

The next chapter will develop the time variable on its own. With the propagation path unchanged, a different speed of change makes the windows of observation, negotiation, and recovery wholly different. This chapter keeps only the necessary interface: time limits must be judged relative to promises and feedback, not pre-decided by the value preference for "fast" or "slow."