Deep in the cultural DNA of humanity, "failure" is a deeply stigmatized word. From our first day at school, we are admonished by red crosses and glaring scores: avoid mistakes, pursue perfection. In the workplace, this fear is amplified. A failed project can mean squandered budgets, dashed hopes of promotion, even a stain on one's career. We build complex processes, rigorous approvals, and layers of oversight, all with a seemingly impeccable goal: to eliminate failure entirely.
Yet an organization that never makes mistakes is often an organization that is dying. Because in its pursuit of 100 percent certainty, it inevitably abandons the exploration of 100 percent possibility. It is like a child overprotected in a sterile environment — safe from all known germs, but deprived of the chance to build its own immune system. When an unknown, novel virus arrives, the outcome is catastrophic.
In the preceding chapters, we explored "optionality" and "available margin" — the "provisions" and "ammunition" for organizations in an uncertain world. In this chapter, we will touch on a deeper, more challenging issue of culture and organizational design: how to properly "spend" these reserves on exploration. And exploration necessarily comes with failure.
The central argument of this chapter is: What we should fear is not failure itself, but "valueless failure." Our goal is not to build a "zero-failure" organization, but to consciously and systematically design an "antifragile" organization that can draw strength from "valuable failures" and continuously evolve.
We will first deconstruct the widely circulated yet profoundly misleading proverb — "failure is the mother of success" — and establish two crucial prerequisites for this "mother." Next, we will delve into the three design principles for building an "antifragile" organization: firewalls, modularity, and rapid iteration. Then, through a deep analysis of Netflix's astonishing practice of "Chaos Engineering," we will witness how proactively creating disorder forges extreme systemic resilience. Finally, we will provide an essential practical tool — how to conduct a proper "failure retrospective" — to help your organization shift from a "blame culture" to a "learning culture," transforming every failure into a valuable cognitive upgrade that drives growth.
Section 1: Failure Is Not the Mother of Success — Controllable, Information-Rich Failure Is
"Failure is the mother of success" — this phrase is like a warm bowl of chicken soup for the soul, often used to comfort those who have suffered setbacks. It implies that if we just keep failing, we will somehow, mysteriously, find our way to success.
This is a dangerous, romanticized misunderstanding.
In the real world, the vast majority of failures are just failures. They are stepping stones to greater failures, poisons that consume resources, sap morale, and ultimately lead to elimination. A catastrophic product launch can bankrupt a company outright. A wrong strategic acquisition can cripple a century-old enterprise. A repeated, elementary mistake proves nothing but organizational foolishness, offering no beneficial insight.
Failure does not automatically and generously give birth to "success." It requires strict screening and transformation. Only a very specific type of failure is qualified to be the "mother of success."
So what kind of failure is "valuable failure"? It must satisfy two demanding conditions simultaneously: it must be "controllable" and "information-rich."
Condition One: Controllability — Installing an Airbag for Failure
Controllability means that the negative impact of a failure must be strictly limited to a known and bearable range. In other words, it is a gamble you can afford to lose. You cannot bet your entire fortune on an uncertain outcome.
This is consistent with the idea of "asymmetry" we discussed in Chapter One. A "controllable failure" has limited, known downside risk, while its potential upside — whether direct success or indirect cognitive gains — is large and open-ended.
Failure without controllability is catastrophic.
The 2010 Deepwater Horizon oil rig explosion at BP is a textbook example of "uncontrollable failure." In pursuit of lower costs and faster timelines, BP had numerous lapses in safety protocols. Ultimately, a small methane gas leak triggered a chain reaction, causing the entire rig to explode, killing 11 workers and unleashing the worst environmental disaster in American history in the Gulf of Mexico. This failure cost BP over $65 billion, destroyed its brand reputation, and nearly brought it to ruin. This is not the "mother of success"; it is the "grave of a company."
The "big bang" product launch is another common form of uncontrollable failure. A team spends years and a huge budget secretly developing a "revolutionary" product, then launches it all at once at a grand event. If the market reaction falls short, years of investment are lost, and team morale suffers a devastating blow. This "succeed or die trying" gamble has very poor controllability.
How to ensure the controllability of failure? The key is to limit the "blast radius" of failure.
Scale of Experiment: Any attempt to explore the unknown should start at the smallest possible scale. Want to test a new business model? Do not build factories and distribution channels from the start. Instead, create a simple web page, run a small amount of advertising, and see how many people are willing to pay for this "idea." This is the essence of MVP (Minimum Viable Product) thinking.
Resource Commitment: Set clear, limited budgets and time windows for exploratory projects. It is like at the poker table, using only a small portion of your chips to call an uncertain hand. Even if you lose, it causes no real harm.
Risk Isolation: Physically or legally isolate high-risk innovation projects from the organization's core business. For example, establish an independent subsidiary or a "Skunk Works," granting it a high degree of autonomy while ensuring its failure does not infect the parent company's stable operations.
Controllability is the "safety prerequisite" for embracing failure. It ensures that failure does not become a disaster, but merely a controllable tuition fee. It gives an organization the courage to experiment, because the worst outcome is already accounted for in the contingency plan.
Condition Two: Information Richness — Purchasing "Cognition" from Failure
Being controllable alone is not enough to make failure "valuable." A controllable but stupid failure is still a waste. For example, spending one hundred dollars to test the idea of "whether a human can jump across a ten-meter-wide canyon." The failure is controllable (you only lost $100), but it is worthless, because it tests a conclusion that physics has already confirmed with certainty. You gained no new information.
Information richness means whether the result of the failure can effectively test an important, uncertain assumption, thereby helping us reduce uncertainty about the future.
Every "valuable failure" is essentially a transaction to "purchase information." The cost we pay (time, money, energy) is the price for intelligence about "the truth of reality." If a failure yields no new, valuable intelligence, the transaction was a loss.
Low-information failures often stem from:
Repeating known errors: Knowingly using a flawed technical solution out of laziness or wishful thinking, leading to project failure.
Lack of clear hypotheses: "Let's just build an app!" — a blind attempt without a clear hypothesis. Even if it fails, you have no idea where the problem lies. Is it that the need does not exist? The design is poor? Or the marketing is ineffective? You get noise, not a clear signal.
Inability to attribute effectively: The experiment is poorly designed, with too many variables changed at once, making it impossible to determine which variable caused the outcome after failure.
How to ensure the "information richness" of failure?
- Start with a "falsifiable hypothesis": Every experiment must begin with a clear, testable hypothesis. For example, "We believe that if we change the registration button from blue to green, the new user registration conversion rate will increase by 10 percent." This hypothesis is specific, measurable, and falsifiable. Whether the result is an increase, a decrease, or no change, we gain valuable information about "the effect of color on user behavior."
- Design "clean" experiments: Control variables as much as possible, ensuring only one core hypothesis is tested at a time. A/B testing is a perfect embodiment of this idea. By showing two groups of similar users pages that are identical except for one variable (such as button color), we can cleanly attribute the change in results to that variable.
- Build a fast "feedback loop": The informational value of failure decays rapidly over time. You must establish a mechanism to see the results of experiments and learn from them as quickly as possible. The "short iteration cycles" in Agile development are designed to shorten the "hypothesis-develop-test-learn" feedback loop.
The Value Matrix of Failure
Combining the two dimensions of "controllability" and "information richness," we can construct a "Failure Value Matrix" to categorize different types of failure:
| High Information Richness | Low Information Richness | |
|---|---|---|
| High Controllability | Valuable Failure: Scientific experiments, MVP testing, A/B testing, the organization's core learning engine | Stupid Waste: Repeating known errors, hypothesis-free attempts, to be avoided through process and training |
| Low Controllability | High-Risk Exploration: Moonshot projects, R&D for new drugs, requires extreme caution, typically only affordable by nations or giants | Catastrophic Failure: BP oil spill, Kodak's strategic errors, must be avoided at all costs |
This matrix clearly tells us that as managers, our task is not to broadly "encourage failure," but to:
Vigorously create the conditions for "valuable failure": This is the core activity of building an antifragile organization.
Resolutely eliminate "stupid waste": This can be achieved through SOPs, knowledge bases, and enhanced training.
Treat "high-risk exploration" with awe and extreme caution: This requires the highest level of strategic decision-making.
Treat "catastrophic failure" as public enemy number one for organizational survival: This must be defended against at every level through risk management, safety redundancy, and other means.
Summary: Safe Failure
The proverb "failure is the mother of success" needs a lengthy footnote to be true. It should be: "A negative result from a rapid experiment conducted under the protection of a safety net (controllability), with the purpose of testing a clear hypothesis (information richness), is the cognitive stepping stone to ultimate success."
With this understanding, we move beyond blind fear or romantic fantasies about failure. We can begin, like scientists, to calmly, rationally, even excitedly, design and embrace those "valuable failures" that make us smarter.
Section 2: Organizational Principles of Fault-Tolerant Design: Firewalls, Modularity, and Rapid Iteration
If we agree that "valuable failure" is a necessity for organizational evolution, the next question is: how do we build an organizational "container" that can accommodate and even encourage this kind of failure?
The answer comes from a concept in engineering — "fault-tolerant design."
In aerospace, nuclear energy, and critical network services, engineers never assume that every component of a system will never fail. Instead, their design philosophy is: Assume that any component can fail at any time, and on that basis, build a system that can continue to operate or safely degrade even when a partial failure occurs.
A modern airliner, even if one engine flames out mid-flight, can still fly and land safely on the other. A well-designed data center, even if one rack goes down due to a power outage or overheating, seamlessly switches user access to other healthy servers.
This "fault tolerance" thinking can be transplanted directly into organizational design. A "fault-tolerant" organization is not a perfect team of "superhumans" who never make mistakes. It is a system whose very structure has the inherent capacity to absorb and digest errors. It acknowledges that humans make mistakes, markets change, and accidents happen.
Building this kind of "fault-tolerant" organization depends on three major design principles: firewalls, modularity, and rapid iteration.
Principle One: Firewalls — Limiting the Contagion of Failure
In the field of cybersecurity, a firewall is a barrier used to isolate and protect an internal network from external threats. In organizational design, the firewall principle refers to establishing a series of isolation mechanisms to ensure that a local, controllable failure does not spread like a virus, infecting and destroying the entire organization.
Its core goal, as discussed in the previous section, is to limit the "blast radius" of failure.
Organizations without firewalls are "tightly coupled." An error in one link quickly triggers a domino effect.
Lack of financial firewall: Placing all the company's money in a single account and using profits from the core business to subsidize a high-risk new project. If the new project fails, the huge losses could drag down the entire company's cash flow.
Lack of brand firewall: Endorsing an unproven new product under the company's main brand. If the new product has major defects or negative press, the main brand's reputation suffers severe damage.
Lack of cultural firewall: Applying the same evaluation criteria and the same incentive approach across the entire company. This forces innovative teams to abandon long-term, disruptive exploration to meet short-term financial requirements, as their "fear of failure" becomes "infected" by the mainstream culture.
How to build organizational "firewalls"?
- Independent Innovation Units: This is the most classic organizational firewall. Separate the small teams responsible for disruptive innovation from the mainstream business in terms of physical space, organizational structure, budget, and evaluation systems. Lockheed Martin's legendary "Skunk Works," operating in extreme secrecy and autonomy, developed a series of revolutionary aircraft — the U-2, the SR-71 "Blackbird" — with minimal manpower and time.
- Independent Legal Entities and Brands: When a company wants to enter an entirely new, high-risk area, it can consider establishing a wholly owned subsidiary with a new brand. This is like lowering a small exploration boat alongside the mother ship. Even if the exploration boat hits a reef and sinks, the mother ship remains unharmed. Google (now Alphabet) spinning off its numerous exploratory projects (such as Waymo for self-driving cars, Verily for life sciences) into independent subsidiaries is a manifestation of this firewall strategy.
- Sandboxed Budgets: Provide "sandbox budgets" for innovation projects. This budget is independent, protected, and has a clear cap. The project team has 100 percent control over resources within this "sandbox" and is free to experiment. But once the budget is used up, the project must either prove its value to enter the next phase, or be terminated. This ensures the cost of failure is controllable.
The firewall principle provides a safe "laboratory" for "valuable failure." It allows an organization to conduct its boldest experiments without endangering its own survival.
Principle Two: Modularity — Building a Lego-Like Organization
Modularity means decomposing a complex system into multiple relatively independent, loosely coupled "modules" that can interact through standardized interfaces.
Imagine building a castle with Lego bricks. Each brick is an independent module. You can replace a red brick with a blue one at any time without affecting the rest of the castle. If a brick is accidentally damaged, you only need to replace that one brick, not tear down the entire castle.
The opposite is "monolithic" design. Like carving a statue from a single block of marble — if you make a mistake while carving the nose, the entire statue may be ruined, beyond repair.
Organizations lacking modularity are those where "one move affects the whole body."
Monolithic technology architecture: All functional code is tangled together. Fixing one small bug can cause unexpected new problems elsewhere in the system. Launching a new feature requires a lengthy, high-risk regression test of the entire system.
Functional matrix organization: A project must go through layers of approval and collaboration from product, design, engineering, testing, marketing, and other departments. Any delay or disagreement in any department brings the entire project to a halt. Power is centralized; decisions are slow.
How to build a "modular" organization?
- Microservices Architecture: This is the ultimate application of modular thinking in the technology domain. Decompose a large application into dozens or even hundreds of tiny, independent services. Each service is owned by a small team and can be independently developed, deployed, and scaled. The failure of one service (like "user comments") does not affect the normal operation of other services (like "payment").
- Amazon's "Two-Pizza" Teams: This is not just a principle of organizational scale, but a profound modular design. Each "two-pizza" team is like an independent "microservice." They are required to interact with other teams through clear APIs, not through endless meetings and emails. Bezos even had a famous "API mandate": all data and function calls between teams must go through network service interfaces; no other forms of backdoor communication are allowed. This forcibly achieved complete "decoupling" within the organization.
- Project-Based "Task Forces": When facing a complex challenge, assemble a temporary, fully empowered "task force" by drawing members from different departments. This task force is like an independent module, with its own goals, resources, and decision-making authority, and dissolves when the mission is complete. This is far more agile than inefficient collaboration across rigid departmental silos.
The modularity principle decomposes one large, uncontrollable failure risk into N small, controllable failure risks. It allows an organization to constantly and cheaply repair and iterate parts of itself, like replacing Lego bricks, enabling continuous overall evolution.
Principle Three: Rapid Iteration — Using Frequency to Combat Uncertainty
Rapid iteration means frequently making adjustments and optimizations to products or strategies through extremely short "develop-test-learn" cycles, moving forward in small, rapid steps.
Its core logic is: Since we cannot predict the future, the best strategy is to "bump into" the future as quickly and cheaply as possible, and then learn and adjust our direction from the "echo" of each collision.
Organizations lacking rapid iteration capability are clumsy and slow.
Waterfall development: Spending a year on requirements analysis and design, another year on development, and half a year on testing, finally releasing a product that may no longer meet market needs. The feedback loop is too long; the cost of failure is too high.
Annual strategic planning: Drawing up a thick, detailed strategic plan covering the next twelve months at the end of each year. Then, for the next twelve months, rigidly executing this plan regardless of how the external environment changes.
How to achieve "rapid iteration"?
- Agile Development and Continuous Delivery: This is a revolution in software development. It breaks large development tasks into small "sprints" of one to two weeks. At the end of each sprint, a usable, deliverable product increment is produced. This allows the team to continuously receive user feedback and make rapid adjustments.
- A/B Testing Culture: Stop arguing in meetings about "what color the button should be." Design two versions, direct 50 percent of user traffic to version A and 50 percent to version B, and let real data decide which version is better. Turn A/B testing from a technical tool into the default method for all organizational decisions — from product design to marketing copy, even pricing strategy.
- "Occam's Razor" and "Good Enough": Occam's Razor states: "Entities must not be multiplied beyond necessity." In the early stages of product design, cut all non-core features. Your goal is not to build a fully featured "Swiss Army knife," but a "bottle opener" that does one core thing exceptionally well. This dramatically reduces development complexity and shortens iteration cycles.
The rapid iteration principle decomposes one large, fatal failure into N small, information-rich failures. It allows an organization to perceive and adapt to a fuzzy, uncertain world with "pixel-level" resolution. It is a survival wisdom that uses "frequency" to trade for "accuracy."
Summary: Fault-Tolerant Design
Firewalls, modularity, and rapid iteration — these three principles together build an "antifragile" organizational operating system.
Firewalls provide the "safe container," allowing failure to be safely detonated.
Modularity provides the "flexible structure," localizing the impact of failure.
Rapid iteration provides the "learning engine," maximizing the extraction of value from failure.
An organization that possesses all three characteristics will no longer fear failure. On the contrary, it will proactively, even cheerfully, embrace "valuable failures," knowing that each controllable, information-rich failure is like a "vaccination" that makes its immune system stronger.
Section 3: Case Study — Netflix's "Chaos Engineering": Proactively Creating Disorder to Enhance Systemic Resilience
If most companies still maintain an attitude of "passive tolerance" toward failure, Netflix has pushed this thinking to an awe-inspiring new level: proactively, systematically, creating failures in its own production environment.
This astonishing practice is called "Chaos Engineering." It is the ultimate, purest expression of the "antifragile" organizational design philosophy. Through a deep analysis of Netflix's "Chaos Engineering," we can see how an organization, by "embracing failure," forges itself from a fragile system into an almost indestructible "phoenix."
Background: From DVD Rental to Cloud Giant — The Growing Pains
In 2008, Netflix suffered a nearly fatal database corruption incident that disrupted its DVD-by-mail service for three full days. This event made then-CEO Reed Hastings and his technical team realize that their centralized, traditional IT architecture was extremely fragile.
At the same time, Netflix was undergoing an even more ambitious strategic transformation: from a company that mailed DVDs to a full-fledged online streaming service. This meant their service needed unprecedented stability, scalability, and global availability. Any extended outage could lead to the loss of millions of subscribers.
To achieve this goal, Netflix made a decision that was considered extremely bold at the time: abandoning its own data centers and migrating all core operations to Amazon's AWS public cloud.
The benefit of this decision was that AWS provided powerful elasticity and global coverage. But it also brought an entirely new, enormous challenge: in a public cloud environment, partial failures are not accidental — they are inevitable. Amazon's servers go down, networks experience jitter, and storage has errors. You cannot control the underlying infrastructure.
Faced with this uncertain new environment, the traditional operations mindset of "pursuing zero failures" was completely bankrupt. Netflix's technical team had to undergo a fundamental paradigm shift. Their new philosophy was:
"The best way to avoid failure is to fail constantly."
The Birth of Chaos Monkey: Raising a "Gorilla" in Your Own Backyard
Based on this new philosophy, in 2010, a team of Netflix engineers developed an internal tool, giving it a vivid and memorable name: "Chaos Monkey."
Chaos Monkey is a program. Its job is simple, but hair-raising: during business hours on weekdays, it randomly, without warning, forcibly shuts down server instances in Netflix's production environment — the real environment serving global users.
Imagine this scenario: you are watching House of Cards online, and suddenly, the server responsible for streaming your video has been "killed" by the company's own Chaos Monkey.
This would be unthinkable in any traditional IT department — a crazy act that would get an engineer fired on the spot. Why did Netflix do this?
The logic of Chaos Monkey is exactly the same as the logic of vaccination.
It is a "controllable" failure vaccine: Chaos Monkey's "destruction" is small-scale and random. It shuts down only a few instances at a time, not the entire service.
It is an "information-rich" stress test: Every "murder" is a real test of the system's "fault tolerance." If Chaos Monkey shuts down a server and the user's video playback stutters, that exposes a "single point of failure" in the system design. That failure is an extremely valuable piece of "information."
It forcibly drives "antifragile" design: The existence of Chaos Monkey forces every Netflix engineer, when writing code and designing services, to take as a default premise from day one that "any dependency can disappear at any moment." They can no longer rely on luck; they must build automated failover, service degradation, and redundancy backup mechanisms for their services.
Chaos Monkey is like a strict fitness coach who is always on patrol, constantly hitting every "muscle group" of the system with small weights. If a muscle group is weak, it will feel sore from the impact (revealing a problem). Engineers must then repair and strengthen that muscle group. Day by day, the entire system, through this continuous, low-dose "pain," grows stronger and more resilient.
From Monkey to Simian Army: The Systematization of Chaos Engineering
The success of Chaos Monkey greatly encouraged Netflix. They realized that "randomly shutting down servers" was just the tip of the iceberg. A complex system can experience all kinds of strange and varied failures.
So they expanded the idea of "Chaos Engineering" into a much larger toolset, playfully called the "Simian Army." Besides Chaos Monkey, this army includes:
"Latency Monkey": Artificially injects latency into network communications to simulate network congestion or slow service response, testing the system's performance in a "slow" environment.
"Conformity Monkey": Checks every instance in the system to see if it follows best practices and configuration standards. If it finds a "non-conforming" instance, it shuts it down.
"Doctor Monkey": Monitors server health. If it finds an instance showing signs of ill health (such as high CPU usage), it proactively removes it from service and notifies the relevant team to fix it.
"Janitor Monkey": Searches the cloud environment for and cleans up redundant resources that are no longer in use, reducing waste and potential security risks.
"Chaos Kong": The ultimate weapon of the Simian Army. Instead of shutting down individual servers, it simulates the complete failure of an entire AWS Availability Zone (a physically isolated data center cluster). By "taking out" an entire region, it tests Netflix's cross-region disaster recovery and resilience capabilities.
The existence of the Simian Army means that Netflix has transformed "creating failures" from an occasional tool into a systematic "immune system" integrated into daily development and operations processes.
The Cultural Foundation Behind Chaos Engineering
It must be emphasized that "Chaos Engineering" is not just a set of technical tools — it is a profound culture. Without the right cultural soil, any attempt to introduce Chaos Monkey would only lead to chaos and disaster.
This culture includes at least two core elements:
- An Absolute "Blameless" Culture: When Chaos Monkey exposes a problem, the entire organization's reaction is not to ask "whose fault is this?" but to collectively celebrate: "Great, we found a vulnerability before a real user did!" Engineers know that exposing weaknesses is encouraged, not punished. This motivates them to build stronger systems, rather than finding ways to hide and cover up problems.
- A Strong "Ownership" Mentality: Netflix engineers are given a high degree of autonomy. They have "cradle-to-grave" responsibility for their own services. This "you build it, you run it" model means engineers have the strongest motivation to ensure their services can withstand Chaos Monkey's tests. Because if a service is brought down by the "monkey" in the middle of the night, the person called to fix it is themselves.
Summary: Case Study Analysis
Netflix's "Chaos Engineering" is the ultimate declaration of the philosophy of "embracing valuable failure." It tells us:
Antifragility is not innate; it must be deliberately "trained." It requires continuous, low-dose pressure and impact.
Rather than passively waiting for failure, proactively design failure. Turn uncertainty from an external, uncontrollable threat into an internal, manageable training tool.
The strongest systems are not those that never make mistakes, but those that can quickly learn from mistakes and recover.
Of course, not every company needs a "Chaos Kong." But the philosophy behind Chaos Engineering is universal. It challenges us to think: in our organization, is there a "mini," "safe" version of "Chaos Monkey"?
Perhaps it is an A/B test, a "red/blue team" exercise, a hackathon where employees are allowed to mess up, or simply a meeting where everyone can discuss failure openly and fearlessly.
Section 4: Toolbox 3 — How to Properly Conduct a "Failure Retrospective": From Blame to Information Extraction
Theory, principles, and case studies must ultimately be translated into daily organizational practice. And in the matter of "embracing valuable failure," the most important, most central, and most culture-reflecting practice is: "how to conduct a failure retrospective."
A failure (whether a project delay, a product defect, or a server outage) is like an "unscheduled" scientific experiment. It produces a wealth of raw data about the vulnerabilities of our systems and processes. The retrospective is the laboratory where we process and analyze this data to "extract information."
Yet in most organizations, the retrospective (a term borrowed from medicine with ominous connotations) turns into something else entirely: a "blame session," a "finger-pointing session," a tacit "political performance."
The Wrong Approach: A Courtroom Looking for a Criminal
Imagine a typical, wrong retrospective:
Atmosphere: Heavy, tense, like a courtroom. Senior leaders sit at the head of the table with stern faces.
Opening: The project lead or involved party begins to describe the "incident" in an apologetic, self-critical tone.
Process: Participants (especially from different departments) start asking questions, but the questions are often loaded: "Why didn't you consider...?" "Wasn't this process already specified? Why wasn't it followed?" "Who approved this plan?"
Focus: The entire discussion centers on "who made the mistake," not on "where the system went wrong."
Outcome: Eventually, one or several "responsible parties" are found. They make commitments and write assurances. The meeting minutes are filled with vague, unenforceable "improvement measures" like "So-and-so will strengthen their sense of responsibility" or "Such-and-such department will strictly implement the process."
Long-term impact: Participants learn that the best strategy next time a problem arises is to hide it, cover it up, or shift the blame to others. The organization becomes more "political," trust declines, and real problems are papered over until they erupt again on a larger scale.
This kind of "blame-oriented" retrospective not only fails to extract any valuable information from failure, but also poisons the organizational culture and systematically kills future "valuable failures."
The Right Approach: An Engineering Workshop to "Fix the System"
A proper retrospective, aimed at "extracting information," has a completely different underlying logic and cultural atmosphere. Its core is not to judge individuals, but, like a group of engineers, to calmly, objectively, and curiously "debug" the entire organizational system.
Google's Site Reliability Engineering team is a pioneer and exemplar of the "blameless postmortem" culture. Their practices provide a gold standard.
Core Principle: The "First Premise"
Before the retrospective begins, the facilitator must read aloud to all participants, in a serious tone, the "First Premise" articulated by Norm Kerth:
"Regardless of what we discover, we understand and truly believe that everyone did the best job they could, given what they knew at the time, their skills and abilities, the resources available, and the situation at hand."
This directive is the cornerstone of the entire blameless postmortem. It shifts the focus of discussion from "people's intentions" (assuming people are lazy and prone to error) to "the state of the system" (assuming people are well-intentioned, but the system is flawed). It creates an environment of "psychological safety" where participants dare to say "I thought..." or "I overlooked..." without fear of blame.
The Five Steps of a Proper Retrospective
An effective blameless retrospective should strictly follow these five steps:
Step One: Set the Scene, Reiterate Principles (5 minutes)
The facilitator opens the meeting, making it clear that the purpose is not to assign blame, but to learn.
Read the "First Premise": Ensure everyone understands and agrees to this basic premise.
Define the scope: Clarify exactly which failure event is being reviewed today.
Step Two: Construct an Objective, Non-Emotional Timeline (25 minutes)
Goal: Collaboratively reconstruct a minute-by-minute, fact-level account of the incident.
Method:
- Use a shared whiteboard or online document.
- Starting from the first trigger point of the incident, proceed in chronological order, having all relevant personnel describe what they observed and the actions they took at the time.
- Strictly prohibit any causal analysis or commentary during this step. State only facts. For example, the correct statement is "10:05, I received an alert that CPU usage exceeded 95 percent"; the wrong statement is "10:05, because so-and-so deployed buggy code, the server went into alarm."
- Cite objective evidence as much as possible: system logs, monitoring screenshots, chat records.
- Output: A detailed, agreed-upon timeline of events without any subjective judgments.
Step Three: Conduct Root Cause Analysis (40 minutes)
Goal: Starting from the surface problems in the "timeline," go deeper layer by layer to find the systemic root causes of the failure.
Core Tool: The "Five Whys":
- For a direct cause, repeatedly ask "why."
- Example:
- Problem: The website went down.
- Why? Because the database was overloaded.
- Why? Because a background task sent a large number of query requests to the database.
- Why? Because the code for this task did not have "rate limiting."
- Why? Because our code review checklist did not include a check for "rate limiting."
- Why? Because we have never encountered a database crash due to excessive traffic before; our system assumed traffic would always be within a controllable range. (This is the systemic root cause!)
Key: Shift the focus from "human error" to "systemic flaws." "Human error" should not be the end point of analysis, but the starting point. If a person made a mistake, we should ask: "Why did our system make it so easy for a well-intentioned person to make this mistake?"
Step Four: Extract Key Lessons and Generate Actionable Improvements (30 minutes)
Goal: Translate root causes into specific, executable improvement items that can prevent similar problems from recurring.
Method:
- For each root cause, brainstorm: "What can we do to fix or improve this system flaw?"
- Avoid ineffective improvement items: Items like "So-and-so will be more careful" or "We will strengthen training" are ineffective because they rely on human states, which are unreliable.
- Pursue effective improvement items: Good improvement items target the "system." For example: "Add a mandatory 'rate limiting' check to the code review checklist," "Add an automatic circuit-breaking mechanism for database overload protection," "Set default resource usage limits for background tasks."
- Output: A list of 3-5 most important improvement items. Each item must clearly state "what to do," "who is responsible," and "when to complete it."
Step Five: Summarize and Publish (10 minutes)
The facilitator summarizes: Quickly recap the key findings and most important improvement items.
Thank the participants: Express gratitude again for everyone's candor and constructive contributions.
Publish the postmortem report: Compile the results (timeline, root causes, improvement items) into a concise report and publish it to the company's internal knowledge base. This step is crucial. Publicity means knowledge sharing, meaning the lessons learned by one team can become an asset for the entire organization.
Summary: Toolbox
A failure is like an unexpected mineral exploration. It exposes valuable veins beneath the surface — the truth about the vulnerabilities of our organizational systems.
A "blame-oriented" retrospective is like, upon discovering a vein, busily punishing the exploration team member who accidentally tripped, then hastily covering up the mine pit and pretending nothing happened.
A "blameless" retrospective is like gathering the best engineers and geologists to conduct a thorough, scientific exploration of the vein. They draw detailed maps, analyze the composition of the ore, and design the most efficient and safest mining plan. In the end, this accidental "trip" becomes a great opportunity for organizational capability upgrade.
Will your next retrospective be a "courtroom" or an "engineering workshop"? Your choice will directly determine whether your organization grows more fragile in the shadow of failure, or becomes truly "antifragile" through the nourishment of failure.