If Chapter 1's "site selection and energy consumption" was about finding a physical-world coordinate for the intelligent computing center, then this chapter's "architecture and supply chain" is about building its "skeleton" and "nervous system" in the digital world. This is a deeper and more challenging topic, because it touches not on physical laws and economic costs, but on great power strategic interplay, technological sovereignty, and national security.
In the golden age of traditional IT construction, architectural design was largely a purely technical problem. The core task of CIOs and architects was to pick the most performant, cost-effective, and ecologically mature components (chips, servers, operating systems, databases) from an open, globalized market like shopping in a supermarket, and then assemble them into an efficient system. The creed of that era was "technology has no borders" and "let professionals do their jobs."
However, we must soberly recognize that era is gone forever.
When the "CHIPS Act" and "entity lists" become realities we face every day; when "decoupling" and "technology blockade" shift from geopolitical strategic interplay terminology to the sword of Damocles hanging over every tech enterprise, we must re-examine the architectural design of intelligent computing centers through a completely new, highly politicized lens.
Architecture is politics.
The meaning of this statement is: Today, what kind of chips, servers, and network solutions we choose for the national intelligent computing center is no longer a mere consideration of technical performance or commercial cost. Every architectural decision is, in essence, a positioning, a vote for a national technology path, a prediction of and hedge against future supply chain risks. It profoundly affects our nation's industrial security, data sovereignty, and even our right to speak in the future digital world.
An intelligent computing center that is entirely dependent on a single foreign supplier, no matter how high its benchmark scores or how low its PUE, is from a national security perspective, a castle built on sand. Its "lifeline" is in the hands of others, and it could be paralyzed at any moment by a single command.
Therefore, this chapter will directly confront two sharpest, most core "political" issues:
- "Bottlenecks" and "Spare Tires": At the hardware level, how to deal with chip blockades and build an autonomous, diversified, and resilient computing power foundation?
- Data Sovereignty: At the network and data level, how to ensure the nation's "right to observe" and "right to decide" are absolutely in its own hands, preventing the outflow or contamination of core data?
2.1 "Bottlenecks" and "Spare Tires": Survival Rules in the Heterogeneous Computing Era
For a long time, the global high-performance computing field, especially AI training, has been dominated by a single company -- NVIDIA of the United States. Its CUDA ecosystem, leveraging first-mover advantage and sustained investment, has built an extremely powerful and difficult-to-surpass moat. The vast majority of mainstream AI frameworks (such as TensorFlow, PyTorch), algorithm libraries, and application software are developed based on CUDA. This has led to a path dependency: to do AI, you must use NVIDIA GPUs.
This single-dominant situation might have been the most efficient in an era of peace and cooperation. But when geopolitical winds shift, and the most advanced GPUs are listed as embargoed items against China, this "optimal solution" instantly becomes the "most dangerous shackle." This is the root of the "bottleneck" dilemma we face.
In the face of this dilemma, we cannot wait helplessly, nor can we indulge in wishful thinking. We must actively seek change, using strategic initiative to compensate for technical passivity. Our core strategy is not to find a perfect "alternative" that can 100% replace NVIDIA, because this is unrealistic in the short term. Our strategy should be to build a powerful "heterogeneous computing power pool" capable of managing and dispatching a variety of different computing resources.
"Heterogeneous" is the key word for China's intelligent computing center construction in the next decade. It means our computing foundation will no longer be uniformly "N-cards," but a "mixed fleet" composed of domestic GPUs, domestic ASICs, domestic CPUs, and a small number of compliantly available foreign chips.
2.1.1 Saying Goodbye to "CUDA-Only Mentality": From Relying on an Ecosystem to Creating One
To build a heterogeneous computing pool, we must first break the "CUDA superstition" in our thinking.
CUDA's success is essentially an ecosystem victory of "hardware plus software plus developer community." The reason we could not do without it in the past was not only because of its powerful chip performance, but because all our software, algorithms, and talent have grown on this "soil." Migrating to a new "soil" comes with enormous costs.
Therefore, our counter-strategy must also be at the ecosystem level. As builders and operators of national intelligent computing centers, we possess immense power to define markets and guide demand. We must use this power to actively "create" a new ecosystem that does not rely on a single technology system.
Specifically, managers need to drive the following key tasks:
Clearly mandate "heterogeneous" orientation in procurement bidding:
In the procurement documents for the intelligent computing center, it must be clearly required that the computing power supply should be diversified. For example, it could stipulate that "domestic computing power shall account for no less than XX%," or "must include at least two or more AI chips with different technical architectures." This kind of procurement strategy acts as a "baton," sending a strong signal to the market: integrators who cling only to NVIDIA's thigh will not be able to obtain national-level projects. This will force the entire industry chain -- from chip manufacturers to server manufacturers to software developers -- to invest resources in adapting to and optimizing domestic computing power.
Invest in building a "Heterogeneous Computing Compilation and Scheduling Platform":
This is the "technical core" of building a heterogeneous computing pool. Imagine we have domestic AI chips from different manufacturers like Huawei Ascend, Cambricon, and Biren Technology, as well as a small number of downgraded NVIDIA GPUs. Their instruction sets, drivers, and memory architectures are all different. If application developers had to adapt to each chip individually, it would be a nightmare, and the ecosystem would never be built.
Therefore, the national intelligent computing center must take the lead in building a "unified programming model and compiler." This platform acts like a "universal interpreter." Application developers at the upper level only need to write their AI models in a unified, high-level language (for example, based on open-source Triton or a domestically led programming framework), and then the platform automatically compiles, optimizes, and dispatches it to the most suitable underlying hardware for execution.
This platform is the core of our own "non-CUDA" ecosystem. It "abstracts away" the complexity of the hardware from the upper-layer applications, allowing developers to "write once, run anywhere." Only when this "middle layer" is built can our "mixed fleet" truly form combat effectiveness, rather than being a heap of loose sand.
Establish a "Domestic Computing Power Application Demonstration Zone" and a "Migration Incentive Fund":
Building an ecosystem requires both "carrots" and "sticks." In addition to hard requirements in procurement (the stick), we need incentives (the carrot). The intelligent computing center can open up a dedicated area, deploy the latest domestic computing power clusters, and provide them at very low prices, or even free of charge, to universities, research institutions, and promising startups.
At the same time, set up a special fund to reward enterprises that take the lead in migrating their applications from the CUDA platform to the domestic platform and achieve good results. Developers who help domestic chip manufacturers discover and fix driver bugs, performance bottlenecks, etc., during the migration process should be generously rewarded. We need to cultivate a community atmosphere of "use domestic, love domestic, optimize domestic," so that the earliest adopters receive the richest rewards.
2.1.2 The Strategic Philosophy of the "Spare Tire": From Pursuing Optimality to Ensuring Survival
In the process of building a heterogeneous computing pool, we will inevitably hear voices saying: "Domestic chips are not as performant as NVIDIA's, the ecosystem is immature, using them will affect efficiency, and might even cause problems."
From a purely technical and commercial perspective, this voice is entirely correct. But if we look at it from the strategic height of "architecture is politics," this view appears short-sighted.
Here, we need to introduce the strategic philosophy of the "spare tire."
The meaning of a "spare tire" is not that it is more comfortable or easier to use than the "primary tire" in normal times. The sole value of a "spare tire" is that when the "primary tire" blows out, it can keep your car moving forward, rather than leaving you stranded in the wilderness. Its core value is to "ensure survival," not to "pursue optimality."
Introducing domestic computing power into our intelligent computing center is to prepare a crucial "spare tire" for the nation's digital economy.
Tolerance and patience with the "spare tire":
We must recognize that no new technological ecosystem can be built overnight. In the early stages of development, domestic chips will inevitably have performance gaps, software bugs, and incomplete documentation. As managers of the national team, we cannot act like picky consumers, demanding perfection. Instead, we should play the role of a "patient practice partner" and "chief trial user."
We need to establish a fault-tolerance mechanism. In the operation of the intelligent computing center, efficiency losses or stability issues caused by using domestic "spare tires" should be regarded as a necessary "strategic cost" or "R&D investment," rather than the "dereliction of duty" of the operations team. We must provide the most valuable large-scale, real-world "testing ground" for the growth of domestic chips. Only after being tempered through countless trials in our intelligent computing centers and solving numerous practical problems, can these "spare tires" gradually mature and eventually gain the strength to compete with the "primary tire."
The strategic deterrence of the "spare tire":
Even in the most ideal scenario, where our "spare tire" never catches up with the top-tier "primary tire" in performance, its mere existence has enormous strategic value.
As long as we have a usable, systematic "spare tire" in hand, the harm of external blockades and sanctions against us is limited. The other side will understand that trying to defeat us with a single blow through "bottlenecking" will not work. This, in turn, increases our bargaining power at the negotiating table and may even force the other side to relax restrictions. A buyer who is always on their knees and has no alternatives will never earn respect.
Therefore, dedicating a budget share in the range of twenty to thirty percent (the specific ratio should follow supply-chain risk assessment; given here as a policy illustration) to building a domestic computing power cluster that may seem "inefficient" in normal times is not a waste -- it is a form of "strategic insurance." This insurance ensures that in extreme situations, the spark of AI development for the entire nation will not be extinguished.
Diversity and complementarity of the "spare tire":
Our "spare tire" strategy should not be a single bet on one domestic manufacturer. We should simultaneously support and adopt domestic chips with different technical paths. Some may excel in AI training (GPU-like), some in AI inference (ASIC-like), and some may have advantages in graphics rendering.
A healthy "spare tire library" should be diverse. This not only prevents us from going from being "bottlenecked by one foreign giant" to being "held hostage by one domestic giant," but more importantly, different technical paths can complement each other. In future applications, a complex task might be decomposed by an intelligent scheduling platform: one part of the graphics processing task is assigned to chip A, which is good at rendering; another part of the model inference task is assigned to chip B, which has the best energy efficiency ratio.
This strategy of "mixing cards from different brands, not putting all eggs in one basket" not only diversifies the risk of any single technical path failing but also, through internal "horse racing" mechanism, stimulates innovation vitality in the domestic chip industry.
2.1.3 The Manager's Role: From "Purchaser" to "Industry Chain Chief"
In summary, facing the "bottleneck" dilemma, the role of the intelligent computing center manager must undergo a fundamental transformation.
We are no longer a passive "purchaser" choosing the best products in a given market. We must actively and consciously play the role of an "industry chain chief."
Every one of our tenders, every technical selection, every application deployment is shaping and cultivating an autonomous and controllable domestic computing power industry chain. Our intelligent computing center is the "aggregate demander" and "aggregate integrator" of this industry chain.
We need to use our orders to feed domestic chip companies; use our scenarios to polish the domestic software ecosystem; use our platform to cultivate talent familiar with domestic computing power. This is a difficult, long, and potentially "thankless" process in the short term. But it is the historical mission that must be undertaken by our generation of digital infrastructure builders.
Building a well-constructed heterogeneous computing pool is like building a "Noah's Ark" for our nation in the stormy global technology competition -- solidly structured, diversely powered, and highly resilient. This is the profound connotation of "architecture is politics" at the hardware level.
2.2 Data Sovereignty: Fortifying the "Border Defense" of Digital Territory
If hardware supply chain security solves the problem of "whether we have the ability to observe and think," then data sovereignty security solves the fundamental question of "whether what we observe is the real world" and "who controls the results of our thinking."
In the digital age, data is the new territory. Every bit of core data is as sacred and inviolable as an inch of our physical territory. The national intelligent computing center, as the hub for the aggregation, processing, and application of massive data, is essentially the "central server" and "strategic intelligence bureau" of our digital territory.
Therefore, ensuring data sovereignty is an overriding, absolute "red line" in the architectural design of the intelligent computing center. On this red line, there is no room for compromise.
The concept of "data sovereignty" sounds grand, but it can be broken down into two very specific "political tasks" that must be locked in place during the architectural design phase:
- Reducing the risk that the "right to observe" flows outward: Use classification, least privilege, encryption, auditing, and cross-domain approval to prevent unauthorized export of core sensitive data where possible, while retaining detection and response for personnel, endpoint, supply-chain, and implementation risks that cannot be eliminated completely.
- Preventing the "contamination of observation": Ensuring that the data and models we use for decision-making are clean, real, and unaltered, preventing external forces from influencing or even manipulating our decisions through "data poisoning" or "model backdoors."
To accomplish these two tasks, we need to deploy two layered "border defense lines" in the network architecture and data governance of the intelligent computing center: physical isolation and network slicing.
2.2.1 Physical Isolation: The "Forbidden City" of Core Data
For a national intelligent computing center, the sensitivity and importance of the data it processes are graded.
- Highest level (core sovereign data): This includes "red line data" involving national security, economic lifelines, and citizens' fundamental information. For example:
- Government data: household registration, social security and medical insurance, taxation, public security surveillance, national geographic surveying and mapping, etc.
- Core financial data: bank core transaction systems, national payment and clearing systems, credit reporting data, etc.
- Critical infrastructure data: operational data of lifeline projects like the power grid, water network, transportation, and communications.
- Non-anonymized personal biometric information: facial images, fingerprints, DNA, etc.
For this type of data, we must adopt the oldest and most effective security measure -- physical isolation.
Physical isolation, as the name suggests, means placing the server cluster that processes this core data in an environment that is physically completely disconnected from the public internet and even from the ordinary business network within the intelligent computing center.
This means:
- Independent server room space: With independent access control, security, and monitoring, unauthorized personnel are absolutely unable to enter.
- Independent network equipment: Using independent switches, routers, and firewalls, with network cables that have absolutely no physical connection to any external network.
- Independent storage system: Data is stored in dedicated, encrypted arrays.
- Strict personnel and media management: Any person, USB drive, or laptop entering this "red zone" must undergo the strictest scrutiny and control. Data import and export can only be conducted through a single, tightly monitored "data ferry" system.
In architectural design, we must, like planning a city, designate a "government affairs zone" or "financial zone" within the intelligent computing center. This zone is the "Forbidden City" or "Zhongnanhai" of our digital territory. In terms of network topology, it is an absolute "isolated island."
Some may question whether physical isolation affects efficiency or makes data sharing difficult.
The answer is yes. Physical isolation is itself a strategy that sacrifices some convenience for ultimate security. But for core sovereign data, this sacrifice is entirely necessary and worthwhile. Security is the 1; efficiency and convenience are the zeros after it. Without the 1, no number of zeros has any meaning.
As managers, our task is not to challenge the necessity of physical isolation, but to design safer and more efficient "data ferry" and "cross-domain computing" processes. For example, we can use privacy computing technologies like "federated learning" and "secure multi-party computation" to enable model capability sharing without leaving the original data domain. For instance, Bank A and Bank B can jointly train an anti-fraud model using secure multi-party computation within their respective physically isolated zones, without exchanging each other's customer transaction data.
Physical isolation is the first and most solid line of defense for safeguarding data sovereignty. It uses the walls of the physical world to create a high-security vault for the "imperial seal" of our digital world -- but the vault itself still requires supporting personnel, media, and process controls; physical isolation reduces the risk along external network paths, not all risk.
2.2.2 Network Slicing: Customized "Highways" for Different Businesses
However, not all data in an intelligent computing center is of the highest level of sensitivity. There are large amounts of secondary sensitive data and public data that need to be processed. For example:
- Industrial data: Equipment operational data collected by the industrial internet, enterprise supply chain data, etc.
- Urban sensing data: Traffic flow videos from traffic cameras, air quality data from environmental sensors, etc.
- Research and internet data: Public datasets used for training general-purpose large models, simulation data from research institutions, etc.
If all of this data were subjected to physical isolation, the intelligent computing center would become a collection of isolated data silos, greatly diminishing its value as a "hub for data element circulation." But if it were placed on the same network plane as core sovereign data, it would create enormous security risks -- once an ordinary business server is breached, attackers could use it as a springboard to move laterally and infiltrate the core area.
How to resolve this contradiction? The answer is network slicing.
Network slicing, a core technology originating from 5G communications, can certainly be applied to the internal network design of intelligent computing centers. Its core concept is: on top of a shared set of physical network infrastructure, use virtualization technology to carve out multiple mutually isolated, end-to-end logical networks.
We can imagine the entire physical network of the intelligent computing center as an 8-lane super highway. Network slicing technology is like using virtual "isolation barriers" on this highway to carve out different exclusive lanes:
- Slice 1: "Government Affairs Express Lane," a 2-lane, highest-security channel. Only authorized government application traffic can enter this lane. It has the highest bandwidth guarantee, the lowest latency commitment, and the strictest security policies. Even if a "car accident" (cyberattack) occurs on the next lane, it absolutely will not affect traffic on this express lane.
- Slice 2: "Financial Express Lane," similarly carving out a highly reliable, highly secure exclusive channel for financial businesses.
- Slice 3: "Industrial Internet Channel," providing a stable and reliable data transmission channel for local manufacturing enterprises, ensuring the continuity of industrial production.
- Slice 4: "Public AI Training Channel," providing a high-bandwidth channel for universities and developers to access public datasets and conduct model training. The security level of this channel can be relatively lower, but it must be strictly isolated from the previous critical business channels.
Network slicing is the "precision border defense" of data sovereignty. Compared to physical isolation, it is a more flexible and efficient isolation method. It achieves:
- Security isolation: When policy is correct, implementation is reliable, and controls are continuously verified, traffic in different slices is logically isolated. Control-plane flaws, misconfiguration, shared-device defects, and side channels can still create cross-slice effects, so logical slicing is not equivalent to physical isolation.
- Resource assurance: Bandwidth, latency, jitter, and other quality-of-service parameters can be configured per slice to reduce the chance that ordinary congestion affects core workloads. Redundancy, monitoring, and degradation plans are still needed for extreme overload, link failure, and configuration error.
- Flexibility and scalability: Network slices can be dynamically created, adjusted, and deleted based on business needs, without modifying physical wiring.
As managers, in the network architecture tenders for the intelligent computing center, we must include "whether it supports end-to-end network slicing capability" as a key technical requirement. We need a powerful "network controller," which is like the "traffic management bureau" of this highway system, capable of visually managing, monitoring, and scheduling all slices.
2.2.3 The Ultimate Goal: Ensuring the "Right to Observe" Is Firmly in Hand
The combination of physical isolation and network slicing ultimately points to one ultimate goal: keeping the "right to observe" our digital territory as completely and controllably in our own hands as possible.
The outflow of the "right to observe" means our core data is stolen. Hostile forces can use it to analyze our economic operations, social trends, and military deployments, making us strategically transparent.
"Contamination of observation" is a more insidious and dangerous attack. By implanting tampered data or AI models with backdoors, it leads us to make wrong decisions based on erroneous "observations." Imagine if the traffic camera data of the city brain were maliciously tampered with, causing the system to set all main roads to red lights during the morning rush hour -- what chaos would ensue? If a financial risk control model were implanted with a backdoor, approving all loan applications in a certain sector at a specific time, what systemic risk would be created?
Therefore, safeguarding data sovereignty is not just about preventing data leakage; it is also about ensuring the integrity and availability of data. Our architectural design must be able to defend against a full spectrum of attacks, from hardware backdoors and operating system vulnerabilities to data poisoning and model adversarial attacks.
This requires that, based on physical isolation and network slicing, we also deploy a complete defense-in-depth system: zero-trust network access, encrypted data storage, data lineage tracking, AI model security detection, and so on.
Conclusion: The Architect's "Political Awareness"
The core of this chapter is for every manager and technical expert involved in building national intelligent computing centers to develop a high level of "political awareness."
When choosing a technical architecture, what should appear in our minds is not just performance curves and cost reports, but also the map of global technology strategic interplay and the domain map of national digital sovereignty.
- Choosing heterogeneous computing power is a vote of confidence in the independence and autonomy of the nation's technology. We should be willing to be "practice partners," using our patience and resources to nurture the seedlings of domestic technology and letting the "spare tire" grow strong.
- Deploying physical isolation and network slicing is about building the "Great Wall" and "border stations" for our digital territory. We must use the strictest standards to guard the core data related to the national economy and people's livelihoods, keeping the nation's "right to observe" and "right to decide" firmly in our own hands.
Every decision in this process is weighty and difficult. It requires a prudent balance between long-term strategy and short-term benefits, between autonomous control and technological advancement. This tests not only our technical judgment but also our strategic composure, historical patience, and sense of responsibility.
Building the "skeleton" and "nervous system" of the intelligent computing center, and sharpening this "eye of observation," is how we can truly "see clearly, see far, and see accurately," and how we can firmly grasp our own destiny in future competition among great powers. This is the full meaning of "architecture is politics."