If Part 1, "Build Well," was about forging a sharp sword, then Part 2, "Manage Well," is about mastering the swordsmanship. No matter how sharp the blade, if the wielder's technique is chaotic and intentions misguided, it can lightly wound the user or wreak havoc on society at large. The same holds true for this "sword of computing power" -- the national intelligent computing center.
An intelligent computing center with the most advanced hardware and the lowest PUE can nonetheless end in disaster if its operations and governance are mismanaged. It could degenerate into a "private mining field" for a few giants, widening the digital divide. It could squander enormous computing capacity through misallocation, producing idle waste. It could even touch off serious social and ethical problems through the misuse of algorithms.
The core of "Manage Well" is to install a precise, reliable operating system on this powerful "convergence engine" -- one aligned with the nation's strategic intent. This operating system is responsible for rationally allocating the underlying hardware resources to the myriad applications above, ensuring that computing power, this "productive force" of the new era, can truly, fairly, and efficiently empower the whole of society.
And the kernel of this operating system is computing power scheduling.
Computing power scheduling -- a term that sounds quite technical -- is at its essence a profound question of economics and governance. It answers a classic proposition of resource allocation: among whom should limited, precious computing resources be prioritized? At what price? Under what rules?
The answer to this question directly determines the nature and destiny of an intelligent computing center. Is it a commercial real estate venture that maximizes short-term profit, or a public infrastructure that undertakes to nurture a regional innovation ecosystem? Is it a digital-age "accelerator," or an amplifier of the "Matthew effect"?
If we adopt a completely laissez-faire market logic, the outcome is all but certain -- and it is the one we least wish to see: the formation of "computing power hegemony."
3.1 Refusing the "Big Players' Monopoly": The "Tragedy of the Commons" in the Computing World and Its Solutions
In a pure market environment, the scarcest resource will always be monopolized by the highest bidder. In today's China, who has the most urgent demand for high-end AI computing power and the strongest ability to pay? The answer is beyond doubt: the handful of leading internet tech giants and AI unicorns.
They are waging an arms race of large-model development, routinely requiring thousands or even tens of thousands of top-tier GPUs running continuously for months on end. To seize first-mover advantage in the market, they are willing to pay a steep premium.
At this point, if we, as the operator of a national intelligent computing center, make "occupancy rate" and "short-term revenue" our core KPIs, then the most rational business choice is to sign long-term exclusive agreements with these "big players," packaging and selling 80% or even 90% of the center's computing power to them in one stroke.
The advantages of this approach are obvious:
- Financial statements look good: instantly high occupancy, stable and predictable cash flow, quickly recovering huge construction and operating costs, and a clean report to superiors and investors.
- Operations become effortless: serving a few large clients is far simpler than serving hundreds of thousands of small ones -- clear requirements, strong technical capabilities, low communication costs.
On the surface this seems like a "win-win" deal. But if we shift our gaze from the short-term financial statement to the long-term healthy development of the regional digital economy, we will see that this is, in fact, a poison wrapped in honey.
The "big players' monopoly" model will inevitably lead to the following three grave consequences:
Stifling innovation among SMEs:
Innovation in the AI era is no longer merely "a flash of genius from a talent in a garage." The iteration of algorithms and the optimization of models depend heavily on large-scale computing power for experimentation and validation. An AI startup with an excellent algorithmic idea but no computing power is like an F1 car designer with no wind tunnel or track on which to test the design -- the idea can only ever remain on paper.
When the intelligent computing center's capacity is monopolized by the giants, the thousands of local tech SMEs, university research teams, and even hospital doctors hoping to use AI for assisted diagnosis will face the desperate reality of "no cards available." They are either priced out, or there is simply no scattered capacity left to rent, and so they are shut out of the AI era.
This effectively erects a "barrier to entry" on computing power. Only the giants can play the AI game, while smaller innovators cannot even obtain a ticket. A regional innovation ecosystem that loses these vibrant "capillaries" will eventually wither into a "desert of innovation" monopolized by the giants.
Intensifying the siphon effect on data and talent:
Compute availability can influence where some data, talent, and projects go, but it is not the only factor. When local SMEs cannot obtain local compute, they may also use remote colocation, multiple clouds, on-demand APIs, compressed models, edge devices, or deferred jobs, each with different costs and capability losses. Moving to a public cloud does not automatically transfer ownership of the data to the provider. What requires audit is contractual rights, key control, data location, access logs, portability, and exit costs.
At the same time, locally trained AI talent, lacking a platform to display their abilities, will drift toward the giants' headquarters cities such as Beijing, Shanghai, and Shenzhen. In the end, the "plane tree" that the local intelligent computing center planted at great cost attracts no "golden phoenix," but instead serves as a wedding garment for others, aggravating regional imbalances in development.
Forfeiting the bargaining power of regional industry:
When a region's digital economy -- from computing power to data to talent -- becomes highly dependent on one or two external giants, that region has in effect lost command of its own digital industrial destiny. The giants' platform rules, pricing strategies, and technical paths become the region's de facto standards.
When the local government sits down to negotiate industrial policy with the giants, it finds itself in an extremely passive position. The intelligent computing center, which should have been the strategic asset serving as the local government's "digital handle," instead becomes alienated into an "economic concession" through which external giants control the region.
This evolutionary path of the "tragedy of the commons" is entirely foreseeable. As managers of the national intelligent computing center, our foremost duty is to use the levers of governance to actively intervene in market failures and break the formation of "computing power hegemony."
3.1.1 The Lesson: The "Available Margin" That Must Be Reserved
Here we return once more to the core theoretical kernel of this book. A healthy, sustainable computing power ecosystem must always retain a certain proportion of "available margin."
This "available margin" is not capacity that sits idle because nobody uses it. Rather, it is a pool of computing power that we proactively, deliberately, and through institutional design set aside -- one that is not committed to long-term locking or commercial monetization.
It is like a nation's strategic petroleum reserve, or a city's emergency shelters. In ordinary times it appears to be "idle" and "yielding no returns," yet it is precisely what guarantees the resilience of the entire system and holds open its future potential.
This "available margin" carries strategic significance on at least three levels:
Preserving the "spark of innovation" for SMEs:
This reserved capacity should be channeled, in the form of "computing vouchers" or at inclusive prices, to vetted local SMEs and research teams. They may need only a few dozen cards, rented for a few hours or days, to validate a new algorithm or train a small model.
This "targeted drip" of computing support is negligible to the giants, but for these startup teams it is "fuel in snowy weather" -- the spark that ignites their dreams of innovation. Whatever the intensity of the giants' demand, we must ensure that this "reservoir of sparks" reserved for the small and micro never runs dry.
Preserving "strategic redundancy" to respond to emergencies:
When a city faces a sudden public health event (such as a pandemic), extreme weather (such as torrential rain and flooding), or a major safety incident, it must summon enormous computing power in an instant for simulation, emergency dispatch, and resource allocation. If 100% of our capacity is occupied by commercial tasks, then at the moment of crisis we will have no "cards" to call upon. Would we then interrupt the giants' model training, or sacrifice the city's emergency response capability? That would be a dilemma with the potential for enormous losses.
The existence of the "available margin" provides us with a buffer. It ensures that the city brain always holds a "strategic reserve force" available for priority dispatch, safeguarding the city's "anti-fragility."
Preserving a "space of possibility" for future uncertainty:
The evolution of technology and markets is shot through with uncertainty. What appears today to be the most mainstream of large-model training tasks may no longer be the principal consumer of computing power five years from now. A wholly new application we cannot yet imagine (such as the full-fidelity internet or brain-computer interfaces) may emerge, with a wholly different pattern of demand for computing power.
If we lock all of our capacity into a single application mode through long-term contracts, we seal off our capacity to adapt to the future. By reserving a portion of flexible, non-locked "available margin," we preserve a precious "experimental ground" for exploring and embracing the uncertainty ahead.
3.1.2 A Manager's Action Guide: How to Institutionalize the "Available Margin"
How, then, can managers translate the "available margin" from a theoretical concept into an executable operational strategy?
Establish a "Three-Tier Computing Power Quota" system:
When formulating computing power allocation policy, explicitly divide the center's total capacity into three quota pools:
- Tier 1 (Strategic Guarantee Pool, 15%-20%): This is our "available margin." In principle, this capacity does not participate in commercial bidding and is not bound by long-term locking agreements. It is supplied, through "computing vouchers" and an application-review process, to local small and micro tech enterprises, universities, research institutes, and government public service departments. Its pricing should be far below market rates, or even free. Its assessment criteria are not revenue but "breadth of empowerment" (how many entities supported) and "innovation outcomes" (how many patents incubated, how many new applications spawned).
- Tier 2 (Commercial Foundation Pool, 50%-60%): This is the "ballast" that keeps the intelligent computing center stable in operation. This capacity may be leased to large clients, including the giants, through medium- to long-term (e.g., 1-3 year) agreements to secure the center's baseline revenue and operating costs. The agreements must nonetheless stipulate that, should the government declare a state of emergency, it has the right to temporarily commandeer this capacity.
- Tier 3 (Flexible Bidding Pool, 20%-30%): This capacity operates on a fully market basis, using time-sliced, bid-based models open to all types of users. Prices fluctuate with market supply and demand, maximizing returns. This pool satisfies temporary and sudden computing needs in the market and is also an important source of the center's profits.
It should be noted that the ratios given above -- Tier 1 (Strategic Guarantee Pool) at 15%-20%, Tier 2 (Commercial Foundation Pool) at 50%-60%, Tier 3 (Flexible Bidding Pool) at 20%-30% -- are policy-design interval illustrations offered by this book for reference and adjustment, not optimal solutions validated by large-scale practice. Each locality should calibrate them against its local industrial structure, compute supply and demand, and fiscal conditions.
Establish a "Computing Power Inclusion Review Committee":
To ensure the fair and impartial distribution of the Tier 1 pool, a neutral review committee should be formed, comprising representatives of government, industry experts, investment institutions, and university professors. This committee reviews the "computing voucher" applications of SMEs and research teams, ensuring that precious inclusive capacity genuinely reaches those innovators with the technology, the potential, and the greatest need, and preventing subsidy fraud and resource waste.
Incorporate "Ecosystem Cultivation" into the core KPIs:
Most critically, the "baton" by which the operations team is assessed must be changed. We cannot look only at financial indicators such as revenue, profit, and occupancy rate. We must introduce a series of non-financial indicators measuring "ecosystem contribution" and give them considerable weight. For example:
- the number of local SMEs empowered
- the number of new AI-related patents and software copyrights
- the number of joint laboratories established with local universities
- the number of AI talents attracted and retained
Only when the "livelihood" of the operations team is deeply bound up with the flourishing of the local innovation ecosystem will they have genuine motivation to resist the short-term temptation of the "big players' monopoly" and to patiently, meticulously do the unglamorous work of serving SMEs.
Breaking "computing power hegemony" is not about making enemies of the giants. The giants remain important participants and contributors to the computing power ecosystem. What we must do is use the "visible hand" of governance to balance the "invisible hand" of the market, ensuring that beneath the towering tree of the giants there is still sunlight and rain enough to nourish a thriving "undergrowth of innovation." This is the breadth of vision and the sense of responsibility a national-level infrastructure should possess.
3.2 Time-Sharing Leasing Strategy: Managing Computing Power Like a Reversible Lane
Having established the strategic principle of "reserving an available margin," we also need a set of more refined tactical tools to raise the efficiency of computing resources across the dimension of time.
A common operational dilemma: by day, a host of online services with high real-time requirements (such as financial transactions, urban traffic dispatch, online customer service) crowd the computing resources, causing network congestion and task queuing; late at night, however, the load of these services falls to a trough, leaving large numbers of servers idle or lightly loaded, with enormous energy waste and opportunity costs.
This "tidal phenomenon" in computing demand is exactly analogous to the morning and evening rush hours of urban traffic. And an effective remedy for urban traffic congestion is the "reversible lane" -- dynamically adjusting the direction of travel of a lane according to the flow of traffic at different hours.
The same line of thinking applies fully to computing power scheduling. This is our "time-sharing leasing strategy."
Its core idea is to intelligently and differentially route different computing tasks to different time windows for execution, thereby "shaving peaks and filling valleys" in capacity utilization and maximizing the overall utilization of the 24-hour day.
3.2.1 Task Profiling: Identifying the "Express Lanes" and "Slow Lanes" of Computing Power
To implement time-sharing leasing, we must first learn to profile computing tasks and identify their key attributes. We can categorize them along two dimensions:
- Dimension 1: Latency Sensitivity
- High sensitivity (fast cars): these tasks demand extremely rapid response; a single second of delay can mean business failure or a sharp deterioration in user experience. They must be handled immediately and hold the highest preemption priority. For example:
- financial risk control and high-frequency trading
- real-time optimization of city traffic signals
- obstacle recognition for autonomous driving
- real-time speech recognition and translation in video conferencing
- Low sensitivity (slow cars): these tasks have requirements on total completion time but are indifferent to whether the intervening process is interrupted or queued. They can be deferred and processed in batches. For example:
- offline training of AI large models
- background rendering for film and animation
- scientific computing simulations in research
- batch data cleaning and ETL (extract, transform, load)
- High sensitivity (fast cars): these tasks demand extremely rapid response; a single second of delay can mean business failure or a sharp deterioration in user experience. They must be handled immediately and hold the highest preemption priority. For example:
- Dimension 2: Throughput Demand
- High throughput (heavy trucks): these tasks require mobilizing a large-scale computing cluster at once for prolonged, high-intensity parallel computation. They are the "big consumers" of capacity. For example, training a large model with hundreds of billions of parameters may require thousands of GPUs running continuously for weeks.
- Low throughput (cars): these tasks are relatively small in computational scale, perhaps needing only a few cards, or even a fraction of a single card, and can be completed in a short time. For example, generating an image with AI painting, or handling an API call.
By cross-combining these two dimensions, we can clearly classify the bewildering variety of computing tasks into four quadrants and devise differentiated scheduling strategies for each.
| High latency sensitivity -- "fast lane" | Low latency sensitivity -- "slow lane" | |
|---|---|---|
| High throughput -- "heavy truck" | Quadrant 1: urgent heavy tasks (e.g., city-scale disaster simulation, financial market crash scenario analysis, national-security-grade emergency decryption) Strategy: highest priority (QoS Level 1), absolute preemption. No regard for cost; ensure access to all required resources at the earliest moment. | Quadrant 2: planned heavy tasks (e.g., training of hundred-billion-parameter large models, film visual-effects rendering, gene sequence assembly) Strategy: planned scheduling, making use chiefly of nighttime or weekend troughs. Offer substantial price discounts to encourage users to "shave peaks and fill valleys." |
| Low throughput -- "car" | Quadrant 3: real-time light tasks (e.g., online AI customer service, real-time video stream analysis, high-frequency trading risk-control API calls) Strategy: guarantee immediate response during peak periods. Resources reserved in a pool; billing may be per invocation or per response time. | Quadrant 4: scattered light tasks (e.g., individual developer experiments, university students' research assignments, small-scale model validation) Strategy: lowest priority, "gap-filling" scheduling. Execute in every fragment of idle time, at inclusive prices approaching cost. |
3.2.2 Designing the Tidal Scheduler: The Twin Engines of Price Leverage and Intelligent Scheduling
With task profiling in hand, we can design an intelligent "tidal scheduler." This scheduler, like the dispatch center of a shrewd taxi company, uses two levers -- "price" and "algorithm" -- to steer and allocate the supply and demand of computing resources.
The price lever: establish a time-sharing pricing model
This is the most fundamental economic instrument. Just as power companies distinguish "peak, flat, and valley" electricity tariffs, we must likewise build a dynamic, time-based price system for computing power.
- Peak hours: for example, 9 a.m. to 5 p.m. on workdays. This window chiefly serves the "fast car" business of Quadrants 1 and 3. Computing prices are at their highest, reflecting scarcity. The high prices naturally "dissuade" the non-urgent "slow car" tasks, inducing them to queue for cheaper time slots on their own initiative.
- Flat hours: for example, 6 p.m. to 11 p.m. on workdays. Prices are moderate, absorbing some of the overflow from the day and some "slow car" tasks with certain time requirements.
- Valley hours: for example, midnight to 7 a.m., as well as weekends and holidays. This is the golden window for the "planned heavy tasks" of Quadrant 2. Prices can be set extremely low -- for instance, at a fraction of peak rates (in pricing strategy, on the order of two to three tenths of peak). Through the dramatic price differential, all high-throughput training and rendering tasks are attracted and incentivized to converge on this time slot.
This time-sharing pricing strategy not only smooths the computing load curve effectively but also brings additional revenue to the intelligent computing center. More importantly, it offers an affordable path to computing power for users who are cost-sensitive but flexible in time (such as research institutions and startups) -- and that in itself is a form of inclusion.
The intelligent scheduling algorithm: from "first come, first served" to "global optimum"
Beyond the price lever, the scheduler also needs powerful algorithmic capacity to execute more refined resource allocation. The traditional "first come, first served" algorithm is acutely inefficient in a complex, heterogeneous environment. A modern intelligent computing center scheduler should, at minimum, possess the following capabilities:
- Task priority and preemption: the scheduler must be able to identify task priorities. When an "urgent heavy task" of Quadrant 1 (such as a disaster warning) is submitted, the scheduler is empowered to suspend or even terminate lower-priority tasks in flight (such as a Quadrant 2 training job), instantly releasing resources to the high-priority task. Naturally, the interrupted task should be able to save its state and resume automatically later, and its user should receive appropriate compensation.
- Resource profiling and intelligent matching: the scheduler must profile not only tasks but also the underlying computing resources. It needs to know, in real time, which server carries domestic chip A and which carries NVIDIA chip B; which rack is air-cooled and which is liquid-cooled; which server has the highest network bandwidth. When a task is submitted, the scheduler, based on the task's characteristics (for instance, whether it requires a CUDA environment or an Ascend CANN environment), intelligently matches it to the most suitable physical server -- "putting good steel on the blade."
- Packing and gap-filling: this is an optimization algorithm. When a large task requiring a thousand cards is scheduled to run at night, the scheduler may detect a two-hour window before it begins. At that point, it automatically searches for those "scattered light tasks" (Quadrant 4) that can be completed within two hours and "fills" them into the gap, preventing idle resources. This ability to "seize every chance" can dramatically improve the utilization of fragmented capacity.
- Energy-aware scheduling: a more advanced scheduler also links with the data center's environmental monitoring system. It knows which rack is running hot and which area's power load is nearing its ceiling. When assigning tasks, it strives to spread the computing load evenly across the entire data center, avoiding local hotspots and thereby reducing cooling energy consumption at the macro level and optimizing PUE.
3.2.3 The Manager's Role: From "Landlord" to "Traffic Planner"
Implementing a time-sharing leasing strategy requires managers to transform their role completely.
We are no longer simple "computing power landlords" who rent out a rack and consider the matter settled. We must become, as it were, a city's "traffic planner" and its "director of the traffic bureau."
- As "traffic planner," we must study in depth the local industrial structure and the pattern of computing demand. Does the finance sector predominate, or manufacturing? Are there more online services by day, or more offline tasks by night? Through precise demand analysis, we scientifically design our time-sharing pricing model and our resource quota ratios.
- As "director of the traffic bureau," we must build a powerful, visual computing operations center. On the large screen of this center, we can observe the entire intelligent computing center's "computing heatmap" in real time: which tasks are running, what resources they occupy, when they are expected to finish; which resources lie idle. We can clearly see the tidal changes in computing load and, in accordance with the contingency plan, dynamically adjust our scheduling strategy.
Conclusion: Managing Computing Power Well Is Managing the Future Well
The core of Chapter 3 is to supply a solution for the operational governance of the intelligent computing center that joins "the Way" with "the Method."
- At the level of the "Way," there is the strategic resolve to "break computing power hegemony." We must recognize that the national intelligent computing center bears a public character and a strategic mission, and must never become the appendage of a few giants. Through an institutionalized "available margin," we preserve precious "sparks" for small and medium innovators and for the public interest, safeguarding the diversity and fairness of the entire digital ecosystem. This is a long-term commitment and a sense of responsibility that transcends short-term commercial gain.
- At the level of the "Method," there is the refined operation of "time-sharing leasing." Through task profiling, time-sharing pricing, and intelligent scheduling, we manage computing resources with the same precision one brings to a reversible lane, "shaving peaks and filling valleys." This not only greatly enhances resource utilization efficiency and economic returns but also, through a differentiated price system, opens a path to computing services for every distinct type of user.
"Building well" is only the prelude; "managing well" is the main text. An intelligent computing center that cannot schedule is like a giant with well-developed limbs but a sluggish mind -- possessing brute force in abundance, yet unable to unleash its true might.
Only when we establish a fair and efficient computing power scheduling mechanism can we ensure that the surging power welling up from the chips flows precisely, in an orderly and harmonious manner, into every corner of the national economy; only then can we ensure that the mighty "wild horse" of the algorithm is bridled with rational reins, galloping always in the direction that serves the overall well-being of society, and never running amok.
To manage computing power well is to manage the power of distribution over the data factor well, to manage the "faucet" of the digital economy well. It is to manage, for ourselves and for the next generation, a future brimming with limitless possibility.