In Chapter 2, we erected the solid operational "skeleton" for the computing platform — defining the organizational structure and institutional norms. However, having only an efficient operations system is like possessing an advanced traffic management system without determining how many lanes to build or what speed limits to set. Traffic flow forecasting and road planning are the prerequisites for determining whether the entire transportation system can support urban development. Similarly, computing planning and capacity management are the "first move" and "lifeline" that determine whether the computing platform can support the grand vision of enterprise intelligent transformation.
This chapter focuses on the planning dimension within the "Five-in-One" overall construction approach. We will deeply explore how to establish a full-process planning management system from demand insight and scientific forecasting to precise investment and elastic growth.
Its core goal is to answer four closely interlocking strategic questions:
- How much computing power do we need? (Demand consolidation)
- What kind of computing power should we build? (Capacity planning)
- How should we build and develop computing power? (Elastic expansion)
- How do we treat existing computing resources? (Resource integration)
Solving these four problems well ensures that the investment in computing resources can both meet current and future business needs, avoiding "resource starvation" hindering innovation, and maximize return on investment (ROI), avoiding "over-investment" causing waste. This is a science and art seeking optimal solutions under uncertainty.
Demand Consolidation: From Business and Back to Business
The starting point of all planning comes from demand. The value of computing power is ultimately reflected in its support for business. Therefore, computing planning must be deeply rooted in the soil of business, following the basic principle of "from business and back to business." Demand consolidation is precisely the process of systematically transforming the scattered, vague, and potential business "ideas" within the enterprise into clear, quantifiable, and plannable computing "demands." If this process is done roughly, subsequent planning will become a "castle in the air."
The Challenge of Demand Consolidation: Why "Guesstimation" Is Not Acceptable?
In enterprises lacking systematic demand consolidation mechanisms, the proposal and aggregation of computing demands often face many challenges:
"Fragmentation" and "Hiddenness" of Demands: Demands are often proposed independently by various business departments or project teams, lacking an overall perspective. Department A wants to do drone inspection, Department B wants to do intelligent customer service, and Department C is exploring marketing data mining. When proposed, these demands often focus only on their own limited scope, making it difficult to form a complete picture of the company's overall computing needs. Moreover, there are a large number of "hidden demands" — business personnel have pain points and ideas but, because they don't know what AI technology can do or what it requires, fail to translate them into clear computing demands.
"Ambiguity" and "Uncertainty" of Demands: When proposing demands, business departments often use qualitative descriptions, such as "I need very strong computing power to train a large model," but find it difficult to provide key quantitative indicators such as the specific parameter scale, training data volume, and training cycle. At the same time, AI projects themselves have high levels of exploration and uncertainty, so initial demand estimates may deviate significantly from final actual consumption.
"Tidal" and "Periodic" Nature of Demands: Computing demands of different businesses show significant peak-valley differences over time. For example, finance department AI applications may peak at month-end or quarter-end; research projects may focus on large-scale training at the beginning of the project. If the peak demands of each department are simply added up linearly, it will inevitably lead to huge redundancy in planning results.
"Gaming" Nature of Demands: In the context of limited resources, to ensure the smooth progress of their own projects, departments tend to "exaggerate" their demands, reporting more and earlier, using computing demands as a bargaining chip in inter-departmental resource gaming. This plants the seeds for artificial "demand bubbles."
It is precisely these challenges that determine that demand consolidation cannot be a simple "homework collection" exercise but must be an active, in-depth, structured management process.
Building a Normalized, Structured Demand Consolidation Mechanism
To address the above challenges, we have established a normalized, structured demand consolidation mechanism, consisting of "two forms, one survey, and one set of processes."
"Two Forms": Standardized Collection of Demands
We designed two standardized online forms as the unified entry points for demand collection, aiming to guide and regulate how users propose demands.
The "Medium-to-Long-Term Computing Demand Planning Form":
- Collection cycle: Twice a year (mid-year and year-end), used to support annual planning and budgeting.
- Collection target: Facing planning or technical leaders of provincial companies and headquarters departments.
- Core content: Requires filling in a list of AI application projects that the unit plans to launch or deepen in the next 1-3 years. For each project, provide:
- Business background: The project's business objectives, expected value, and current stage (exploration, pilot, promotion).
- Application scenario description: Whether it is for image recognition, natural language processing, predictive analysis, or others.
- Preliminary technical solution: The planned model type (e.g., YOLO, BERT, Transformer), approximate parameter scale (e.g., hundreds of millions, billions, tens of billions), training data source and volume (TB, PB level).
- Computing demand estimate: Differentiate between training computing power (estimated required GPU card count, training duration, frequency) and inference computing power (estimated peak QPS, number of services). The form provides computing estimation reference templates and auxiliary calculation tools for different scenarios to help users with preliminary quantification.
- Timeline: Estimated project launch and go-live dates.
The "Temporary/Emergency Computing Demand Application Form":
- Collection cycle: On-demand, immediate.
- Collection target: Facing all project leaders or developers.
- Core content: Focuses on short-term, specific computing demands with finer granularity. In addition to the above technical solution details, it emphasizes:
- Urgency and reason for the demand.
- Clear list of resource requirements: such as "needs 8 A100 GPUs, exclusive use for 7 days, XX TB of high-speed storage space."
- Clear delivery time requirement.
Through these two forms, we guide and transform originally vague, qualitative business ideas into structured, semi-quantified technical demand descriptions, laying the data foundation for subsequent analysis and planning.
"One Survey": From Passive Reception to Active Mining
Relying solely on online forms may still miss a large number of "hidden demands." Therefore, we organize at least one specialized "AI Application and Computing Demand" survey led by the headquarters intelligent computing operations team, going deep into the frontline each year.
Survey format: Combining "online questionnaire census + offline in-depth interviews."
Survey content:
- Survey existing applications: Comprehensively inventory the current status of AI applications that have been launched or are being developed by various units, including their technical architecture, resource usage, bottlenecks encountered, etc. This is not only for collecting demands but also for "taking stock" of existing computing usage — a key step in "guiding demand submission." We use the "Computing Resource Usage Analysis Report" (generated by the operations platform) for each unit as input for interviews, jointly analyzing the reasonableness of their resource utilization with users and correcting deviations in their demand submissions.
- Explore potential demands: Conduct in-depth interviews with business experts from core departments (such as dispatch, inspection, marketing, safety supervision), informing them about the latest AI technology capabilities and application cases, and jointly exploring "What new problems can AI solve for us?" and "Which long-standing business pain points now have new technical solutions?" This process is an active "demand stimulation," transforming business pain points into future AI projects and computing demands.
- Collect feedback and suggestions: Listen to frontline users' opinions and suggestions on the existing computing platform and services as important references for platform optimization and planning adjustments.
"One Set of Processes": Ensuring the Closed Loop and Authority of Consolidation Work
- Regular release and promotion: The headquarters regularly (e.g., quarterly) releases the "Notice on Computing Demand Consolidation Work," specifying the collection period, scope, requirements, and submission entry for this round of demand collection, forming a stable work rhythm.
- Hierarchical review and consolidation: The provincial operations team is responsible for the preliminary collection, review, and consolidation of the unit's demands, ensuring the accuracy and completeness of the information, then uniformly reporting to headquarters.
- Headquarters centralized review and confirmation: The headquarters intelligent computing operations team, together with relevant business and technical experts, organizes a demand review meeting to centrally review the reported major demands, evaluating their business necessity, technical feasibility, and demand reasonableness, removing "water" and identifying key points.
- Form a unified demand list: After review, form a company-level "Computing Demand Resource Summary Table." This list is not only the basis for total computing capacity planning but also a reference for subsequent resource allocation priorities.
- Feedback and communication: Provide feedback on the review results to the demand units. For demands not approved or adjusted, give clear reasons and suggestions.
Through this systematic method of "coming from business," we are able to depict a comprehensive, true, and dynamic "heat map" of the company's AI applications and "trend chart" of computing demands, providing solid and reliable input for the scientific planning of "going back to business."
Capacity Planning: Scientific Forecasting and Precise Investment
Demand consolidation answers "what we need," while capacity planning answers "what we should build" and "how much to invest." This is a complex process of "translating" business demand language into technical configuration language and financial budget language. Its core challenge is how to make relatively optimal planning decisions under the uncertainty of rapid technological iteration and dynamic demand changes, achieving a balance between technological advanced, business fit, and investment economy.
The Methodology of Capacity Planning: The Art of "Translating" from Demand to Configuration
Our capacity planning is not simple linear superposition but a multi-dimensional, multi-level comprehensive assessment method.
Demand Classification and Priority Ranking
First, we classify and rank the vast amount of consolidated demands based on their nature and maturity.
- By business criticality: Divided into core critical business (such as grid security and stability analysis), important business (such as marketing anti-theft detection), and general business.
- By project maturity: Divided into applications already in production and promotion, applications in pilot verification stage, and applications in preliminary research and exploration stage.
- By computing type: Clearly differentiate between training computing power and inference computing power. Generally, training computing power has high requirements for single-card performance and parallel computing efficiency and can tolerate some queuing; inference computing power has high requirements for latency and concurrent processing capability and needs to ensure service real-time performance.
Based on the above classification, we assign a priority to each demand. For example, the expansion demand for inference computing for a core business related to grid security that is already online will certainly have higher priority than a training computing demand for a non-core business in the research stage. This priority ranking is the basis for subsequent resource allocation and investment decisions.
Computing Benchmark Testing and Selection
"Card-hours" or "FLOPS" are only theoretical peak values of computing power. When running real AI tasks, the effective computing power of GPUs with different architectures and models varies tremendously. Therefore, before large-scale procurement and planning, establishing a standard computing benchmark testing system is crucial.
- Build a typical model library: We select the most widely used and representative types of AI models within the company (such as YOLOv5/v7 for the transmission field, BERT for the marketing field, GNN for the dispatch field) to form our "benchmark model library."
- Conduct horizontal performance evaluation: We regularly obtain the latest products from mainstream GPU manufacturers on the market and conduct horizontal performance evaluations using the benchmark model library under a unified software and hardware environment, focusing on key indicators such as training throughput, inference latency, and energy efficiency ratio (performance/watt).
- Form "Computing Selection Guidance White Paper": Based on evaluation results, combined with factors such as procurement cost, supply chain security, and ecosystem maturity, we form a dynamically updated "Computing Selection Guidance White Paper," providing scientific and objective decision-making basis for headquarters and various units in their procurement, avoiding blind pursuit of the newest or being locked into a single vendor.
Computing Demand Conversion and Quantification
This is the most core and technically intensive step in capacity planning. We need to accurately "translate" upstream business demands into downstream hardware and software configuration lists.
Training Computing Conversion:
- Formula estimation: We have established an estimation model based on model parameters and training data volume. For example, for dense GPT-like Transformer models, the commonly used industry rule of thumb (traceable to OpenAI's analysis of GPT-3 training compute, i.e., total compute on the order of
6 x model parameters x training data tokens) gives the order of magnitude of compute required for training. Based on this formula, combined with the target training cycle, we can inversely deduce the required number and model of GPUs (based on their effective FLOPS performance). This formula is an order-of-magnitude estimate rather than a precise prediction, and requires separate correction for newer architectures such as MoE. - Empirical benchmarking: For some mature application scenarios, we refer to industry or internal best practices. For example, "training a 10-billion-parameter BERT model using 8 A100s takes approximately XX days."
- Reserve redundancy: Considering uncertainties such as algorithm optimization and model retries, we typically add 15%-30% redundancy to the estimation results.
Inference Computing Conversion:
- Stress testing method: This is the most reliable method. We perform stress testing on the model to be deployed in a benchmark environment, measuring the maximum QPS that a single GPU card can support while meeting the business SLA (such as P99 latency < 100ms).
- Matching demand and capability: Divide the total QPS demand proposed by the business department by the per-card QPS carrying capacity to obtain the required number of GPU cards. Similarly, high availability (such as N+1 redundancy) and future business growth should be considered, reserving a certain capacity buffer.
Supporting Resource Conversion:
- Storage planning: Based on the volume and performance requirements of training data, model files, logs, etc., plan the capacity and bandwidth of high-performance parallel file systems (for training), object storage (for data archiving), and local SSDs (for inference caching).
- Network planning: The network of the training cluster (especially multi-node multi-GPU training) is key to performance. Based on the cluster scale and GPU communication bandwidth, plan the topology and bandwidth of the computing network (such as InfiniBand or RoCE). Additionally, plan the storage network and management network.
Forming Standardized Planning Outputs
After the above series of scientific analyses and conversions, we finally form two core planning output documents to guide the investment and construction of headquarters and provincial units.
"Computing Planning Standards"
This is a programmatic document that does not involve specific hardware and software lists but defines the principles, standards, and methods the company should follow in computing planning and construction.
- Technology roadmap principles: For example, adhere to open architecture, embrace cloud-native, training and inference resource pools physically separate but logically unified, etc.
- Configuration standards: Define standard configuration models for different scales and levels of intelligent computing centers (such as minimum construction unit, GPU-to-CPU ratio, storage-to-computing ratio, network convergence ratio, etc.) for reference by various units.
- Cost measurement model: Define the company's unified TCO (Total Cost of Ownership) measurement model, including all cost elements such as hardware procurement, software licensing, data center, electricity, and personnel operations.
"Computing Planning Hardware and Software List" and Investment Cost Details
This is a specific executable list, the direct basis for annual budget and procurement. Listed by unit and project: Details of computing resources that need to be added or expanded for headquarters and provincial companies in the upcoming fiscal year to support specific projects.
- Detailed hardware and software configuration:
- Computing servers: Model, CPU configuration, memory, local storage, GPU model and quantity.
- Storage devices: Type (distributed file storage/object storage), capacity, bandwidth, IOPS.
- Network equipment: Switch model, port speed, quantity.
- Supporting software: License fees for operating systems, virtualization software, container platforms, scheduling software, AI frameworks, monitoring tools, etc.
- Clear investment cost details: For each hardware and software item, list the estimated unit price, quantity, and total price, aggregating to form the annual computing investment total budget for each unit and even the entire company. This list serves as the core material for applying for budget from the company's finance department.
Through this scientific planning method from qualitative to quantitative, from business to technology, we ensure that every computing investment has "a clear purpose and is well understood," achieving the precise investment goal of "good steel used on the blade."
Elastic Expansion: Moderate Advancement and Rapid Response
The market and technology are accelerating their changes. Any one-time, static planning cannot solve problems once and for all. Elastic expansion is precisely the dynamic adjustment mechanism designed to address this uncertainty. It pursues an art of balance: both "moderately forward-looking" to reserve space for future business explosions, avoiding missing opportunities due to insufficient resources, and "rapidly responsive" to efficiently complete resource expansion in an orderly manner when unexpected urgent needs arise.
The "Moderate Advancement" Strategy: Establishing Resource Level Warning Lines
Just-in-Time expansion is not feasible in the field of computing, where hardware procurement cycles are long and deployment is complex. We must establish a set of early warning and expansion trigger mechanisms based on resource levels.
Define resource level model: We treat the entire company's logical computing resource pool as a "reservoir." "Total capacity" is the reservoir's design capacity, and "allocated capacity" is the current water level. We define several key warning water lines:
- Safety level (e.g., 70%): When allocated capacity reaches 70% of total capacity, the system triggers a "yellow warning." At this point, the next round of regular expansion planning and budget application process should be initiated.
- Danger level (e.g., 85%): When the level reaches 85%, the system triggers an "orange warning." At this point, the regular expansion process must be accelerated, and evaluation of using reserved resources or adopting temporary solutions should begin.
- Limit level (e.g., 95%): Triggers a "red warning." Resources are about to be depleted, possibly affecting the acceptance of new businesses. Emergency expansion plans must be activated.
Headquarters computing resource reserve: To address network-wide sudden demands and strategic tasks, the headquarters intelligent computing center consciously reserves a portion of unallocated "strategic reserve force" computing resources during planning. These resources are not allocated to daily business under normal circumstances and are only launched by headquarters unified decision when major emergencies occur.
The "Rapid Response" Mechanism: Standardized Processes and Problem-Driven Approach
When the expansion demand (whether planned or urgent) is clear, an efficient mechanism must be in place to ensure its rapid and smooth implementation.
Formulate "Standard Process for Computing Expansion": We decompose the expansion process into a series of standardized phases and tasks, clearly defining the responsible department, inputs/outputs, and time nodes for each link.
- Phase 1: Technical solution design (computing operations team, architects)
- Phase 2: Procurement and tendering (material procurement department)
- Phase 3: Machine room preparation and site survey (infrastructure operations department)
- Phase 4: Equipment arrival and racking (vendors, operations team)
- Phase 5: System installation and deployment (operations team)
- Phase 6: Integration testing and stress testing (operations team, vendors)
- Phase 7: Acceptance and resource pool integration (operations team)
By standardizing the process, multiple tasks can be advanced in parallel, and the progress of the entire expansion project can be clearly tracked.
Compile "Computing Expansion Guidance Manual" and Typical Plans: To help inexperienced provincial companies complete expansion efficiently and with high quality, the headquarters led the compilation of the "Computing Expansion Guidance Manual." The manual not only contains standard processes but also provides typical expansion plans for different scales (such as one cabinet, one cluster), including recommended hardware configurations, network topology diagrams, deployment scripts, test cases, etc. This is like providing "renovation showrooms" — units can "move in with just a suitcase" or make slight modifications, greatly reducing the technical threshold and implementation cycle of expansion.
Establish expansion problem collaborative resolution mechanism: Various problems inevitably arise during the expansion process, such as equipment compatibility, software bugs, and substandard performance. We have established a collaborative mechanism centered on the "Expansion Problem List Summary Table."
- Problem-driven: All problems encountered during expansion must be recorded in this list, clearly describing the problem, responsible party (vendor/internal team), priority, and expected resolution time.
- Regular "consultation": The headquarters intelligent computing operations team regularly (e.g., weekly) holds expansion coordination meetings, inviting all relevant internal teams and vendors to participate, going through the open items in the problem list one by one, coordinating resources, clarifying action plans, and driving problem resolution.
- Closed-loop management: After a problem is resolved, the solution and verification results must be recorded in the list, forming a closed loop. These records also become an important part of the knowledge base.
Through the "moderate advancement" early warning mechanism and the "rapid response" standardized execution system, we have transformed computing expansion from a passive, firefighting-style emergency response into an active, planned, and controllable engineering management activity.
Resource Integration: Achieving Unity and Intensive Use of Existing Assets
In the intelligent transformation process of any large enterprise, there will inevitably be a large number of "legacy" computing assets. These assets may have been purchased by different departments, different projects, at different times, before unified planning. They are of different brands, architectures, and standards, like scattered "pearls" unable to form a synergy. Resource integration is the process of stringing these scattered "pearls" together, conducting unified "inventory, assessment, connection, and transformation," and ultimately integrating them into the unified resource pool. This is a complex but highly valuable "existing stock reform" effort.
The Dilemma of Existing Resources: Why Must Integration Be Done?
- Resource waste: Existing resources are often bound to their original specific projects. After the project ends or the load decreases, these resources become idle, but other departments cannot use them.
- High management costs: Each existing resource requires an independent operations team, monitoring system, and management process, resulting in high personnel and management costs.
- Technology silos: Heterogeneous technology stacks make it difficult to migrate and share applications and data between different resource pools, hindering collaborative innovation.
- Security risks: Lack of unified security baselines and management policies, these "floating" assets may become shortcomings in enterprise information security.
Therefore, the integration of existing resources is an inevitable choice for realizing the "integration" vision, improving overall resource utilization, and reducing TCO.
The Path of Integration: Step-by-Step Implementation, from Easy to Difficult
Resource integration is a systematic project that cannot be accomplished overnight. We have formulated a four-step work plan: "comprehensive survey, formulate standards, batch connection, continuous optimization."
Step 1: Comprehensive Survey, Know Your Inventory
We first need to conduct a thorough "census" of the company's existing AI computing resources.
- Survey current status: Through questionnaires, interviews, automated scanning, and other means, comprehensively collect detailed information on the existing resources of various units, including:
- Hardware information: Server manufacturer and model, CPU/memory/disk configuration, GPU model/quantity/memory.
- Software information: Operating system, virtualization technology, container platform, AI framework version.
- Network information: Network architecture, bandwidth.
- Management information: Owner department, operations responsible person, core applications hosted, current load status.
- Form "Computing Resource Survey Analysis Report": Aggregate and analyze the survey data to form an overall view of existing resources. The report points out the total amount, distribution, utilization status, degree of fragmentation of technical architecture, and the main challenges and opportunities that integration will face.
Step 2: Formulate Standards, Define Path
Based on the survey results, we need to formulate a clear "Computing Integration Standard and Work Plan."
- Define technical standards for integration:
- Minimum hardware requirements: Define the minimum hardware configuration standard for what can be integrated. For equipment that is too old or has too low performance, suggest direct retirement.
- Unified operating system and kernel version.
- Unified containerization base: Clearly require that all integrated computing nodes must install a unified version of Docker/Containerd and Kubernetes Agent (Kubelet).
- Unified GPU driver and device plugin: Require installation of unified device management plugins such as NVIDIA GPU Operator, enabling K8s to identify and schedule GPUs.
- Design multiple integration modes:
- Full Control: For resources that meet standards and have low load, reinstall and standardize, fully integrating into the headquarters' unified Kubernetes cluster, achieving unified resource scheduling. This is the ideal mode.
- Federation: For some clusters with special security requirements or that cannot be fully standardized, use K8s Federation or similar technologies to achieve logical unified management and cross-cluster task distribution while retaining their independence.
- Monitoring Only: For resources too old or heterogeneous to transform, the minimum is to uniformly connect their core monitoring indicators (such as GPU utilization) to the headquarters' monitoring platform, at least achieving "visibility" and providing data basis for subsequent retirement and replacement.
- Formulate work plan: Based on the status of existing resources and transformation difficulty of each unit, formulate a batch-by-batch, phase-by-phase work plan, clearly define the goals, responsible units, and timeline for each phase, achieving "integrate one batch as soon as it's ready."
Step 3: Batch Connection, Steady Progress
Follow the work plan, establish joint working groups with relevant units, and steadily advance the integration implementation.
- Application migration assessment: Before integration, assess the applications hosted on existing resources, formulate detailed migration or transformation plans, ensuring smooth business transition.
- Standardization transformation: The headquarters provides standardized deployment scripts and tools to guide or assist units in completing the reinstallation of operating systems, deployment of containerization bases, and other transformation work.
- Connection verification: After node transformation is completed, add them to the unified cluster and conduct a series of functional and performance verification tests, ensuring they meet expectations.
Step 4: Continuous Optimization, Realize Value
Integration is not an endpoint but the beginning of renewal. The integrated resources, merged into the unified resource pool, will undergo unified operations management and scheduling like other resources, and their usage efficiency and value will be greatly improved. At the same time, we will continuously track the operational status of these integrated nodes and, based on technological development, continuously optimize and upgrade them.
Chapter Summary
Chapter 3 has fully elaborated the complete picture of computing planning and capacity management. Starting from demand consolidation, we ensure the input for planning is true and comprehensive through systematic mechanisms; then, through scientific capacity planning methods, we accurately transform business demands into technical configurations and investment budgets; next, we design an elastic expansion system to address future uncertainty; finally, through resource integration, we revitalize existing assets. These four links together form a closed-loop management system from strategy to execution, from incremental to existing, ensuring that State Grid's computing "arteries" can synchronously strengthen and coordinatedly pulse with the "heart" of intelligent business, providing continuous core power for the digital transformation of the entire enterprise.