In the previous chapter, we addressed how to allocate computing power "fairly" and "efficiently." But this is only the first step of "managing well." A successful national intelligent computing center cannot be content merely with being an efficient "computing power plumber," ensuring smooth and equitable flow through the pipes. It must reckon with a deeper question: what exactly is flowing through those pipes?
If we remain at the level of providing Infrastructure as a Service -- selling bare computing power -- we are essentially "selling flour." Clients buy computing power (flour) from us, but still need to find their own recipes (algorithms), chefs (AI engineers), and kitchen (development environment), enduring a long and expensive process before they can produce the "bread" (AI application) they want.
Under this model, the intelligent computing center occupies a low position in the value chain. Its product is highly commoditized -- computing power itself offers little differentiation -- so competition revolves around price and energy consumption, squeezing margins. More importantly, the center has almost no control over what the computing power is used for or what value it generates, nor can it share in the higher added value that results.
To break out of this pattern and achieve a value leap, the national intelligent computing center must undergo a deep revolution in its business model and operational philosophy: upgrading from "selling flour" to "selling bread," and even "opening a bakery."
This is the core concept of this chapter -- building a "Model Supermarket."
The essence of a "Model Supermarket" is MaaS (Model as a Service). The center is no longer merely a computing power provider, but a "central kitchen" and "brand chain store" that curates and delivers high-quality AI models.
4.1 MaaS: From Computing Power Wholesaler to Value Co-Creator
Imagine a local mid-sized manufacturing enterprise that wants to use AI for product quality inspection, replacing manual labor.
Under the traditional "selling flour" model, the path is:
- Apply to the center to rent GPU servers.
- Hire an expensive AI algorithm team.
- Spend months collecting data, selecting algorithms, setting up environments, and training a model.
- Engage in constant trial and error and parameter tuning, consuming vast computing power and time.
- After the model is preliminarily usable, go through another lengthy deployment, integration, and optimization process.
The cycle may take a year or more, cost millions, and remain fraught with uncertainty. For the vast majority of SMEs, this threshold is dauntingly high.
Now see how the path changes under the "Model Supermarket":
- The manager logs into the local center's "Model Supermarket" platform.
- In the search box, they type "industrial quality inspection," and the platform displays a dozen related pre-trained models.
- They view detailed introductions: which university or company developed it, accuracy on public datasets, industries it has been applied to, and cost per invocation.
- They select a suitable "general industrial quality inspection foundation model" and click "Try it."
- The platform provides a simple interface for uploading a small number of representative product images from their factory.
- Guided by the platform, they fine-tune the model on their own small sample data. This takes just a few hours, consuming very little computing power.
- Once fine-tuning is complete, the platform generates a customized quality inspection API endpoint.
- The enterprise's IT staff embed this API into their production system. Camera images are automatically sent to the API, and the model returns a "pass" or "fail" judgment within milliseconds.
- The enterprise pays a monthly service fee based on API calls -- far less than the cost of building their own team.
Comparing the two models, the difference is revolutionary.
The "Model Supermarket" dramatically lowers the barrier to AI. It lets thousands of SMEs enjoy AI's efficiency gains in a "plug-and-play" manner. Its significance for regional digital transformation cannot be overstated.
4.1.1 The Supermarket Shelves: Three Tiers of Products to Meet Different Needs
A successful "Model Supermarket" must offer products that are rich and clearly stratified, like a real supermarket with basic commodities, branded processed goods, and fresh local specialties.
We can divide the shelves into three tiers:
Tier 1 (Bottom Shelf): Foundation Large Models These are the "staple goods" of the "Model Supermarket": ultra-large pre-trained models with general capabilities, trained by the national team, top research institutions, or tech giants on massive general data. Examples include:
- General language models: Similar to GPT-4, Pangu, or ERNIE Bot, with strong natural language understanding, generation, translation, and summarization.
- General vision models: Capable of recognizing and understanding objects, scenes, and actions in images and videos.
- Multimodal models: Capable of processing text, images, sound, and other information simultaneously.
These are the "public goods" the center must provide, serving as the "technology base" for all downstream industry innovation. The center should cooperate strategically with their developers, or deploy and optimize open-source models on its own computing power, so local users can access these advanced capabilities at low cost and high efficiency. Selling them is not about profit, but about securing an "ecological niche."
Tier 2 (Middle Shelf): Industry Fine-Tuned Models These are the "branded prepared foods" and "semi-finished dishes," built by fine-tuning foundation models on specialized industry data for vertical domains. Examples include:
- Financial risk control: Fine-tuned from a language model on regulations, credit reports, and transaction data to detect fraud in loan applications.
- Medical imaging diagnosis: Fine-tuned from a vision model on large-scale CT and X-ray image data, with the goal of helping physicians flag suspected early lesions for physician review.
- Legal contract review: After fine-tuning, able to quickly read and analyze a hundred-page contract and flag risky clauses.
These models are key to the center's value creation. The center should act as an "incubator" and "connector." On one hand, it should leverage its computing resources and data advantages -- while ensuring security and compliance -- to co-develop these models with local leading enterprises, universities, and hospitals. On the other hand, it should introduce mature external providers and invite them to "list" their models in the supermarket.
Tier 3 (Top Shelf): Localized Application Models These are the "local snacks" and "community group buys": lightweight, customized models for specific local needs and scenarios. Examples include:
- Local tourism guide: Trained on local history, culture, and attractions, conversing with tourists in the local dialect.
- Local government service: After learning local service guides, answering citizens' questions 24/7 about "how to apply," "where to go," and "what to bring."
- Local crop pest and disease identification: Tailored to local crops and pests, enabling photo-based diagnosis and control recommendations.
These models best embody the center's service to and rooting in the local community. They may not be the most technically advanced, but they are the most "humane" and best at solving local problems. The center should encourage local developers and students to create these "small but beautiful" applications through AI competitions and by providing "computing vouchers" and "data sandboxes."
4.1.2 The Manager's Role Shift: From "Property Manager" to "Shopping Mall General Manager"
To operate a "Model Supermarket" well, the manager's role must evolve. We are no longer a "property manager" maintaining the building's electricity, water, and HVAC. We must become a strategic "shopping mall general manager."
The core work of a "shopping mall general manager" is no longer about server CPU utilization, but about the "efficiency per square meter" and "customer traffic" of the "mall."
- Investment attraction (model onboarding): The general manager needs a keen eye to judge which models are "rising stars" and which are "established brands," then actively negotiate and bring them in. We need clear "model listing standards" to evaluate models from multiple dimensions: technical advancement, security, compliance, and commercial prospects.
- Brand promotion (model marketing): Even great models need promotion. The general manager should organize events -- industry solution launches, developer salons, case-sharing sessions -- to promote "star products" to local enterprises.
- Category management (model optimization and iteration): Closely monitor "sales data" (invocation volumes, user feedback) and conduct regular "inventory checks." Invest more in popular models; decisively remove outdated ones.
- Customer service (model application support): The center needs a professional "model application consultant" team to help clients -- especially SMEs -- understand model scenarios and provide support for fine-tuning and API integration.
- Rule-making (platform governance): Establish fair, transparent "rules of the game" covering pricing, revenue sharing, data privacy, security, and intellectual property. A healthy ecosystem depends on sound governance.
Through this shift, the center transforms from a passive, cost-driven infrastructure provider into an active, value-driven ecosystem builder. Revenue sources expand from a single "computing power rent" to diverse streams: "model invocation fees," "solution consulting fees," and "ecosystem revenue sharing." This is the path to self-sustaining, sustainable development.
4.2 Avoiding "Evaluation Alienation": Guarding Against the Trap of "PPT Models"
As the "Model Supermarket" shelves grow richer, a new governance challenge emerges. How do we judge whether a model is truly "good" or merely "looks good on paper"?
In the current AI field, there is a dangerous tendency -- "evaluation alienation," or "leaderboard-chasing culture."
Many AI model releases come with impressive "battle reports": on internationally authoritative public datasets, the model has achieved "State-of-the-Art" (SOTA) results, setting new records.
These "leaderboards" and "scores" are valuable in academic research and technical communication. They provide an objective yardstick for measuring algorithmic progress.
However, when these "scores" become the sole criterion for evaluating model value, "alienation" occurs. Some developers no longer aim to solve real-world problems, but instead do whatever it takes to "climb the leaderboard."
This spawns "PPT models" or "benchmark-chasing models." Such models:
- Overfit: They are "pixel-level" optimized for specific datasets, performing stunningly on them but collapsing on slightly different real-world data.
- Lack generalization: They perform well in the lab but fail under complex network conditions, hardware constraints, and real-time requirements.
- Ignore total cost: To improve accuracy by 0.1%, they adopt complex architectures that make inference costs prohibitive.
- Contain bias: Multiple studies of mainstream facial recognition datasets and commercial systems (e.g., the Gender Shades audit of error rates in commercial facial analysis systems) document pervasive ethnic and gender sampling bias in public datasets; high-scoring models trained on them inherit and amplify these biases.
If our "tenant recruitment" criteria are about whose "PPT" looks best and who ranks highest on leaderboards, the supermarket will be filled with exhibits rather than products that create value.
The consequences are devastating: when users excitedly purchase a "world champion" model only to find it cannot be deployed, they lose confidence in the center and AI technology itself.
4.2.1 Establishing New "Combat-Oriented" Evaluation Standards
To avoid the trap of "evaluation alienation," we must establish a new set of model evaluation standards centered on "actual combat effectiveness." This standard no longer asks, "What score did your model achieve on dataset X?" but rather, "What problem did your model solve in a real scenario, and what value did it deliver?"
This means our evaluation metrics must shift from academia's "intermediate metrics" (such as accuracy and F1-score) to industry's "outcome metrics."
Let us examine some concrete examples:
| Domain | Traditional "Leaderboard" Metric | Combat-Oriented Evaluation Metric |
|---|---|---|
| Smart Transportation | Vehicle recognition accuracy | Average reduction in regional traffic congestion index, minutes saved in peak commute time, year-over-year percentage reduction in traffic accident rate |
| Smart Healthcare | AUC value for tumor lesion detection | Percentage reduction in doctors' average image reading time, actual reduction rate of early cancer misdiagnosis, consistency between model-assisted diagnosis and expert consultation |
| Industrial Quality Inspection | Precision/Recall for defect classification | Actual percentage point increase in product yield rate, reduction in unit product manual inspection cost, reduction in customer return rate due to quality issues |
| Financial Risk Control | F1-score for fraudulent transaction detection | Actual percentage point reduction in credit card bad debt rate, fraud amount intercepted per million transactions, reduction in complaint rate from legitimate users due to misjudgment |
| Government Services | Intent recognition accuracy | Reduction in citizens' average waiting time, online completion rate of "one-stop service" items, reduction in transfer rate to human operators on government service hotlines |
The contrast in this table makes the core idea unmistakably clear: what we evaluate is no longer the model's technical parameters, but its "transformation efficiency" and "value contribution" to real-world business processes as a productivity tool.
4.2.2 A Manager's Action Guide: Implementing "Combat-Oriented Evaluation"
To truly put this "combat-oriented" evaluation standard into practice, managers need to undertake a series of institutional designs and process re-engineering efforts:
1. Establish a "Model Proving Ground"
Alongside the "Model Supermarket," we must build a "proving ground." Any model hoping to be "listed," in addition to submitting its academic leaderboard results, must enter this proving ground and undergo a real-world "road test."
This proving ground should collaborate deeply with local government departments (such as the Transportation Bureau and Health Commission) and leading enterprises to obtain authorized and anonymized "localized real test datasets." A traffic model should not simply be scored on the public KITTI dataset; it must perform real-time vehicle recognition and traffic flow prediction tests on video streams captured by real, complex intersection cameras in our city, under various weather and lighting conditions.
2. Promote a "Try Before You Buy, Pay by Results" Model
For industry models entering the supermarket -- especially those claiming to deliver significant business value -- we can promote a more demanding business model. Allow enterprise clients to "try" the model for free or at very low cost for a certain period. After the trial, instead of paying per invocation, the fee is determined by the quantifiable business gains the model brought during the trial period (such as cost savings or increased revenue), through tiered pricing or revenue sharing.
This "pay by results" model is a powerful "filter." It deeply aligns the interests of model providers with those of end users. Only models that are true "doers" capable of creating real value can survive and profit under this arrangement; benchmark-chasing "theorists" will be mercilessly eliminated by the market.
3. Build a "Red Team/Blue Team Evaluation Team"
In addition to routine performance testing, the intelligent computing center should assemble a professional "Red/Blue" team.
- Blue Team (Defense): Represents the model provider, responsible for articulating the model's advantages, performance, and applicable scenarios.
- Red Team (Attack): Composed of the intelligent computing center's security experts, industry consultants, and potential client representatives. Their job is to act as a "devil's advocate," "attacking" the model from various challenging angles. For example:
- Robustness attacks: Intentionally input data with noise, blur, or adversarial perturbations to see whether the model's performance degrades sharply.
- Fairness audits: Check whether the model exhibits bias or discrimination against specific groups. For instance, does a facial recognition model have a significantly lower recognition rate for dark-skinned individuals compared to light-skinned individuals?
- Explainability questioning: When a model makes an important judgment -- such as rejecting a loan application -- can it provide a logical explanation that humans can understand? Or is it merely an inscrutable "black box"?
- Economic analysis: Comprehensively evaluate the model's "total cost of ownership," including the computing power required for inference, the personnel costs for deployment and maintenance, and so forth, to determine whether it is economically viable.
Only by passing such rigorous "Red/Blue team" stress tests can a model be awarded the "Intelligent Computing Center Select" certification label and receive priority recommendation in the "Model Supermarket."
4. Establish a "User Real Feedback and Rating System"
Finally, and most importantly, the ultimate power of judgment must be returned to the market and to users.
The "Model Supermarket" platform must establish a public and transparent user feedback system, similar to "Dianping" or "Taobao reviews." Any user who has invoked a model can rate and review it. The evaluation dimensions should not be limited to a simple "good/bad" dichotomy, but should be more specific, such as:
- Ease of deployment: Is the API interface clear? Is the documentation complete?
- Performance stability: Is the response still fast during business peaks?
- Cost-effectiveness: Is the service fee reasonable relative to the business value it brings?
- Technical support response speed
This massive, continuous stream of feedback from real users constitutes "live data" on model reputation. It reflects a model's comprehensive value far more truthfully than any static, one-time leaderboard score. The intelligent computing center's operations team should use user ratings and feedback as the most important basis for "inventory management" -- recommendation, sorting, and delisting -- of models.
Conclusion: Be a "Referee of Value," Not a "Clerk of Scores"
The core of Chapter 4 charts an essential path for the operational governance of the national intelligent computing center, from "resource provider" to "value enabler."
- In terms of business model, we must drive a profound transformation from IaaS to MaaS. Establishing a "Model Supermarket" upgrades the intelligent computing center from a primary raw material supplier "selling flour" to a "central kitchen" capable of providing "prepared dishes" and "set meals." This is not only a leap in the center's own value chain, but also a key move in democratizing AI benefits for thousands of small and medium-sized enterprises and igniting digital economy innovation across the entire region.
- In terms of value evaluation, we must resolutely resist the temptation of "evaluation alienation." We must recognize clearly: the national intelligent computing center is not an ivory tower of academia, but the main battlefield for serving the national economy and people's livelihoods. Our mission is not the vanity of chasing one or two "world records," but to use AI technology to tangibly solve real problems -- traffic congestion, unequal distribution of medical resources, low industrial efficiency, and more.
Therefore, the manager of the intelligent computing center must play the role of a "referee of value," not a passive "clerk of scores." We need to use the strictest yardstick of "actual combat effectiveness" to measure, screen, and cultivate good models that can truly "take root, blossom, and bear fruit."
Only when our "Model Supermarket" is stocked with "trustworthy products" that have passed real-world tests and can create genuine value for users can we truly say we have "managed" computing power well. Only then can we truly transform this powerful technological force into a reliable, constructive force for social progress.
This is not only being responsible to the taxpayers who have invested in us, but also fulfilling the historical mission entrusted to us by this great era in which we live.