China’s AI Models Are Catching Up—But Its Compute Gap Is Getting Worse
China’s artificial-intelligence industry is entering a more difficult phase.
For the past several years, competition has largely been measured through model benchmarks: parameter counts, reasoning performance, context length, training efficiency, and API pricing. Chinese developers have demonstrated that architecture optimization, mixture-of-experts models, quantization, distillation, and increasingly efficient training methods can narrow the performance gap with global leaders.
But the next stage of competition will not be decided by models alone.
As models become more capable, more widely used, and increasingly integrated into search, coding, document processing, data analysis, and autonomous-agent workflows, the underlying demand for computing resources is expanding faster than the efficiency gains delivered by software.
The central contradiction facing China’s AI industry is therefore becoming clearer:
Model capabilities and user demand are growing rapidly, while the supply of high-end GPUs, AI accelerators, HBM, advanced packaging, high-speed networking, and data-center power cannot expand at the same speed.
This is no longer a single-chip shortage. It is a system-wide supply-chain constraint.
The success of models such as Kimi K3 may intensify this pressure. Lower-cost and higher-efficiency models can rapidly expand market adoption, but broader adoption also produces substantially more inference demand. Once AI services move beyond simple question-and-answer interactions, each user may trigger multiple model calls, long-context retrieval, tool usage, memory operations, and multi-step reasoning.
The cost of an individual inference may fall, but aggregate token consumption can still rise dramatically.
As a result, competition in China’s AI market is shifting from “who can offer the cheapest model” toward two more difficult questions:
Who can secure stable computing resources?
And who can convert constrained hardware into the greatest amount of effective computing output?
China’s recent breakthroughs in large AI models have drawn worldwide attention. The discussion gained even more momentum this morning when “the founders behind China’s four leading AI model companies” began trending on Chinese social media. (Deepseek, Kimi, Zhipu, Minimax)
Their career paths reveal how the role of a technology founder is changing in the AI era. The job is no longer simply about writing the most code or personally solving every technical problem. Increasingly, the founder’s real value lies in setting the direction, assembling the right team and building a system in which 10% of their direct execution can unlock the other 90% of the company’s growth.
Beyond the AI model itself, people have also begun discussing why Yang Zhilin turned down an offer from Apple after completing his PhD in the United States and instead chose to return to China to start a company. The decision reflects a complex combination of personal, professional, and policy considerations.
Prominent investor Vinod Khosla attributed part of the issue to the Trump administration’s immigration policies. Responding to discussions about the success of Kimi K3, he argued: “The bigger issue is that our immigration policies targeting highly skilled talent are driving away exceptional people from other countries.”
People choose where to build their lives and careers for many different reasons. For the United States, however, growing uncertainty surrounding immigration and international academic and professional exchanges could make the country appear less welcoming—and ultimately weaken its long-term competitiveness.
SemiVision spotted an interesting chart on OpenRouter that clearly illustrates how the U.S.–China technology rivalry is evolving—from a competition over semiconductor chips to a competition over AI models.
Beyond Semiconductor Materials and Equipment, Should the U.S. Also Restrict Chinese AI Models?
The White House is reportedly examining whether leading Chinese AI models have benefited from U.S. large language model technology and is considering sanctions against overseas AI models found to infringe intellectual property. Officials are also discussing whether companies using Chinese AI models should be required to disclose that fact to their customers.
However, open-weight models present a fundamentally different challenge from semiconductor hardware. Once released, they can be freely downloaded, modified, and deployed locally, making them far more difficult to control through tariffs, export restrictions, or outright bans.
The reality is that Chinese AI models have already gained significant traction in the U.S. market. On the AI platform OpenRouter, roughly 60% of the tokens consumed by U.S. users reportedly come from Chinese models. Companies such as DoorDash and Airbnb have also adopted them as cost-effective alternatives to models from OpenAI and Anthropic.
A blanket ban on Chinese AI models could therefore create unintended consequences for Silicon Valley startups, research institutions, and AI infrastructure providers. Rather than imposing an outright prohibition, Washington is more likely to rely on legal liability, regulatory uncertainty, and reputational pressure to encourage companies to move away from Chinese models voluntarily.
China, meanwhile, has positioned universal AI access and open-weight ecosystems as part of its global strategy, arguing that every country should have the freedom to choose its AI technologies without being forced to align with either Washington or Beijing.
This marks an interesting reversal in the U.S.–China technology competition. The United States originally sought to keep China dependent on Nvidia GPUs and the broader U.S. technology ecosystem. Today, however, China’s combination of low-cost, highly capable AI models is creating a new form of dependence—this time among American enterprises.
For countries that lack the resources to develop frontier AI models on their own, Chinese AI offers an attractive combination of affordability and ease of deployment. Even if Washington increases political and regulatory pressure, it is likely to find it difficult to prevent these models from continuing to spread across global markets.
Because of U.S.–China semiconductor export restrictions, China continues to face limited access to NVIDIA’s latest AI chips. Existing Hopper-based GPUs remain in short supply, while next-generation GB-series platforms and their complete rack-scale systems are even more difficult to obtain in meaningful deployment volumes.
Building an AI Supercluster Requires Far More Than GPUs
GPUs are only one part of the equation. A production-scale AI supercluster requires a complete technology stack.
Beyond AI accelerators, it also depends on high-bandwidth memory (HBM), advanced packaging, high-speed network interface cards (NICs), Ethernet or InfiniBand switches, optical transceivers, server racks, power conversion systems, cooling infrastructure, and a stable, reliable electricity supply.
A bottleneck in any one of these components can delay the deployment of the entire AI cluster.
Even when Chinese companies succeed in acquiring AI accelerators, there can still be a significant gap between nominal computing capacity and effective computing capacity. Real-world performance depends on much more than the number of GPUs installed. It is determined by networking efficiency, software compatibility, memory constraints, hardware reliability, cluster scheduling, and workload utilization.
As a result, two AI clusters with the same theoretical teraFLOPS may deliver dramatically different levels of real-world computing performance.
Chinese AI developers have become highly skilled at extracting more output from constrained hardware.
Low-precision computing, quantization, model sparsity, knowledge distillation, speculative decoding, and optimized inference engines can all improve the number of tokens generated per unit of hardware.
According to Huawei, the system is scheduled to reach the market in the fourth quarter of 2026, marking China’s entry into the era of AI clusters built with more than 10,000 accelerator cards. Unsurprisingly, it became one of the biggest attractions at this year’s exhibition.
The upgraded supernode architecture stands out for three reasons.
-
First, Huawei has found a way to link as many as 1,024 AI chips into a single system. Together, they can deliver 1 exaFLOPS of FP8 performance and 2 exaFLOPS at FP4, making it one of the largest supernodes introduced by the industry.
-
Second, all 1,024 chips can access a shared memory pool of up to 256 TB. Rather than treating each accelerator as a largely isolated resource, Huawei is trying to make the entire cluster operate more like one enormous computing system.
-
The third advantage is communication speed. Huawei says the interconnect can move data at terabyte-per-second speeds, while reducing round-trip latency to just three microseconds. At that level, information can move between chips almost instantaneously, addressing one of the biggest problems in large AI clusters: the time accelerators spend waiting for one another instead of doing useful work.
These technological advances are strategically important. They extend the useful life of existing AI accelerators and reduce dependence on the latest generation of GPUs.
However, higher efficiency does not necessarily translate into lower infrastructure demand.
As AI inference becomes cheaper, more users adopt AI services. Existing users also consume more compute and run increasingly complex workloads. An AI agent, for example, may make dozens of model calls to complete a workflow that previously required only a single prompt.
This closely mirrors the idea Jensen Huang has repeatedly emphasized: when the cost of computing falls, demand rises even faster. Lower prices unlock entirely new use cases rather than simply reducing spending. AI inference follows the same economic principle. As token costs decline, developers and enterprises become willing to build more AI-native applications and call models far more frequently.
Token costs are still substantial today. SemiAnalysis Dylan Patel previously disclosed that it spent approximately US$7 million on tokens in 2025, and that expenditure continues to grow rapidly. As inference becomes more affordable, overall token consumption is likely to increase even faster rather than decline.
This is a classic example of the Compute Rebound Effect: reducing the cost of each inference ultimately stimulates even greater demand for total computing capacity.
This is why AI inference accelerators have become one of the industry’s most closely watched markets. Unlike general-purpose training GPUs, inference chips can be optimized for specific models, data formats, operators, and memory-access patterns.
It also explains why nearly every major cloud service provider and AI leader—including OpenAI, SpaceX, and Anthropic—is investing in custom AI inference silicon. Designing chips specifically for inference offers greater control over cost, performance, power efficiency, and infrastructure optimization than relying solely on merchant GPUs.
A successful inference processor does not need to replicate every capability of a cutting-edge training GPU. Instead, it must strike the optimal balance between compute performance, memory capacity, software compatibility, and energy efficiency for its target workloads.
Chinese AI companies are continuing to expand their use of computing resources in the Middle East, Southeast Asia, South Korea, and other overseas markets. Advanced chips are generally easier to obtain outside China, which is one of the main reasons many Chinese AI companies have established offshore entities and overseas operations.
These regions may offer access to more advanced GPUs, reliable electricity, available land, capital, and newly constructed data-center capacity.
For Chinese model developers, overseas infrastructure can shorten computing-resource procurement cycles and reduce the substantial capital expenditure required to build large domestic computing clusters.
Where data-compliance requirements and network latency allow, inference workloads can be distributed across overseas facilities. Model-training teams may also establish offshore operating environments that include data, model weights, computing infrastructure, and engineering personnel.
However, this strategy is facing growing regulatory and operational risks.
Export controls are gradually expanding beyond the direct sale of physical chips to cover resale, transshipment, end use, and access to cloud-computing resources.
Data-center operators may increasingly be required to verify the identities of their customers, their ultimate beneficial owners, and the final purpose of their computing workloads.
Recently, a group of servers equipped with Nvidia’s advanced AI chips was allegedly exported illegally to China and other destinations. The case involved approximately US$25 million, and the individuals involved were arrested by Taiwanese prosecutors, investigators, and the Coast Guard Administration. The case illustrates how US chip restrictions have expanded significantly and how scrutiny of advanced-chip exports has become exceptionally strict.
Therefore, obtaining computing resources through an overseas legal entity may not provide durable protection if the underlying customer or ultimate end use remains subject to restrictions.
From an operational perspective, overseas computing resources are also difficult to integrate fully with infrastructure located inside China.
Cross-border network latency, data-security regulations, model-weight protection, engineering collaboration, and regulatory compliance all add to operating costs.
Overseas computing capacity will therefore remain an important supplementary resource, but it is unlikely to become a complete or sustainable long-term substitute for domestic computing infrastructure.
China has demonstrated that innovations in model architecture, improvements in engineering efficiency, and software optimization can partially offset the disadvantages created by hardware restrictions.
However, software efficiency cannot fully substitute for physical infrastructure.
The more successful Chinese AI models become, the more urgently they will require AI accelerators, memory, advanced packaging, networking equipment, electricity, cooling systems, and data-center capacity.
To address these challenges, China will expand the use of Huawei Ascend and other domestic chips, develop inference-specific accelerators, build more supernodes, lease overseas computing capacity, expand domestic semiconductor fabs, localize HBM and advanced packaging, and improve software for heterogeneous computing.
These efforts will reduce China’s dependence on individual foreign suppliers, but they will not eliminate the technological and supply-capability gap in the short term.
Advanced-process capacity, HBM yields, packaging integration, high-speed interconnects, and software ecosystems are all capabilities that require years of cumulative development. These challenges cannot be resolved through a single product launch or the construction of one new fab.
SemiVision believes that China’s computing shortage will remain an important driver of semiconductor and power-infrastructure investment. One relevant example is China’s national “Eastern Data, Western Computing” initiative, a large-scale infrastructure project designed primarily to address the country’s geographic mismatch in resources.
China’s data demand is concentrated mainly in the eastern regions, where land, water, and electricity supplies are increasingly constrained. To address this imbalance, China has established eight national computing-hub nodes covering the Beijing–Tianjin–Hebei region, the Yangtze River Delta, the Guangdong–Hong Kong–Macao Greater Bay Area, the Chengdu–Chongqing region, Inner Mongolia, Guizhou, Gansu, and Ningxia.
Multiple national data-center clusters have also been designated within these hubs to strengthen coordination between the eastern and western regions.
The United States continues to lead in high-end AI accelerators, AI software, and system architecture. Taiwan remains central to advanced wafer manufacturing, advanced packaging, and server production. South Korea retains its advantages in HBM and advanced memory, while Japan continues to play an indispensable role in semiconductor materials, equipment, and critical components.
China, meanwhile, will leverage its enormous domestic market, policy support, manufacturing investment, and system-level innovation to build a more autonomous and controllable domestic computing ecosystem. China is particularly effective at bringing AI applications to market rapidly, with ByteDance’s Doubao serving as a strong example.
The global supply chain will not fully decouple, but access to the highest-performance products will become more tightly controlled. At the same time, regional manufacturing and second-source strategies will become increasingly important.
SemiVision believes that AI competition is no longer determined solely by model intelligence or the specifications of an individual chip.
The real competitive advantage increasingly lies in whether an industry can coordinate chip design, wafer manufacturing, memory, advanced packaging, networking, power, cooling, software, and data centers—and integrate these elements into a complete system that can scale efficiently.
The global AI-computing gap is not simply the absence of a particular flagship chip.
The real constraint lies in the ability of the entire supply chain—and the power infrastructure supporting it—to deliver sufficient computing capacity at the required speed, efficiency, yield, and scale.
This gap will be difficult to close in the short term.
Over the next several years, however, it will remain one of the most important drivers of global semiconductor and infrastructure investment.

















