Back to Chinese AI Companies
BL
Bhushan Lodha
5/27/20260 comments
DeepSeek

How are chinese labs training their models?

Chinese AI labs are training their models through a strategic pivot that compensates for restricted access to advanced NVIDIA chips by maximizing system-scale efficiency, algorithmic innovation, and domestic hardware utilization-5-10. The unified goal is to achieve competitive AI capabilities using what is available—primarily Huawei Ascend chips and a "more chips, more cleverness" philosophy.

Here is a summary of the three core strategies Chinese labs are employing:

StrategyCore ApproachReal-World Examples
1. Massive Cluster DeploymentLink thousands of less-powerful domestic chips (e.g., Huawei Ascend) to function as a single, powerful system, compensating for lower per-chip performance with sheer scale.Huawei's 10,000-card Ascend 910C cluster; Alibaba's 10,000-chip "Zhenwu" cluster-5.
2. Algorithmic & Hardware OptimizationRedesign model architecture (e.g., linear complexity, 1.58-bit quantization) and system software (e.g.,北大's parallel framework) to drastically cut computation and memory needs, boosting efficiency on domestic hardware.中科院 "瞬悉1.0" (linear-time), Tsinghua & 鹏城实验室 "开元-2B" (Sandwich Norm), OpenBMB BitCPM (ternary weights)-1-2-4.
3. Complete Domestic Stack IntegrationUse a full ecosystem of Chinese technology for entire model lifecycle—from training to deployment—ensuring autonomy and bypassing Western hardware dependencies.Z.AI's GLM-Image (Ascend + MindSpore); China Unicom's enhanced DeepSeek model[CITATION:3]-10.

🏗️ Strategy 1: Building "Mega-Clusters" with Domestic Chips

Since importing NVIDIA's most advanced chips is restricted, Chinese firms have shifted their focus from chasing the fastest single chip to building the largest possible clusters of domestic chips-5-10.

  • The 10,000-Card Barrier: Both Huawei and Alibaba have recently activated major clusters built around their own Ascend and "Zhenwu" AI chips, respectively-5. A 10,000-card cluster of Huawei Ascend 910C chips is designed to train models with hundreds of billions of parameters by operating as a single, massive system-5.
  • System Over Silicon: This approach directly acknowledges the U.S. export controls. The strategy is to compensate for lower per-chip performance through scale, advanced networking, and software innovation-10.

🧠 Strategy 2: Redesigning Models for Efficiency

To get the most out of domestic hardware, Chinese labs are also rethinking the AI models themselves to be far more efficient.

  • Pursuing Linear Complexity: Transformer models have a "quadratic" cost, meaning they slow down significantly as sequences get longer. The Chinese Academy of Sciences (CAS) tackled this by developing "瞬悉1.0" (SpikingBrain-1.0), a "linear-time" model architecture. This allows it to handle sequences up to 4 million tokens long with over 100x faster initial response times on domestic GPUs from a company called MetaX, not NVIDIA-1.
  • Extreme Compression with 1.58-bit Models: A team called OpenBMB pioneered a method called "BitCPM-CANN," which trains models using ternary weights (-1, 0, +1) on Huawei Ascend NPUs. This reduces memory usage by roughly 6x and achieves over 95% of the performance of a full-precision model-2.
  • Training Stability Tricks: Because older domestic chips like the Ascend 910A are less powerful and support only limited-precision calculations (FP16, similar to a V100), training can be unstable. Researchers from Tsinghua University and 鹏城实验室 (Peng Cheng Lab) introduced techniques like "Sandwich Norm" and "Soft Clipping" to stabilize training and prevent errors from derailing the process on this more challenging hardware-6.
  • Smarter Parallel Processing: At a systems level, Peking University developed a new "fine-grained" parallel training framework for Ascend chips. This more intelligently splits a model across different chips, leading to a 15% improvement in training efficiency by reducing idle time and improving communication-4.

🏛️ Strategy 3: Building the Full Domestic Stack

China is pushing for complete technological independence by developing and integrating every layer of the AI stack, from the chip up to the application.

  • End-to-End Ecosystems: The goal is to use Chinese chips (Huawei Ascend) with Chinese AI frameworks (MindSpore) to train and run models. This has been demonstrated by Z.AI, a blacklisted company, which trained its "GLM-Image" model entirely on Huawei's Ascend Atlas 800T A2 servers using the MindSpore framework-10.
  • Government-Led Initiatives: This push is not just by private companies. National labs like Peng Cheng are releasing fully open-source models trained "only on国产算力" (domestic computing power), creating blueprints and toolkits for others to follow-6. Meanwhile, Huawei's own PanGu-Σ model, a trillion-parameter giant, was trained on a cluster of 512 Ascend chips-7.

🔬 Practical Impact: Training vs. Inference

The real-world impact of these strategies varies depending on the task:

  • For Training Giant Models (The Hard Part): Training a massive model from scratch remains challenging. It requires huge clusters (like the 10,000-card ones from Huawei and Alibaba) and sophisticated engineering to keep them stable-5-7. However, labs are succeeding, proving domestic clusters can get the job done, albeit with more effort and hardware.
  • For Fine-Tuning & Inference (The Practical Part): This is where Chinese labs are finding easier success. China Unicom, for example, took the open-source DeepSeek-R1 model and fine-tuned it to be more efficient and secure for its own needs, without needing to train a new model from billions of dollars-3.

I hope this gives you a clear picture of how Chinese labs are innovating within their constraints. Are you more interested in the specific technical details of the "linear-time" model from CAS or the parallel training framework from Peking University?

DeepSeek
Source: US China GPU Data Center Capacity Comparison

Comments

No comments yet. Readers can leave comments directly from the expanded post on the board page.