chinese foundational models trained completely on chinese chips
Chinese AI labs have achieved a significant milestone by successfully training foundation models entirely on domestic Huawei Ascend chips. Zhipu AI has emerged as the leader in this push for technological self-sufficiency, with several other major players building the broader ecosystem.
Below is a summary of the key models and the infrastructure that supports them.
🏆 Fully Domestic Models: The Leading Examples
A few foundational models have now been publicly confirmed as trained entirely on Chinese chips, primarily using Huawei's Ascend series. Zhipu AI has been at the forefront of this movement.
| Model Name | Developer | Type / Size | Domestic Hardware & Software | Significance |
|---|---|---|---|---|
| GLM-Image-1-4-6 | Zhipu AI | Multimodal (Image Generation) | Hardware: Huawei Ascend Atlas 800T A2 serversFramework: Huawei MindSpore | First major multimodal model trained entirely on a domestic stack; topped Hugging Face's trending leaderboard-1-6. |
| GLM-5-5 | Zhipu AI | 744B parameters (MoE architecture) | Huawei Ascend chips | A massive, open-source model proving domestic chips can scale to frontier-level LLMs; MIT license-5. |
| Qianfan-VL-9 | Baidu Smart Cloud | Vision-Language (3B, 8B, 70B) | Baidu's in-house Kunlun P800 chip | Demonstrates an alternative to Huawei, using Baidu's own chips for full-stack domestic training-9. |
🏗️ The Ecosystem: Building the Domestic Infrastructure
Single model releases are just one part of the story. A broader infrastructure is rapidly developing to support this transition.
- Huawei's Full-Stack Solution: Zhipu's success is built on using Huawei's complete ecosystem, which includes Ascend AI processors, MindSpore (the AI framework), and CANN (the computing architecture)-1-4. This integration is key to achieving competitive performance.
- National Support Platforms: The Chinese government is actively supporting this shift. An open innovation platform led by China Mobile, called "AI Huanxin," is designed to provide state-owned enterprises and other companies with access to domestic chips and models to accelerate development-3.
- Optimization from Within: Companies like SenseTime and research teams at Huawei are not just using the chips but are also developing advanced parallel training and low-precision (FP4) techniques to close the efficiency gap with Western hardware-2-10.
In summary, Chinese labs have definitively proven it is possible to train state-of-the-art foundation models using only domestic chips. The focus has now shifted from a question of "if" to one of "how efficiently" they can scale this infrastructure.
Are you more interested in the technical details of how they optimize training on these domestic chips, or would you like a deeper look at the performance benchmarks of these models against their NVIDIA-trained counterparts?