Back to Chinese AI Companies
BL
Bhushan Lodha
5/26/20260 comments
DeepSeek

How are Chinese model able to serve such a huge token demand given they have limited hardware

Chinese AI companies have successfully turned the challenge of limited hardware into a major competitive advantage through relentless architectural and software innovation. They serve enormous token demand by being radically more efficient, not by having more powerful chips-5.

Here are the key strategies bridging the hardware gap:

⚙️ 1. Architectural Efficiency: The "Mixture of Experts" (MoE) Revolution

Instead of activating a whole model for every task, MoE models only use the "expert" sub-networks needed for a specific query.

  • How it works: A model like MiniMax M2.5 has 230 billion total parameters, but only activates 10 billion per token. This makes it as smart as a massive model but as agile as a small one-5.
  • The Result: This allows for ultra-low pricing. For instance, MiniMax M2.5 can cost as little as $0.15 per million input tokens, which is roughly 1/30th the price of comparable Western models-5.

📉 2. Model Distillation: Creating "Mini-Me" Specialists

Distillation is the process of training a smaller, faster "student" model to replicate the complex reasoning of a much larger "teacher" model-3.

  • The Strategy: Instead of running a massive, slow model for every request, companies create compact, specialized models (as small as 0.6B to 3B parameters) that run extremely fast on limited hardware-3-8.
  • The Benefit: These smaller models can achieve over 44,000 tokens per second on a laptop, making real-time AI on phones and edge devices possible without cloud dependency-8. Crucially, new techniques (like TCOD) solve the problem of "distilled" models failing on complex multi-step tasks, making them viable for real work-7.

🏭 3. Infrastructure & Hardware Innovation: "Good Enough" at Scale

Chinese firms cannot rely on high-end Nvidia H100s, so they have redesigned the software stack to maximize the chips they can get (like the restricted H800 or domestic Huawei Ascend chips)-1-5.

  • GPU Pooling (Alibaba's Aegaeon): Alibaba found that most models sit idle on dedicated GPUs. Their "pooling" system serves hundreds of models from a single pool, allowing them to reduce GPU usage by 82% (saving nearly 1,000 GPUs) while increasing output 9x-9.
  • Inference Optimization (Huawei's UCM): AI inference involves a massive memory bottleneck (KV Cache). Huawei's Unified Cache Manager (UCM) cleverly shuttles data between different types of memory (HBM, DRAM, SSD), resulting in 2-22x improvements in tokens-per-second in long tasks-6.

🏆 A Real-World Example: GLM-5

Z.ai's GLM-5 is a flagship model that demonstrates the success of this approach. It was trained entirely on Huawei Ascend chips (a domestic, less powerful alternative to Nvidia) using the MindSpore framework, with zero dependency on Nvidia hardware-1.

Despite the hardware constraint, GLM-5 outperforms GPT-5.2 on complex coding benchmarks (SWE-bench) and is priced 5-8x cheaper than comparable US models-1.

In summary, the US focuses on brute-force single-chip power, while China excels at system-level efficiency and clever software. For many real-world applications, "good enough" hardware with world-class software is beating "best-in-class" hardware with standard software.

Are you interested in a deeper look at how these specific strategies, like knowledge distillation or GPU pooling, could impact your own projects?

DeepSeek
Source: China vs US AI Market Share Comparison

Comments

No comments yet. Readers can leave comments directly from the expanded post on the board page.