+May 27, 20260How are chinese labs training their models?Chinese AI labs are training their models through a strategic pivot that compensates for restricted access to advanced NVIDIA chips by maximizing system-scale efficiency, algorithmic innovation, and domestic hardware utilization-5-10. The unified goal is to achieve competitive AI capabilities using what is available—primarily Huawei Ascend chips and a "more chips, more cleverness" philosophy.Here is a summary of the three core strategies Chinese labs are employing:StrategyCore ApproachReal-World Examples1. Massive Cluster DeploymentLink thousands of less-powerful domestic chips (e.g., Huawei Ascend) to function as a single, powerful system, compensating for lower per-chip performance with sheer scale.Huawei's 10,000-card Ascend 910C cluster; Alibaba's 10,000-chip "Zhenwu" cluster-5.2. Algorithmic & Hardware OptimizationRedesign model architecture (e.g., linear complexity, 1.58-bit quantization) and system software (e.g.,北大's parallel framework) to drastically cut computation and memory needs, boosting efficiency on domestic hardware.中科院 "瞬悉1.0" (linear-time), Tsinghua & 鹏城实验室 "开元-2B" (Sandwich Norm), OpenBMB BitCPM (ternary weights)-1-2-4.3. Complete Domestic Stack IntegrationUse a full ecosystem of Chinese technology for entire model lifecycle—from training to deployment—ensuring autonomy and bypassing Western hardware dependencies.Z.AI's GLM-Image (Ascend + MindSpore); China Unicom's enhanced DeepSeek model[CITATION:3]-10.🏗️ Strategy 1: Building "Mega-Clusters" with Domestic ChipsSince importing NVIDIA's most advanced chips is restricted, Chinese firms have shifted their focus from chasing the fastest single chip to building the largest possible clusters of domestic chips-5-10.The 10,000-Card Barrier: Both Huawei and Alibaba have recently activated major clusters built around their own Ascend and "Zhenwu" AI chips, respectively-5. A 10,000-card cluster of Huawei Ascend 910C chips is designed to train models with hundreds of billions of parameters by operating as a single, massive system-5.System Over Silicon: This approach directly acknowledges the U.S. export controls. The strategy is to compensate for lower per-chip performance through scale, advanced networking, and software innovation-10.🧠 Strategy 2: Redesigning Models for EfficiencyTo get the most out of domestic hardware, Chinese labs are also rethinking the AI models themselves to be far more efficient.Pursuing Linear Complexity: Transformer models have a "quadratic" cost, meaning they slow down significantly as sequences get longer. The Chinese Academy of Sciences (CAS) tackled this by developing "瞬悉1.0" (SpikingBrain-1.0), a "linear-time" model architecture. This allows it to handle sequences up to 4 million tokens long with over 100x faster initial response times on domestic GPUs from a company called MetaX, not NVIDIA-1.Extreme Compression with 1.58-bit Models: A team called OpenBMB pioneered a method called "BitCPM-CANN," which trains models using ternary weights (-1, 0, +1) on Huawei Ascend NPUs. This reduces memory usage by roughly 6x and achieves over 95% of the performance of a full-precision model-2.Training Stability Tricks: Because older domestic chips like the Ascend 910A are less powerful and support only limited-precision calculations (FP16, similar to a V100), training can be unstable. Researchers from Tsinghua University and 鹏城实验室 (Peng Cheng Lab) introduced techniques like "Sandwich Norm" and "Soft Clipping" to stabilize training and prevent errors from derailing the process on this more challenging hardware-6.Smarter Parallel Processing: At a systems level, Peking University developed a new "fine-grained" parallel training framework for Ascend chips. This more intelligently splits a model across different chips, leading to a 15% improvement in training efficiency by reducing idle time and improving communication-4.🏛️ Strategy 3: Building the Full Domestic StackChina is pushing for complete technological independence by developing and integrating every layer of the AI stack, from the chip up to the application.End-to-End Ecosystems: The goal is to use Chinese chips (Huawei Ascend) with Chinese AI frameworks (MindSpore) to train and run models. This has been demonstrated by Z.AI, a blacklisted company, which trained its "GLM-Image" model entirely on Huawei's Ascend Atlas 800T A2 servers using the MindSpore framework-10.Government-Led Initiatives: This push is not just by private companies. National labs like Peng Cheng are releasing fully open-source models trained "only on国产算力" (domestic computing power), creating blueprints and toolkits for others to follow-6. Meanwhile, Huawei's own PanGu-Σ model, a trillion-parameter giant, was trained on a cluster of 512 Ascend chips-7.🔬 Practical Impact: Training vs. InferenceThe real-world impact of these strategies varies depending on the task:For Training Giant Models (The Hard Part): Training a massive model from scratch remains challenging. It requires huge clusters (like the 10,000-card ones from Huawei and Alibaba) and sophisticated engineering to keep them stable-5-7. However, labs are succeeding, proving domestic clusters can get the job done, albeit with more effort and hardware.For Fine-Tuning & Inference (The Practical Part): This is where Chinese labs are finding easier success. China Unicom, for example, took the open-source DeepSeek-R1 model and fine-tuned it to be more efficient and secure for its own needs, without needing to train a new model from billions of dollars-3.I hope this gives you a clear picture of how Chinese labs are innovating within their constraints. Are you more interested in the specific technical details of the "linear-time" model from CAS or the parallel training framework from Peking University?
+May 27, 20260how many gpu chips huawei made in last 3 years yoy and future projection for 2026 and how much compute capacity it will add?Despite severe U.S. export controls cutting off access to TSMC's advanced manufacturing nodes, Huawei has managed a remarkable, albeit constrained, ramp-up of its Ascend AI processors over the last three years. By partnering closely with China’s domestic foundry, SMIC, Huawei has effectively become the primary supplier of AI compute for China's tech giants like Alibaba, Tencent, Baidu, and DeepSeek. Here is the breakdown of Huawei’s AI chip production (focusing on its top-tier Ascend data center chips) over the last three years, its aggressive 2026 targets, and the compute capacity this will unleash.Production Volume: The Last 3 Years (YoY)Huawei’s transition from early-stage survival to mass production has been gated entirely by SMIC’s ability to manufacture 7nm chips using older Deep Ultraviolet (DUV) lithography machines.2023: Low volume. This was a transitional year. Huawei had depleted its pre-sanction stockpile of TSMC-made chips and began the early ramp of the Ascend 910B using SMIC’s 7nm (N+1) process. Yields were notoriously low. 2024: ~400,000 units. Production began scaling meaningfully with the Ascend 910B. However, SMIC’s yield rate hovered around a dismal 20%, heavily restricting total output and driving up costs.2025:~700,000 to 805,000 units. A massive leap. SMIC improved its 7nm (N+2) yields to an estimated 30–50%. This volume consisted of a mix of 910B chips and the introduction of the dual-die Ascend 910C. 2026 Projections: The "SuperPoD" EraAs of mid-2026, Huawei is moving aggressively to supply China's hyperscalers, who are now formally mandated by the Chinese government to pivot away from NVIDIA. Unit Volume: Huawei projects it will produce up to 1.6 million Ascend dies across its entire lineup in 2026. High-End Shipments: Within that total, Huawei is targeting 600,000 to 750,000 shipments of its top-tier chips—specifically the Ascend 910C and the newly launched Ascend 950PR (which entered mass production in March 2026). Revenue: Huawei expects its AI chip revenue to jump at least 60% year-over-year, reaching roughly $12 billion in 2026, up from $7.5 billion in 2025. How Much Compute Capacity Will This Add?To understand the impact of these ~750,000 top-tier chips hitting Chinese data centers in 2026, we have to look at their raw throughput.Huawei's hardware strategy relies on the "Tau Scaling Law" (LogicFolding)—essentially stacking older 7nm silicon and relying heavily on advanced interconnects (like its UnifiedBus) to make up for the fact that they cannot build smaller 4nm or 3nm transistors like NVIDIA. Per-Chip Performance: The Ascend 910C delivers roughly 780 to 800 TFLOPS of BF16/FP16 compute. The newer Ascend 950PR pushes this to 1 PFLOPS in FP8 and 1.56 PFLOPS in FP4 precision. Total 2026 Capacity Added: Injecting ~750,000 of these chips into the market will add approximately 500 to 600 ExaFLOPS of BF16/FP16 compute capacity to China's AI infrastructure in a single year.The NVIDIA Equivalent: While a single Ascend 910C has roughly 30% to 40% of the raw BF16 throughput of an NVIDIA B200 (or about 50% to 60% of an H100), volume makes up for the deficit. The compute capacity Huawei is adding in 2026 is roughly equivalent to a U.S. hyperscaler deploying a fleet of 250,000 to 300,000 NVIDIA H100 GPUs.While this hardware requires drastically more power and physical space to run than an equivalent NVIDIA cluster, it proves that China has established a viable, entirely domestic pipeline capable of sustaining frontier-level AI development.
+May 27, 20260chinese foundational models trained completely on chinese chipsChinese AI labs have achieved a significant milestone by successfully training foundation models entirely on domestic Huawei Ascend chips. Zhipu AI has emerged as the leader in this push for technological self-sufficiency, with several other major players building the broader ecosystem.Below is a summary of the key models and the infrastructure that supports them.🏆 Fully Domestic Models: The Leading ExamplesA few foundational models have now been publicly confirmed as trained entirely on Chinese chips, primarily using Huawei's Ascend series. Zhipu AI has been at the forefront of this movement.Model NameDeveloperType / SizeDomestic Hardware & SoftwareSignificanceGLM-Image-1-4-6Zhipu AIMultimodal (Image Generation)Hardware: Huawei Ascend Atlas 800T A2 serversFramework: Huawei MindSporeFirst major multimodal model trained entirely on a domestic stack; topped Hugging Face's trending leaderboard-1-6.GLM-5-5Zhipu AI744B parameters (MoE architecture)Huawei Ascend chipsA massive, open-source model proving domestic chips can scale to frontier-level LLMs; MIT license-5.Qianfan-VL-9Baidu Smart CloudVision-Language (3B, 8B, 70B)Baidu's in-house Kunlun P800 chipDemonstrates an alternative to Huawei, using Baidu's own chips for full-stack domestic training-9.🏗️ The Ecosystem: Building the Domestic InfrastructureSingle model releases are just one part of the story. A broader infrastructure is rapidly developing to support this transition.Huawei's Full-Stack Solution: Zhipu's success is built on using Huawei's complete ecosystem, which includes Ascend AI processors, MindSpore (the AI framework), and CANN (the computing architecture)-1-4. This integration is key to achieving competitive performance.National Support Platforms: The Chinese government is actively supporting this shift. An open innovation platform led by China Mobile, called "AI Huanxin," is designed to provide state-owned enterprises and other companies with access to domestic chips and models to accelerate development-3.Optimization from Within: Companies like SenseTime and research teams at Huawei are not just using the chips but are also developing advanced parallel training and low-precision (FP4) techniques to close the efficiency gap with Western hardware-2-10.In summary, Chinese labs have definitively proven it is possible to train state-of-the-art foundation models using only domestic chips. The focus has now shifted from a question of "if" to one of "how efficiently" they can scale this infrastructure.Are you more interested in the technical details of how they optimize training on these domestic chips, or would you like a deeper look at the performance benchmarks of these models against their NVIDIA-trained counterparts?
+May 27, 20260compare huawei chips to nvda chipsHere's a detailed comparison between Huawei's Ascend AI chips and NVIDIA's GPUs. The short answer: NVIDIA leads decisively in raw performance, ecosystem maturity, and interconnect technology, while Huawei is making credible progress as a domestic alternative for the Chinese market, compensating for single-chip weaknesses through system-level innovation .📊 Technical Specifications: Ascend 910C vs. NVIDIA H200 & B200This table compares the flagship chips currently in production or deployment as of early 2026.FeatureNVIDIA H200 SXMNVIDIA B200Huawei Ascend 910CArchitecture / ProcessHopper / 4nm (TSMC N4)Blackwell / 4nm (TSMC 4NP)DaVinci / 7nm (SMIC N+2) -2-5FP16~1,979 TFLOPS (sparse)~2,250 TFLOPS~750-800 TFLOPS-5FP8~3,958 TFLOPS~9,000 TFLOPSNot natively supported (uses INT8) -5Memory141GB HBM3e192GB HBM3e~64-96GB HBM2e/HBM3 -2-5Memory Bandwidth4.8 TB/s8 TB/s~1.2 - 1.6 TB/s -5Chip-to-Chip Link900 GB/s (NVLink)1.8 TB/s (NVLink Gen5)~392 GB/s (HCCS) -5Power (TDP)700W1000W (Up to 1200W)~450-600W -2-5Primary AdvantageMature software & memory capacityExtreme compute & memory bandwidthDomestic production & supply chain⚡ Key Performance GapsThe table above highlights a consistent pattern: NVIDIA's chips are faster, have significantly more memory bandwidth, and benefit from more advanced manufacturing processes. The gap can be broken down into three main areas:Raw Compute Power (Single Chip): Huawei's flagship Ascend 910C, a dual-die design combining two 910B chips, offers around 800 TFLOPS of FP16 performance -5. This makes it competitive with the previous-generation NVIDIA H100 (756 TFLOPS) -2 and roughly 30-40% the performance of a B200 (2,250 TFLOPS) in this precision . The gap is much wider in the crucial FP8 precision used for training large models, which the 910C does not natively support, unlike H200 and B200 . One analysis suggests the single-chip performance gap will widen to 17x by 2027 as NVIDIA continues to advance -1-4.Memory Bandwidth (The Hidden Bottleneck): For large language models, getting data to the compute engines is as important as the compute itself. Here, the gap is stark. NVIDIA's H200 and B200 use advanced HBM3e memory delivering up to 8 TB/s of bandwidth-2-5. The Ascend 910C is limited to around 1.6 TB/s, a roughly 5x disadvantage that severely impacts its ability to efficiently run the largest models .System Interconnect (Scaling Up): For massive AI clusters, chips must communicate at high speed. NVIDIA's NVLink provides up to 1.8 TB/s of bidirectional bandwidth, allowing hundreds of GPUs to function almost as a single giant chip -5. Huawei's HCCS interconnect is significantly slower, at roughly 400 GB/s. This means Huawei clusters are less efficient at scale; one analysis estimates their performance in a "supercluster" is only 20-33% as efficient as an equivalent NVIDIA cluster for training trillion-parameter models .🏗️ A Tale of Two StrategiesThe performance gap leads to two fundamentally different market strategies:NVIDIA's Approach: "Maximum Performance" : NVIDIA produces the world's most powerful single chip. This is the default choice for global tech giants (OpenAI, Google, Meta, ByteDance) training the largest, most advanced frontier models . Their dominance is locked in by the CUDA ecosystem, a software platform with millions of developers that has become the industry standard. Switching away from it requires rewriting massive amounts of code .Huawei's Approach: "System-Level Innovation" : Aware of its single-chip limitations, Huawei is innovating at the system level. Its flagship CloudMatrix 384 architecture links 384 Ascend 910C chips into a single, massive "supernode." By pooling memory and compute, Huawei claims this system can match or even exceed the performance of NVIDIA's GB200 NVL72 system on certain large-scale inference tasks, such as running the DeepSeek-R1 model . This clever workaround is a direct response to its inability to produce a single ultra-powerful chip.🔭 The Future RoadmapThe race is not static.Huawei has a clear, state-backed roadmap. Next up is the Ascend 910D (expected late 2026), which will introduce native FP8 support to close a key gap . The Ascend 950 (targeting 1 PFLOPS FP8) and Ascend 960/970 are planned for 2027-2028 -5. In the longer term, Huawei has announced an ambitious new "Tau Scaling Law" design philosophy and an alternative chip architecture aiming for the equivalent of a 1.4nm process by 2031, a direct challenge to conventional chip design -3. Success is not guaranteed, but it shows a long-term vision to circumvent U.S. sanctions.NVIDIA continues to execute on its breakneck product cycle. Following the B200, the B300 (Blackwell Ultra) is already in testing, offering a modest 1.16x performance increase in training workloads due to architectural and memory optimizations . The Rubin architecture is already on the horizon, promising to extend its lead further in 2027. Each new generation widens the absolute performance gap -1-4.🔬 Practical Impact: Training vs. InferenceThe real-world meaning of these numbers is best understood by looking at the specific task:For Training Frontier Models (e.g., GPT-5 level): This requires massive, highly efficient clusters. The combination of slower single-chip performance, lower memory bandwidth, and less efficient interconnects means a Huawei cluster would need to be 3 to 5 times larger to match the training speed of an equivalent NVIDIA cluster. The cost, complexity, and power consumption make this a significant disadvantage .For Inference (Running the Model): This is where Huawei is more competitive. For tasks like running a chatbot, the memory capacity and bandwidth per chip are critical. With clever system design like the CloudMatrix 384, Huawei can pool resources to deliver competitive token-generation speeds. This makes the Ascend series a viable, if less efficient, option for deploying models once they are trained .💎 Summary: Who Wins Where?NVIDIA wins hands-down for any organization that can buy its chips. For global AI labs, financial institutions, and tech giants outside of China, NVIDIA's B200 and H200 remain the only viable choice for building state-of-the-art AI systems. The performance and software moat is simply too wide .Huawei's "win" is a more localized one. In the Chinese market, where U.S. sanctions prevent the sale of NVIDIA's most advanced chips (H100, B200, H200 restricted), the Ascend series is the best and most scalable domestic alternative -6. With strong government backing and a long-term roadmap, it will continue to capture market share in China. For Chinese companies, the choice isn't "NVIDIA or Huawei," it's "Huawei or nothing at all for top-tier performance" -6.I hope this detailed comparison is helpful. Would you like to dive deeper into the software ecosystems (CUDA vs. CANN) or the specific system-level architectures like Huawei's CloudMatrix 384?
+May 26, 20260USA vs China GPU data centers in gwThe comparison between U.S. and Chinese GPU data center capacity reveals a fascinating strategic divergence: the U.S. currently leads in total computing power and infrastructure scale, but China holds a decisive structural advantage in energy capacity, which is increasingly the bottleneck for future AI growth -1-8.Here is the data on their current standing and future potential:Metric🇺🇸 United States🇨🇳 ChinaCurrent Data Center Capacity29.2 GW (43% of global total) -28.5 GW (13% of global total) -2Future Potential (Pipeline)~102 GW if all planned projects are built -6Massive headroom of 58.9 GW without needing new power plants -2Power Grid HeadroomLimited (facing 49 GW shortfall by 2028) -4-8Vast surplus (400 GW effective reserve by 2030) -1-8Chip/Compute Power~75% of global AI compute; 2025 adds >25 ZFLOPs -4-9~1-4% of U.S. compute; 2025 adds ~1 ZFLOP -4-9🏗️ Divergent Paths to AI DominanceThe data shows that the two nations are not playing the same game. The U.S. strategy focuses on maximizing chip performance, while China's strategy uses energy scale to compensate for less advanced chips.The U.S. "Super Factory" Model: American tech giants like Microsoft, Meta, and OpenAI are building "AI Super Factories" -5. These are single campuses demanding massive power (often 1-2 GW each, with some targeting up to 12 GW) -10. This model depends on cutting-edge chips like Nvidia's Blackwell, which are largely restricted from export to China -9. The main hurdle is that the U.S. electrical grid has been stagnant for 20 years, leading to permitting delays of up to five years -7-10.China's "National Grid" Model: China leverages its state-directed "Eastern Data, Western Computing" initiative, building renewable energy hubs in the west to serve the east's AI needs -10. With a national grid capable of adding over 430 GW of wind and solar in a single year (8x the U.S. rate), China is solving the AI problem with abundant, cheap energy -1-10. This allows them to use domestically produced chips (like Huawei's Ascend), which are less powerful than Nvidia's best but are deployed in vast numbers -4-10.💡 What This Means for the FutureThe core competition may be shifting from chip design to energy infrastructure. Here is what the analysts are saying:The Bottleneck is Shifting: The bottleneck for AI is rapidly moving from "do we have enough chips?" to "do we have enough power to run them?" -8-10. While the U.S. leads in GPUs, China’s massive power surplus gives it a unique strategic lever -1-7.The Timeline: China’s chip industry is rapidly advancing toward self-sufficiency and could potentially meet 76% of domestic AI chip demand by 2030 -4. If that happens, and they combine it with their energy advantage, the current global balance could shift significantly.In summary, the United States is winning the battle of maximum compute power today, but China is positioning itself to win the war of scale and energy sustainability for the next decade.I hope this breakdown is helpful. Is there a specific aspect of this competition, such as the technology behind the chips or the investment figures, that you would like to explore further?
+May 26, 20260How does Chinese ai labs see monitizing ai?Chinese AI labs have shifted from an initial focus on technical benchmarks ("SOTA") to building sustainable, revenue-generating businesses. Their strategy centers on three interconnected priorities:1. API Platform Monetization (MaaS)The core strategy is Monetization-as-a-Service (MaaS), transitioning from project-based delivery to standardized, usage-based API models -2-7. This approach has shown strong traction:Zhipu AI reached ~$250M USD in API ARR (Annual Recurring Revenue) by March 2026, up 60x year-over-year -7.MiniMax doubled its ARR from 100Mtoover100M to over 100Mtoover150M in just two months in early 2026 -3-8.Moonshot AI (Kimi) saw API calls surge, driven by developer tools like OpenClaw, generating more revenue in 20 days than all of 2025 -4.The value of this model is pricing power. Unlike a broad price war, top labs are raising prices. For example, Zhipu increased token prices by 83% in early 2026, indicating strong demand in high-value areas like coding and agents -7.2. Embracing an "AI Platform Company" RoleLabs want to be platform companies that define new AI paradigms. This depends on two factors -8:Intelligence Density: The raw capability of the model (e.g., MiniMax's upcoming M3, Zhipu's GLM-5).Token Throughput: The scale at which tokens are processed.Platform value emerges when a new intelligence breakthrough creates new use cases (coding, video, office automation), bringing in developers and users to form an ecosystem. This is a "model evolution speed" race -3-8.3. Shifting from Open Source to Proprietary ModelsTo monetize directly, companies are pivoting from open-source to closed, proprietary models -1. Cloud APIs are now the primary access method, as customers move away from self-hosting. As a Zhipu executive stated, "There is no longer a need to rely on open source to raise overseas recognition" -1. Kuaishou's Kling AI is a profitable example of a product built on a non-public model -1.4. Dual Focus: Enterprise APIs and Consumer ProductsMost successful labs operate a "dual-engine" model -8:Enterprise (API): Provides scalable, predictable recurring revenue.Consumer (C-end): Consumer apps provide cash flow and serve as a training ground for models via real user feedback.However, monetization in China is not easy. Tencent notes that AI delivery has high variable costs, unlike the internet era. In China, paid subscriptions are in the "single-digit percentages," so companies are cautious about over-investing without a clear return -6.5. Realistic Product RoadmapsLabs are targeting specific areas with clear ROI:Coding (L4-L5 Agents): Moving from tools to "colleague-level" collaboration -3-8.Office Productivity (Next Frontier): Poised to replicate coding's rapid growth -8.Multimodal Creation: Moving towards real-time, publishable long-form content -8.In summary, Chinese AI labs are treating AI as a utility metered by API usage, building platform ecosystems, and prioritizing profitability over open-source openness, despite a challenging domestic market for direct consumer subscriptions.
+May 26, 20260Then what are us models applying efficieny for what are they achieving but applying those efficienciesUS companies are applying the same efficiency techniques (like Mixture of Experts and model distillation) to achieve a fundamentally different goal. Because they have access to cutting-edge chips (Nvidia B200/GB200) and abundant capital, they optimize for maximum capability and solving the world's hardest problems, rather than for survival or cost reduction -3.Here is what US models are achieving by applying those efficiencies:🚀 Pushing the Frontier of PerformanceUS firms use efficiency gains not to reduce costs, but to reinvest all savings into building even larger, more capable models. This is about reaching human-expert or superhuman levels on the hardest benchmarks -9.DomainAchievementEvidenceSoftware EngineeringNear-human performance in one yearSWE-bench Verified went from ~60% to ~100% of human baseline in just 12 months -4-6-9.Postgraduate ScienceSurpassed expert baselineGPQA Diamond (PhD-level science) hit 93%, exceeding the 81% expert benchmark -9.ReasoningConverged at the topTop 4 US models are now within 25 Elo points of each other (Anthropic 1503, xAI 1495, Google 1494) -6-7.🧪 Accelerating National Scientific DiscoveryThe Genesis Mission, led by the Department of Energy (DOE), is the most significant example of what US efficiency is unlocking. It’s a $320 million national initiative to build the "American Science and Security Platform," aiming to double the productivity of American science within a decade-1-5.Strategic GoalDescriptionFoundation Models for ScienceTraining specialized AI on unique DOE data (e.g., particle physics) to create models that don't exist in the commercial sector -1-5.Robotic LabsDeveloping autonomous systems to run experiments 24/7, moving from human-led to AI-led lab work -1-5.Solving National ChallengesTackling 26 "Grand Challenges" in energy, materials, semiconductors, and national security .🛡️ Pursuing World-Leading AI AgentsEfficiency also enables the creation of agents that complete complex, multi-step tasks reliably. This pushes AI from "chat" to "action" -2.Computer Use: Success rate on the OSWorld benchmark (real-world computer tasks) jumped from 12% to 66% (human baseline is ~72%) -2-4-6.Cybersecurity: Problem-solving success rate skyrocketed from 15% to 93% in a single year -2.💡 Key Takeaway: Different PhilosophiesSimply put: US models use efficiency to explore the frontier of intelligence; Chinese models use efficiency to win on price and scale -3.The US Strategy: Apply efficiency to reinvest in scale → This pushes world-leading benchmarks in science, coding, and complex reasoning.The Chinese Strategy: Apply efficiency to cut costs → This wins global market share on price and dominates industrial robotics -2-3-9.Are you interested in a deeper dive into the specific "Grand Challenges" of the Genesis Mission, or how to benchmark these agent capabilities for a specific use case?
+May 26, 20260When user spends token on workbuddy how does dollar change hands from user to further players in chainWhen a user spends tokens on Tencent’s WorkBuddy, the actual money flows through a clear, multi-step chain. The user pays the Distributor (carrier or Tencent), who then settles costs with the Model Provider (Tencent Hunyuan or a third party) and the Infrastructure Providers (data centers and chip suppliers).Here is the step-by-step path of a dollar:💸 The Flow of Money: A 5-Step JourneyUser Pays the Distributor (The Retailer)You start by buying "credits" or "tokens" from a retail front. In WorkBuddy’s case, you pay Tencent directly (e.g., $9.95/month for a subscription) or pay a carrier like China Mobile via your phone bill for a token package -1-3-8.Distributor Aggregates DemandThe platform (WorkBuddy or the carrier’s hub) collects millions of these small payments. It acts like a wholesaler, pooling your $9.95 with everyone else’s. The core business logic here is that WorkBuddy is designed to be a "Token粉碎机" (Token Crusher), converting simple interactions into complex, multi-step tasks that burn through tokens much faster than standard chat .Distributor Pays the Model ProviderWorkBuddy processes your task. If you use the default setting, Tencent’s own Hunyuan model handles the work. If you use a premium external model (like GLM-4.7), Tencent uses your money to pay Zhipu AI or another third-party API provider -3. The settlement system (like China Mobile’s MoMA platform) handles "统一结算" (unified settlement) to split the revenue -1.Model Provider Pays the InfrastructureThe entity running the AI model (Tencent or Zhipu) must cover its costs:Compute (GPU Cloud): They pay Alibaba Cloud, Tencent Cloud, or a specialized AI data center (AIDC) to rent the servers that actually run the calculations -2-7.Hardware & Power: The cloud providers use their revenue to buy chips from Nvidia, Huawei (Ascend), or海光 (Hygon) , and pay massive electricity bills to state grid operators to keep the machines running -2-7.Full Loop CompletionThe infrastructure providers pay their suppliers (chip designers pay fabs, data centers pay construction firms), completing the cycle. The user gets the AI task done, and every player in the chain gets their slice.🤝 Key Players in the ChainLayerRoleExample EntitiesApplication LayerThe "Retailer" selling access to youTencent (WorkBuddy), ByteDance, Carriers -3Aggregator/Distribution LayerThe "Wholesaler" managing billing/routingChina Mobile (MoMA), China Telecom (TokenHub) -1-5Model LayerThe "Engine" generating the intelligenceTencent Hunyuan, Zhipu AI (GLM), DeepSeek -2Infrastructure LayerThe "Factory" providing power & chipsAlibaba Cloud,浪潮 (Inspur), 华为, 英伟达 -2-7💡 WorkBuddy’s Strategic MechanismIt is important to note that WorkBuddy is not just a tool; it is a strategic mechanism designed to multiply consumption. By enabling complex agentic workflows rather than simple chat, it inherently consumes orders of magnitude more tokens, which directly drives higher revenue for the ecosystem .Would you like to see how this financial chain differs when using a WorkBuddy-like service through a telecom provider's phone bill instead of a direct app store subscription?
+May 26, 20260What are Chinese model companies saying about revenue generationBased on official financial reports and executive commentary from Q1-Q2 2026, Chinese AI model companies are shifting from an investment-heavy phase to a monetization-driven phase. While strategies vary by company, a clear consensus is emerging: revenue growth is now the primary metric of success.Here is what the major players are saying about their revenue generation strategies and results.🏢 Alibaba: The Scale & ARR PlayAlibaba is betting on full-stack integration to drive high-margin recurring revenue. CEO Wu Yongming has declared that Alibaba AI has formally entered its "commercialization return cycle"-2-7.The Strategy: Drive revenue through Model-as-a-Service (MaaS) via the "Bailian" platform and token sales via the new ATH (Alibaba Token Hub) business group -1-2.Key Metrics: AI-related revenue now accounts for 30% of cloud external revenue (~¥8.97 billion in Q4) and is expected to exceed 50% next year -2-6-7. The ARR for its AI models and MaaS services has already exceeded ¥8 billion ($1.1B) and is on track to hit ¥30 billion by the end of 2026 -2-7.Efficiency: The strategy is to use its in-house "Pingtouge" GPUs (47k chips delivered) to improve margins and lower costs, creating a flywheel where more apps drive more token consumption -2-6.🐧 Tencent: The Ecosystem IntegratorTencent is focusing on embedding AI into its vast ecosystem (WeChat, Ads, Games) and creating Agent products before aggressively monetizing them.The Strategy: Prioritize internal R&D and user growth over immediate cloud revenue. Tencent has deliberately delayed renting out its GPUs via Tencent Cloud to prioritize internal products like WorkBuddy and CodeBuddy -1-8.Key Metrics: Q1 revenue hit ¥196.5B (up 9%). Cloud & Enterprise revenue grew 20%, driven by AI -3-6. Marketing services (ads) grew 20% thanks to AI-driven recommendation upgrades -3.Cautious Monetization: President Martin Lau noted that unlike the internet, AI has high variable costs per query, so they are focusing on "high-value scenarios" rather than just acquiring DAUs -8.🐻❄️ Baidu: The Tipping PointBaidu reports that its AI business has officially passed the "tipping point," now contributing the majority of core revenue.The Strategy: Monetize via a full-stack approach: AI Cloud (selling infrastructure), AI Applications (subscriptions), and AI Native Marketing.Key Metrics: Core AI revenue (Cloud + Apps + Marketing) reached ¥13.6 billion, accounting for 52% of core revenue for the first time -4-9.Profitability Challenge: While AI revenue is growing fast (Cloud up 79%), overall group net profit fell 55% as high-margin traditional search ads declined -9. They are pushing overseas Robotaxi expansion for future revenue -4.🧠 Zhipu (智谱): The Pricing Power PlayZhipu is aggressively proving that "pricing power" (raising prices without losing customers) is the true measure of value.The Strategy: Move from the 2024 "price war" to a value-based pricing model. CEO Zhang Peng emphasizes monetizing the MaaS platform -10.Key Metrics: Full year 2025 revenue was ¥724M (up 132%) -10. MaaS API ARR reached ¥1.7 billion (up 60x) -10.The "GLM-5" Test: They raised API prices by 83% in Q1 2026. Management reports that demand (calls) still grew 400%, proving customers pay for results, not the lowest price -1-5-10.🔬 MiniMax & DeepSeek: The Different PathsMiniMax is focusing on "Three-Legged Stool": maintaining its C-end user base (Talkie), expanding B-end enterprise sales (hiring ex-Huawei cloud execs), and competing in AI coding -1.DeepSeek is playing the "Long Game." Instead of pushing for immediate token revenue, they are reportedly prioritizing domestic chip (Ascend) adaptation. They are sacrificing short-term growth to maintain long-term strategic independence from Nvidia's CUDA ecosystem -1.📝 Summary Table: Revenue Generation FocusCompanyPrimary StrategyKey Metric TrackedCustomer FocusAlibabaFull-stack Synergy (Cloud + Chip + Model)MaaS ARR (Targeting ¥30B)Enterprise & Developers -2TencentEcosystem Integration (WeChat/Ads)Agent Adoption & Cloud RevenueInternal Ecosystem/Enterprise -3-8BaiduHybrid (Cloud Infra + Apps)AI Revenue % (Surpassed 50%)Hybrid (Gov/Enterprise/Consumer) -4ZhipuValue-based MaaS (API)Pricing Power (83% price increase)Developers & Enterprises -10MiniMaxC-end Monetization (Talkie) + B-endUser-base ConversionConsumer (Global) -1In summary, the Chinese AI industry has moved past the phase of pure technological competition into a "revenue and efficiency" validation phase. Success is no longer measured by a model's benchmark score, but by its Annual Recurring Revenue (ARR) and ability to turn tokens into profit.Would you like a deeper look into the specific financial performance of Alibaba Cloud or Baidu's AI segment?
+May 26, 20260Then how do they spend those tokens?Once Chinese consumers purchase a token package from a telecom carrier, "spending" tokens is an automatic process. You don't actively do anything—the AI application deducts tokens from your balance every time it performs a task for you -1-2.The best way to understand this is to think of tokens exactly like mobile data. You buy a 10GB data plan, then your phone automatically uses up that data when you watch videos or browse the web. The only difference is that tokens are the "fuel" for AI tasks, not internet browsing.🔄 The "Set It and Forget It" WorkflowHere is the exact process of how tokens are spent once a consumer buys a package:Account Binding (The Link): After purchasing the token package (e.g., via the China Mobile or China Telecom app), the user links their phone number to an AI application. This is done either by entering the phone number directly in the app or by configuring an "API Key" (a unique code provided by the carrier) in advanced software -2-6.Task Execution (The Burn): When you ask an AI to do something complex—such as "generate a PowerPoint presentation" or "summarize this 100-page document"—the app automatically deducts tokens from your balance. This happens invisibly in the background, just like data being used when you load a webpage -1-7.📊 What Does "Spending" Look Like? (Usage Table)To visualize how fast these tokens go, here is the breakdown of token costs for different tasks using the carrier packages -2-5-7:TaskApproximate Token CostCost with Pay-As-You-GoCost with Token PackageGenerate a 1-minute AI Video~1 Million Tokens~7.00−7.00 - 7.00−8.00 USD (¥50-60)~$0.14 USD (¥1)-2-5Professional Coding or Data AnalysisTens of thousands per hourExpensive per API callVery cheap (bulk rate) -2-6Simple Chat (e.g., "What's the weather?")~50 - 200 TokensFree (usually)Free (wrapped into the bundle) -6🛒 Where Do They Spend Them? (The "App Store" for AI)You don't just spend tokens in one place. The carrier packages function like a "universal adapter" that works across hundreds of different applications -3-6.The "Super App" Ecosystem: Once you buy a China Mobile package, your tokens are valid in apps like Tencent WorkBuddy (work collaboration), OpenClaw (AI agents), and TeleClaw (Chinese telecom's agent system). You use the same "balance" across all of them -2.The "Token Supermarket" Model: The carrier acts as a middleman. You pay the carrier, and the carrier pays the specific AI model provider (like DeepSeek or Zhipu) behind the scenes every time you use their service -3-6.📈 The "Heavy User" ScenarioWhile casual users chat for free, "spending" tokens becomes essential for "Heavy Users" (Prosumers/Businesses).The Automation Leap: Users are now deploying "AI Agents" (like OpenClaw) that can work autonomously for hours to complete tasks like cleaning data or booking travel. These agents "burn" millions of tokens automatically in a single night -7.The Cost Math: Without a package, running a complex AI agent could cost a fortune. With a package, a task that cost ~7.00(¥50)nowcostsroughly 7.00 (¥50) now costs roughly ~7.00(¥50)nowcostsroughly0.14 (¥1) -2.In short, once you buy the package, spending tokens is automatic: The AI works, the math happens in the cloud, and your token balance drops silently in the background—just like your phone data.
+May 26, 20260Telecom ai capex and what are they doing with this capex?Based on the latest financial reports and infrastructure plans, China's three major telecom carriers are making a strategic pivot away from traditional network spending and towards AI infrastructure. While their total capital expenditure (capex) is declining, the proportion dedicated to AI and computing power is rising sharply.Here is the breakdown of their 2026 AI capex and what they are building with it:💰 2026 Telecom AI Capex & Strategic FocusCarrierTotal 2026 CapexAI/Computing InvestmentKey Projects & StrategyChina MobileRMB 136.6 bn (-9.5% YoY)Computing +37% of total62.4% increase for computing networks; 19.8% for AI networks -1. Building AIDC centers and the "Jiutian 3.0" foundation model -6.China TelecomRMB 73.0 bn (-9.2% YoY)Computing 35% of total26% increase for computing infrastructure -7. Focusing on "Token-based business" and managing 1.2 GW of AIDC projects -2-7.China Unicom~RMB 50.0 bn (-7.7% YoY)Computing >35% of totalAI revenue grew 140%-3-5. Planning a RMB 51 billion AI hub in Suzhou (30,000 racks) and 7 AIDC parks -10.🏗️ What They Are Building with This CapexThe money is being deployed into physical infrastructure to create a "computing network" that rivals the scale of their traditional phone networks. Here is what that looks like on the ground:1. AI Data Centers (AIDCs)Instead of standard data centers, carriers are building "AIDC" facilities designed specifically for the high-density power and cooling needs of GPUs.China Telecom is leading this charge. Management stated that AIDCs will be the "major source of future revenue," with over 20 high-power projects totaling 1.2 gigawatts and contracts exceeding RMB 100 million already in place -7.Specific Example: The China Telecom (Jinan) Smart Computing Center involves a total investment of RMB 3.36 billion for its first phase alone, set to deliver 500 PFLOPS of computing power specifically for AI model training and industrial simulation -4-9.2. Mega-Scale AI Hubs (The "GPU Farms")The carriers are building massive, centralized hubs to host tens of thousands of servers.China Unicom's Suzhou Project: The most eye-catching single project is China Unicom’s planned RMB 51.27 billion AI hub in Suzhou. This single facility will house approximately 30,000 standard server racks and cover 93.3 acres -10. For context, this represents a massive bet on future AI demand in the Yangtze River Delta.3. Vertical Integration (From Chip to Token)Unlike US hyperscalers who strictly sell cloud access, Chinese carriers are building a vertically integrated stack.China Mobile is developing its own "Jiutian 3.0" foundation model and 29 industry-specific intelligent agents -1.China Unicom is integrating its "Lingxi" platform and accumulating massive datasets (over 400TB) to make the infrastructure useful out of the box -5.Shift to "Tokenization": Executives explicitly stated they are moving from a "traffic-based" economy (selling bandwidth) to a "Token-based" economy (selling AI processing power) -2-6.💡 The Strategic LogicThis shift explains a lot about the earlier questions regarding how Chinese consumers spend on AI. The carriers are not just passive "pipes" for internet traffic; they are building themselves into the AI utility layer-6.By controlling the hardware (AIDCs), the chips (domestic GPUs), and the billing (phone bill tokens), they are positioning themselves to capture the revenue from every AI query. The 35-37% allocation of capex to AI is the fuel for the "Token as Currency" economy you asked about earlier.Would you like to dig deeper into the specific types of AIDC projects or the domestic chips powering them?
+May 26, 20260How much percent of their ocf are they putting in capex ?Based on the most recent complete financial data, the percentage of operating cash flow (OCF) being allocated to capital expenditures (capex) varies significantly among the major Chinese hyperscalers:腾讯控股 (Tencent): 26.1% (as of FY2025)阿里巴巴 (Alibaba): 45.7% (as of HY2025, annualized estimate) -4-7百度 (Baidu): Calculating (First half of FY2025 reported OCF was negative) -3-5Here is the detailed breakdown of the numbers:1. 腾讯控股 (Tencent): High Profitability, Low Capital IntensityTencent shows a relatively low capex-to-OCF ratio, characteristic of a mature, high-margin business.Data: For the full fiscal year 2025, Tencent reported an Operating Cash Flow of RMB 303.05 billion and Capex of RMB 79.20 billion .Result: This means capex consumed about 26.1% of their operational cash generation .Context: Despite being low relative to its cash flow, Tencent's capex is still massive in absolute terms and the company has stated it plans to roughly double its AI investment in 2026, signaling a strategic shift toward higher capital intensity .2. 阿里巴巴 (Alibaba): High Strategic InvestmentAlibaba is in a higher spending phase, pouring a larger chunk of its cash flow back into infrastructure.Data: For the first half of fiscal year 2025 (ended Sep 30, 2025), Alibaba generated RMB 30.77 billion in OCF and spent RMB 51.32 billion on capex -4.Result: This puts their capex at approximately 166.7% of their half-year OCF. On an annualized basis (using full-year FY2025 estimates), the figure is roughly 45.7%-4-7.Context: This massive spending reflects their strategic focus on scaling cloud and AI capabilities, as cloud revenue growth of 17.7% YoY in early 2025 was directly attributed to rising enterprise AI demand -6.3. 百度 (Baidu): A Different Financial ProfileBaidu's situation is distinct due to cash flow volatility. For the first half of 2025, their OCF was negative RMB 3.0 billion-3-5. A company cannot meaningfully calculate a "percentage of OCF" for capex when OCF is negative.Context: However, this is a recent challenge as OCF turned positive in Q3 and Q4 2025, generating RMB 3.9 billion in the second half of the year -3. This turnaround suggests their financial ability to invest is improving.💡 Context vs. Global HyperscalersFor perspective on these numbers, it is helpful to compare them to global peers:Chinese Hyperscalers: Historically spent ~10-15% of revenue on capex .US Hyperscalers (The "Big 5"): More aggressive, spending roughly 25% of their revenue on cash capex .The data shows Tencent operating a highly efficient model, Alibaba aggressively reinvesting cash flow into growth, and Baidu emerging from a period of operational cash flow recovery.Would you like to see how these spending levels compare to their actual AI-driven revenue growth?
+May 26, 20260Explain me token as currency, how Chinese consumers are paying for ai?The "Token as Currency" model is fundamentally changing how Chinese consumers pay for AI. Instead of paying a fixed monthly fee for a specific software (SaaS), they buy pre-paid packs of "Tokens"—the basic unit of AI computation-1-4.Treating AI like a utility (electricity or mobile data) has led to three major ways Chinese consumers are now paying for AI.💳 How Chinese Consumers Pay for AIThe most significant shift is that AI tokens are now sold like standard top-up plans by the country's three major telecom carriers-1-7.ProviderPlan TypePrice (RMB/USD)Token AllowanceChina TelecomConsumer (Monthly)9.9 RMB (~$1.40)10 million tokensConsumer (Monthly)49.9 RMB (~$7.35)80 million tokensDeveloper (Monthly)39.9 – 299.9 RMB15M – 150M tokensChina MobileEntry (One-time)1 RMB (~$0.14)400,000 tokensPremium (Monthly)40 RMB (~$5.90)20 million tokensChina UnicomEntry (Monthly)15 RMB (~$2.20)6 million tokensStrategic Importance: This is a pivotal shift. By allowing AI tokens to be paid for via monthly phone bills, the carriers are expanding the market beyond tech-savvy developers to potentially hundreds of millions of ordinary mobile users-1-4.💻 Platform-Specific Token Plans: Subscription & ConsumptionBeyond the carriers, major tech platforms directly sell token plans. These often involve a hybrid model: a base monthly subscription plus a "Credit" system for consumption.Alibaba Cloud (Bailian Token Plan): This plan uses a unified "Credit" system. For ¥198/month, users get 25,000 Credits; for ¥1,398/month, they get 250,000 Credits-3-6. A typical request might cost ~3.18 Credits-3-6.Xiaomi (MiMo API): They adopt a pure pay-as-you-go consumption model. Their API is priced at just 0.7 RMB per million tokens, which is about 2.5% the cost of leading global models-8.Leading AI Models (Zhipu & MiniMax): These companies have adopted a regional pricing strategy. Their subscription plans are significantly cheaper for domestic Chinese users (e.g., ¥49 RMB) than for international users ($18 USD), creating massive price advantages at home-2-5.🤖 The Future: Agent-to-Agent (AI-to-AI) PaymentsThis is the most advanced layer, where tokens enable machines to transact directly on behalf of users.Alipay's AI Payment: Alipay has launched a specific product for AI-to-AI payments. An AI agent acting on your behalf can negotiate with another service's AI to buy a skill (like a premium web search) and complete the transaction automatically using pre-set permissions-10.Autonomous Agents (e.g., OpenClaw): This technology powers the "OpenClaw" agent ecosystem. An agent can autonomously book a trip, manage a server, or buy computing power, with the "token" acting as the fuel for the transaction logic-10.Ultimately, what's happening in China is the shift from "paying for software" to "paying for computational value." Whether it's a student topping up via their phone bill or an AI agent negotiating a data purchase, the token has become the universal currency of this new economy-1-4-10.Would you like to dive deeper into how the math works on these token plans (e.g., what 10 million tokens can actually do)?
+May 26, 20260Chinese vs us ai models market shareBased on the latest data (as of May 2026), the competitive landscape between Chinese and US AI models is rapidly shifting. While the US maintains a dominant lead in investment and certain enterprise metrics, China is demonstrating significant strength in global usage, model performance, and key industrial applications-1-2-7.Here is the detailed breakdown of their market share across different metrics:📊 1. Global Model Usage (Token Call Volume)Winner: ChinaAccording to OpenRouter (a major global API aggregation platform), China overtook the US in February 2026.Weekly Call Volume (Late Feb 2026): China reached 5.16 trillion tokens, compared to the US's 2.94 trillion (Chinese volume nearly double the US).Global Top 5 Models: Chinese models occupy 4 out of 5 spots (by actual developer usage, not just capability), including MiniMax, DeepSeek, and Zhipu-2-10.Key Driver: Chinese models offer drastically lower costs (approx. 0.3/milliontokens∗∗vs.US∗∗0.3/million tokens** vs. US **0.3/milliontokens∗∗vs.US∗∗5/million tokens) and efficient "Mixture of Experts" architecture-2.⚡ 2. Model Performance & CapabilityWinner: Tie (Race effectively closed)Stanford University’s 2026 AI Index report confirms the performance gap has nearly vanished.Performance Score: The top US model (Anthropic) leads by only 2.7% (39 points). Back in May 2023, the US lead was over 300 points.Trading Places: Chinese models (like DeepSeek-R1) have traded the #1 rank with US models multiple times since early 2025-1-4-7.🇺🇸 3. US Consumer Market (Direct User Traffic)Winner: USDespite China's global rise, the US consumer market remains dominated by local giants.Top Players (March 2026):ChatGPT (OpenAI): 33.86 million users (Dominant leader).Google Gemini: 10.66 million users.Microsoft Copilot: 5.02 million users.Chinese Presence: DeepSeek ranks 6th with 410,000 users, showing a strong growth rate (46.2%) but still a niche player in the US consumer space-3-6.🏭 4. Research & Industrial MetricsDifferentiated StrengthsThe US and China lead in different areas:Total AI Patents & Citations: China leads (20.6% of global citations vs. US 12.6%)-1.Top-Tier "Notable" Models (2025): US leads (50 models produced vs. China's 30)-1-9.Industrial Robots: China dominates (295k installations vs. US 34k, nearly 9x more)-1.Private Investment: US leads overwhelmingly (285.9billionvs.China′s285.9 billion vs. China's 285.9billionvs.China′s12.4 billion). The US funds more new AI companies than the next 10 countries combined-1-9.💡 5. Developer Mindshare (Enterprise Preference)Trending Toward China (Specifically Alibaba)Global Developers: Andreesen Horowitz notes that 80% of AI startups in Silicon Valley now rely on Chinese open-source models for their core tech-2.Chinese Market: Among Chinese CIOs surveyed by Morgan Stanley, Alibaba (Qwen) is the top pick to win the AI spending race (41%), surpassing DeepSeek (18%)-8.📝 SummaryIf "Market Share" means Global Usage (APIs): China has taken the lead due to cost and accessibility-2.If "Market Share" means US Consumer Front-end: US remains the absolute leader (ChatGPT, Gemini)-3.If "Market Share" means R&D Investment: US outspends China by a staggering 23:1, ensuring an ongoing battle for top-tier foundational breakthroughs-9.Are you more interested in a specific sector, such as which model leads in coding or mathematical reasoning? Let me know, and I can look up those performance benchmarks.
+May 26, 20260So us models are better than Chinese models?Based on the latest data from May 2026, the short answer is: No, US models are not categorically "better" anymore. It depends entirely on what you measure.For most general-purpose tasks, Stanford University's authoritative 2026 AI Index Report confirms that the performance gap between top US and Chinese models has shrunk to a razor-thin 2.7%-1-8-9.The reality is that these models are effectively competing in the same tier. Here is the detailed breakdown of where each side leads.🏆 The Verdict: A Detailed BreakdownMetricWinnerKey Data & WhyTop-Tier PerformanceTie (Statistically)The gap is just 2.7% (e.g., US ~1503 vs CN ~1464 Elo). In 2025, Chinese models (DeepSeek-R1) temporarily matched top US models, and six models from both nations are now clustered in the same elite tier -1-4-7.Coding & MathMixed (Depends on Task)US (GPT-5.5) leads in real-world bug fixing (81% vs 74%). China (DeepSeek) is competitive in math olympiads (96-97%) and PhD science (90% vs 91%) -2-6.Cost EfficiencyChinaChinese models are drastically cheaper (e.g., DeepSeek is ~75% less expensive than US equivalents for similar output) and often open-source, leading in global usage volume -2-10.Research & ScaleSplitUS leads in investment (285Bvs285B vs 285Bvs12B) and foundational models. China leads in real-world robotics (54% of global share), patents, and industrial application -1-4-8.🔬 Why the Answer is "It Depends"The "better" label has become situational because the two countries have diverging priorities:US Strength: The US focuses on brute-force capability and raw intelligence. It produces more foundational breakthroughs (50 models vs. China's 30 in 2025) and leads in private investment -1-4.China Strength: China focuses on efficiency and application. Facing hardware limits, Chinese engineers optimized software and architecture to create models that are roughly as smart but run much cheaper, allowing for massive deployment in factories and apps -8-10.💡 The Bottom LineIf you need the absolute highest raw score on a specific, niche benchmark, a US model might win by a hair. However, for 99% of real-world business applications—coding, reasoning, or translation—the Chinese models are competitive enough that price, speed, and specific features become the only real differentiators.Are you looking for the best model for a specific task, like coding or customer service?
+May 26, 20260Who is building chinas data center?China's data center construction landscape is a dynamic ecosystem driven by three major forces: telecommunications giants, cloud hyperscalers, and third-party IDC specialists. Here is who is building the country's digital backbone.🏗️ The Main Players in China's Data Center SectorThe industry is led by three core groups, each playing a distinct role -4:Player TypeKey EntitiesPrimary Role & FocusTelecom OperatorsChina Telecom, China Mobile, China Unicom -8-10The backbone of national infrastructure. They own vast network resources and land banks, leading in colocation data centers with over 1,500 MW of power capacity nationwide -10. China Mobile, for example, is pushing large-scale projects like the Yangtze River Delta (Yangzhou) center -1-9.Cloud & Tech GiantsAlibaba Cloud, Huawei Cloud, Tencent Cloud, JD Cloud -2-6-8The primary drivers of AI demand. They build massive "hyperscale" facilities optimized for their own ecosystems (e.g., Alibaba's "Tongyi" Qwen, Huawei's Ascend chips). These companies are also aggressively expanding their global footprint -6.Third-Party IDC SpecialistsConstruction & Engineering: China Construction, China Electronics Engineering Design Institute -10Operators: GDS, Chindata, VNET, Sinnet, 21Vianet -4-8-10The expert builders and niche operators. Firms like China Construction provide the heavy lifting for physical construction. GDS, Chindata, and VNET operate high-performance data centers, often serving as critical "carrier-neutral" hubs for enterprises that do not build their own -4-8-10.⚙️ Specialized Roles: The Builders and EngineersBeyond the headline investors, specialized firms are crucial to building the actual facilities.The Infrastructure Architects: China Construction and its subsidiary China Construction Technology are critical engineering forces. China Construction Technology has designed and built over 2,000 projects (including for China Construction Bank and JD.com) and is leading the charge on green technologies like "zero-carbon" data centers and liquid cooling systems -5. Other major contractors include AECOM and China Railway Construction-10.The Supply Chain (Inside the Box): While not building the walls, hardware vendors are essential. Huawei, Inspur, Lenovo, and ZTE provide the servers, network gear, and cooling systems. For instance, ZTE led a ¥1.01 billion project for China Mobile, handling 30% of the engineering work -1-10.📊 Ranking the Market Leaders (2026)Based on an industry ranking of AI Data Center (AIDC) capabilities, the competitive landscape for 2026 is as follows -8:RankCompanyKey Strength1Huawei CloudFull-stack integration (Ascend chips + cloud)2China TelecomVast national network and colocation assets3Alibaba CloudMassive scale, leading cloud market share4China MobileStrong government-backed infrastructure projects5ZTELeading infrastructure engineering partner🤔 How They Differ from the U.S.?State-Coordinated Strategy: China's "Eastern Data, Western Computing" project deliberately steers investment, shaping a more distributed national grid rather than the purely market-driven US model -10.Domestic Supply Chains: Due to US chip sanctions, Chinese builders (like Huawei) focus on creating self-sufficient "stacks" of domestic hardware, whereas US builders heavily rely on NVIDIA and international vendors.Does this help clarify the major players? Are you interested in the specific investment plans of a company like Alibaba or China Mobile?
+May 26, 20260Specific investment plans of major Chinese hyperscalersHere are the specific AI infrastructure investment plans of China's major hyperscalers, based on the latest official announcements and financial reports.🏢 Alibaba: The All-In BetAlibaba has announced the most aggressive investment plan, with a clear focus on full-stack self-sufficiency.Investment Scale: Plans to invest at least 380 billion yuan (approx. $56 billion USD) in cloud and AI infrastructure from 2026 to 2028 -1-10. This sum exceeds its total cloud/AI spending over the entire prior decade, and the final figure is expected to be far higher-6.Current Spending: Capex reached 126.1 billion yuan in FY2026, up 46.6% year-over-year, with Q4 spending at 26.9 billion yuan -5.Why the urgency? Alibaba's CEO confirmed every available AI server is already running at full capacity-10.Strategic Priorities:Self-Developed Chips: Its in-house "Pingtouge" AI chips (Zhenwu series) are already deployed and driving a growing portion of cloud services -10. Over 60% of its AI compute capacity now serves external business customers -6.MaaS (Model-as-a-Service): The "Bailian" MaaS platform is the key revenue driver, with customer count growing 8x year-over-year-5. AI model ARR (annual recurring revenue) is on track to hit 30 billion yuan by year-end -6.Full-Stack Synergy: The "TongYi Lab (models) + Cloud + Pingtouge (chips)" structure creates a vertically integrated feedback loop from silicon to applications -1.💻 Tencent: The Pragmatic AcceleratorTencent is ramping up spending in a more phased manner, heavily focused on deploying domestic chips to unlock cloud capacity.Investment Scale: Q1 2026 capex reached 37 billion yuan (payment basis), with 31.9 billion yuan recorded in the quarter—up 63% sequentially and hitting a single-quarter record-2-5-7.Full-Year Projection: Analysts forecast total 2026 capex could nearly double 2025's level to exceed 165 billion yuan by 2027 -10.Strategic Priorities:Ending the GPU Shortage: The biggest bottleneck is insufficient GPUs to meet customer demand. The strategy is to aggressively integrate domestic GPUs (from Huawei and others) starting in H2 2026 to expand capacity -2-7.Monetizing the Investment Portfolio: CFO has noted the company is "accelerating the monetization process of some investment portfolios" specifically to fund AI expansion -2.Embedding AI into Ecosystem: The "Hunyuan 3.0" model is being integrated across products (Yuanbao, WorkBuddy), with token usage increasing tenfold from the previous version -2.🔍 Baidu: The AI-First PivotBaidu has transformed into an AI-first company, with over half its revenue now coming from AI-related businesses.Investment Scale: AI business revenue reached 13.6 billion yuan in Q1 2026, accounting for 52% of core business revenue -4.Kunlun Chip Deployment: Its in-house Kunlun P800 chips are now powering production. In Q1 alone, Baidu delivered multiple 10,000-card clusters, and these domestic chips completed the core training for the Wenxin 5.1 model -4.AI Cloud Leadership: Baidu AI Cloud revenue grew 79% year-over-year to 8.8 billion yuan, winning 25 major government/enterprise projects worth 1.25 billion yuan—nearly 5x the amount of its closest competitor -4.🐝 Huawei & The Rise of Domestic InfrastructureWhile Huawei Cloud is a hyperscaler, Huawei's broader role as the foundational chip supplier is shaping all the above plans.The Ascend Engine: Huawei is aggressively ramping production of its Ascend AI chips, which are the primary "domestic GPU" that Alibaba and Tencent are adopting -3-8.Production Targets for 2026:Ascend 910C: Approximately 600,000 chips (roughly double 2025's output) -8.Ascend 950 PR: Approximately 750,000 chips, with mass production starting mid-2026 -8.The "Token Factory" Model: Huawei is commercializing the "Token Factory" concept in partnership with local governments. A flagship project in Wuxi, backed by a 16.5 billion yuan government initiative, will use Huawei's Ascend 384 super-node clusters to provide token-generation-as-a-service -3.📝 Summary TableCompanyKey Investment (2026)Primary Hardware StrategyCommercial FocusAlibaba380B+ yuan over 3 years"Pingtouge" self-developed chipsMaaS platform (Bailian) & full-stack synergy -1-6TencentRecord Q1 capex (37B yuan); 63% sequential growthSwitching to domestic GPUs (Huawei)Embedding AI into ecosystem (WeChat, WorkBuddy) -2-7BaiduAI business now 52% of revenue"Kunlun" chips (P800 series)AI Cloud dominance & model training -4Huawei1.35M+ Ascend chips in 2026Ascend 910C/950 (domestic supply chain)"Token Factory" infrastructure & chip provider -3-8Would you like a deeper look into the financial performance of any of these companies, or more detail on the specific "Token Factory" model Huawei is deploying?
+May 26, 20260Chinese people are not willing to pay for saas, how does Chinese models plan to monetize ?The observation that Chinese consumers have a low willingness to pay for software subscriptions is accurate, and AI labs have built their monetization strategies around this reality. Rather than fighting this habit, they are bypassing the consumer subscription model almost entirely.Instead of asking users to pay, Chinese AI labs are turning their models into a "Token Factory" —a utility that monetizes at the infrastructure level, not the application level -7.Here is how they plan to monetize:1. MaaS APIs: The Primary Engine (B2B)This is the "default" successful model in China. Instead of selling software, they sell intelligence as a utility (tokens) via APIs. This allows them to capture value from enterprises and developers, bypassing the individual consumer entirely -6.Proof of Success: Zhipu AI is leading this charge. Its MaaS API platform reached an ARR of RMB 1.7 billion (approx. $250M) in 2025, growing 60x year-over-year -2-8.Pricing Power: Unlike typical price wars, strong models are raising prices. Zhipu raised API prices by 83% in early 2026, yet demand remained "supply exceeds demand," with calls growing 400% -4-9.2. The "Trojan Horse" Strategy (Consumption Zones)Rather than charging for the AI chat, labs integrate the AI into shopping or productivity scenarios where users are already willing to spend money.E-commerce Integration: Doubao (ByteDance) and Kimi embed shopping links (e.g., to Douyin/TikTok shops or Taobao) directly into chat replies -5. They take a cut of the transaction. The AI becomes a sales channel, not a subscription fee -1.Productivity Bundling: Alibaba and Baidu are using MaaS to boost their Cloud revenue. Enterprises pay for cloud computing, and the AI is the value-added feature driving that spend -1.3. Vertical Agents with Clear ROI (Coding & Enterprise)For B2B, they are selling labor replacement, not just software. They charge for tasks completed, not just time used.AI Coding: This is the "killer app." Zhipu’s coding plan saw developer subscriptions hit 242,000, and they raised prices by 30% (usage still surged) because it saves engineering hours -4-8.Long-Horizon Tasks: New "Claw Plans" (AI agents that work for hours autonomously) saw subscriptions exceed 400,000 in 20 days—validating that businesses will pay for results -8.4. Consumer "Freemium" to Filter Power UsersWhile most users stay free, labs are quietly adding high-end tiers for heavy users to capture surplus value from the top 1% without scaring away the masses.Doubao’s Tiered Plan: Despite massive free usage, Doubao just launched three paid tiers (up to $70/month) focused on complex tasks like data analysis and video generation—areas a casual user doesn't need -7.Telecom Bundles: Even traditional carriers like China Telecom are joining, offering token packages (e.g., 10M tokens for ~$1.45) bundled with connectivity -10.💡 The VerdictChinese AI models are monetizing by turning intelligence into an invisible utility (tokens) rather than a visible subscription (SaaS). They make money when a developer builds an app, a shop sells a product via AI recommendation, or a coder automutes a task—not when a user simply "chats."Does this distinction between "utility monetization" (China/APIs) and "subscription monetization" (US/ChatGPT) clarify the strategy for you?
+May 26, 20260How much ai revenue is generated from telecom selling?Based on the available data, there is no precise, single figure for "AI revenue generated from telecom selling" because the carriers have not yet broken out "Token sales" as a standalone line item in their financial reports. However, we can measure the rapid growth of the broader AI and computing services categories that include these token packages.The table below summarizes the relevant AI-driven revenue reported by each carrier in their latest financial results:CarrierRelevant AI/Computing Revenue (Latest)Growth RateReported PeriodChina MobileRMB 90.8 billion (from AI services)+5.3% YoYFull Year 2025China UnicomRMB 15.4 billion (from Computing Power Business)+8.3% YoYQ1 2026China TelecomNot separately disclosed (included in "Smart Revenue")+39.4% YoYQ1 2026📈 Market Trends and Strategic FocusWhile the specific revenue from token packages is not yet public, the strategic importance of AI to their business models is unmistakable. The carriers see the sale of AI services as the primary engine for future growth, as traditional communication services face challenges -3-6.Revenue Diversification: The pressure to find new revenue streams is real. Carriers are facing headwinds like VAT hikes on core services and slowing growth in traditional telecom -1-3. For instance, China Mobile's combined revenue from computing and AI services already accounted for 20.2% of its principal business revenue in 2025-8.Massive Infrastructure Investment: All three carriers are reallocating their capital expenditures (CAPEX) to prioritize AI.China Mobile plans to increase its computing network investment by a staggering 62.4% in 2026 -5-8.China Unicom will allocate over 35% of its RMB 50.0 billion CAPEX to computing -3.China Telecom is increasing investment in its intelligent sector by 26%, with AIDC investment rising 28%-7.The Shift to a "Token" Business Model: The carriers are explicitly pivoting to a "Token-first" strategy. China Telecom, for example, has stated it aims to become a leading AI service provider with "Token services as its main business focus" -1-4. This indicates that selling AI access via tokens is seen as a core long-term business, even if its current revenue is embedded in broader categories.💡 Conclusion: What We Know and Don't KnowTo summarize, while you can't yet point to a specific "token revenue" number, the figures above show that China Mobile is generating tens of billions from AI services, and all three carriers are growing these segments rapidly. They are transforming from simple "pipeline" connectivity providers into "platform" AI service providers -4-8.This means the revenue from the token packages you asked about is included within the reported growth numbers for "AI services," "Computing Power," and "Smart Revenue," and is a key driver of their optimistic forecasts.Would you like a deeper look into China Mobile's reported AI revenue structure to understand what is included in that RMB 90.8 billion figure?
+May 26, 20260Where is the demand for Chinese models coming from geographicallyGeographically, demand for Chinese AI models comes from a two-speed global market. They have captured a dominant share of price-sensitive global developers (led by the US) via API aggregators, while facing significant barriers to large-scale Western enterprise adoption.Here is the geographic breakdown of demand:🌍 Tier 1: Global Developers (The "Spot Market" - Dominant)This is the primary source of current demand. Through platforms like OpenRouter, developers can access Chinese models at extremely low prices.User Base: OpenRouter's users are 47% American, while Chinese users are only 6%-8-10.Market Share: Chinese models captured 36% of global token volume on aggregators by April 2026, surging from just 10% in January -1.Why: They offer 80-90% of the performance of US models for about 20% of the cost-1. API prices can be 1/16th of US rivals-8-10.🏢 Tier 2: Direct Enterprise Adoption (Limited)Direct use by large Western companies remains restricted, though opportunities exist elsewhere.Western Markets (EU/US): Growth is blocked by compliance, data residency, and geopolitical friction. Western firms are hesitant to use Chinese APIs directly -1.Emerging Markets: Adoption faces less friction in regions like the Middle East, Southeast Asia, and the Global South, which have closer ties to China -1-2. Chinese models have captured over 10% penetration in 30 countries and 20% in 11 countries-2.🌏 Tier 3: The Domestic Chinese MarketThis remains a foundational pillar, but the current global surge is driven by overseas users.Local Dominance: Chinese models lead in Chinese language and local business workflows -3.Recent Data: As of May 2026, China's weekly AI token consumption was 9.2 trillion vs. the US's 4.9 trillion-4.💡 Geographic Insight: The DeepSeek AnomalyDeepSeek is unique. According to a16z data, its web traffic distribution is:China: 33.5%Russia: 7.1%United States: 6.6%This shows it is the only model truly spanning the US, China, and Russian markets -5.In summary, the current demand surge is a global developer phenomenon driven by cost, not a broad-based enterprise shift. While US developers are heavy users of Chinese models via aggregators, most large Western companies are not yet directly deploying them.Would you like to explore the specific cost advantages that make Chinese models so attractive to global developers?
+May 26, 20260China ai path to revenue and profitabilityChina's AI industry has rapidly shifted from a "technology chase" to a "commercialization imperative," and the path to revenue and profitability is now taking clear shape -1-6. While profitability remains elusive for many, top-tier firms are starting to deliver real financial results, and analysts predict a potential inflection point could arrive sooner than the internet era did -1.Here is a breakdown of how Chinese AI companies are building their path to profitability.🚀 1. The Monetization Playbook: Key StrategiesCompanies are moving beyond free offerings with a multi-pronged approach centered on four main strategies:Paid Subscriptions (Freemium 2.0): The era of completely free access is ending. ByteDance’s Doubao, China's most-used AI app with 345 million monthly users, has rolled out a tiered subscription model at ¥68, ¥200, and ¥500 per month to cover high inference costs -4-9. This strategy filters casual users to focus on those with genuine productivity needs -9.Model-as-a-Service (MaaS) & Cloud Bundling: The most immediate revenue stream is integrating AI with cloud computing. Alibaba Cloud’s MaaS platform "Bailian" saw its customer base grow 8x year-over-year, and its AI model & application ARR has already exceeded ¥8 billion (approx. $1.1 billion), on track to hit ¥30 billion by year-end -3-8. Similarly, Tencent is transforming its cloud offerings from raw computing power to a full suite of "AI industrialization capabilities" like agent development platforms -5-10.AI-Powered Ads & Existing Business Boost: Integrating AI into mature, massive user bases provides an immediate profitability boost.Tencent upgraded its ad recommendation model with AI, and its AI-powered "AIM+" smart ad matrix now covers about 30% of advertiser spending, helping drive a 20% year-over-year growth in marketing services revenue -5-10.Alibaba is integrating its "Tongyi Qianwen" AI into its e-commerce ecosystem, offering shopping assistants for consumers and customer service agents for merchants -8.Vertical & Enterprise Agents (The "Agentic AI" Bet): The biggest bet is moving from tools to "agents" that complete tasks autonomously. Tencent's WorkBuddy (productivity) and CodeBuddy (coding) agents are seeing strong traction, with WorkBuddy’s DAU (Daily Active Users) paid conversion rate significantly higher than traditional products. This validates that Chinese enterprises and professionals are willing to pay for AI that directly delivers labor or results, not just software -10.📈 2. Key Financial Metrics: From Investment to ReturnsThe narrative is supported by concrete financial data signaling a clear transition.MetricAlibaba (Q4 FY2026) -3-8Tencent (Q1 2026) -5-10Cambricon (FY2025) -2AI Revenue Contribution30% of Cloud external revenue (≈ ¥8.97B Q4). Forecast to exceed 50% in next year.N/A (Rapid growth in Cloud & Enterprise, up 20%)¥6.5B total annual revenue (up from ¥1.2B in 2024)Profitability BenchmarkAI Model ARR: ¥8B, targeting ¥30B by Dec 2026. (High-margin revenue)Operating Profit Impact: Explicitly stated that without AI product investments, Op Profit would have been ¥84.4B (+17%), not ¥75.6B (+9%).Net Profit: First annual profit of ¥2.1B (vs. a ¥452M loss in 2024).Investment & StrategyFull-stack AI: In-house "Pingtouge" GPUs (47k chips delivered), "Tongyi" models, cloud infra.Heavy R&D/CAPEX: ¥22.5B R&D (+19%) & ¥31.9B CAPEX (+16%) in Q1. Prioritizing internal AI use over short-term compute rental.Chip Self-Reliance: Riding gov't push for domestic AI chips; plans to deliver 500k AI accelerators in 2026.💡 3. Why This Path is AcceleratingRadical Cost Efficiency: Chinese AI firms have successfully driven down inference costs to only 15–20% of their U.S. counterparts through architectural innovations like Mixture-of-Experts (MoE) and model distillation. This efficiency enables competitive pricing while protecting margins -6.Shift from "Price War" to "Pricing Power": After a brutal 2024 price war, the market is rationalizing. Leading players (Alibaba, Baidu, Tencent, Z.ai) are now raising prices on new, more capable models. API input prices have climbed about 80% between Q2 2025 and Q1 2027, reflecting a move from commoditization to value-based pricing. For example, Z.ai's GLM-5 model is priced over 200% higher than its predecessor -6.The Profitability Funnel Effect: The AI industry is consolidating. As capital shifts from "technology worship" to "commercialization capability," a winner-take-most dynamic is emerging. Smaller players are being squeezed, while market leaders with massive user bases (Alibaba, Tencent, ByteDance) and unique value propositions are capturing an ever-larger share of revenue and profits -1.In summary, China's AI path to profitability is no longer hypothetical. It is being forged through pragmatic subscription models, deep integration with profitable cloud and advertising ecosystems, and a strategic pivot to high-value enterprise agents. The shift from technology chase to a disciplined business focus is already translating into real financial results for market leaders.Would you like a deeper look into the specific financials of Alibaba, Tencent, or another player in this space?
+May 26, 20260Then how does Chinese consumers spend tokens when they are sold from telecom?Chinese consumers don't need a separate credit card or app to buy AI tokens. By bundling tokens into monthly phone bills, the three major telecom carriers have turned AI into a utility that ordinary people can pay for just like data or electricity -1-4.Here is how the mechanics actually work.⚙️ The Mechanics: "Buying" vs. "Spending"The process mirrors how you buy mobile data, creating a seamless cycle:Purchase (The "Top-Up") : You buy a "Token Package" from China Mobile, China Unicom, or China Telecom via their official app or website. Payment is deducted directly from your phone bill balance or added to your monthly invoice -2-3.The Transfer (API Key) : After purchase, you receive an API Key (a unique code) or link your phone number to specific AI applications -3-9.Spending (The "Meter") : When you use an AI agent (like OpenClaw, TeleClaw, or Tencent WorkBuddy), the software automatically deducts Tokens from your package balance for every task you run. For context, a single AI-powered PPT generation or document analysis can cost thousands of tokens -3.📊 What Consumers Get: The 2026 Price ListThe market has settled into clear tiers. Here is what the "Token as Currency" exchange rate looks like right now -1-6:ProviderPlan NameMonthly Cost (USD)Token AllowanceChina TelecomLight~$1.40 (9.9 RMB)10 MillionChina TelecomPro~$7.35 (49.9 RMB)80 MillionChina MobileEntry (Pay-as-you-go)~$0.14 (1 RMB)400,000 (One-time)China MobileStandard~$3.70 (24.99 RMB)10 MillionChina UnicomStarter~$2.20 (15 RMB)6 Million💡 Why This "Telecom Model" WorksSince Chinese consumers are historically reluctant to pay for software subscriptions (SaaS), this model bypasses that resistance entirely -7.Psychological Comfort: You aren't "buying software"; you are "topping up your phone," a behavior Chinese users have performed for decades -4.No Friction: Payments are deducted from the pre-paid phone balance, eliminating the need for credit cards or complex app store setups -2.The Math: For 9.9 RMB ($1.40), you get roughly 10 million tokens. This is enough for about 5,000 simple Q&A sessions with an AI, or roughly 500-700 pages of translated text -7-4.🏢 The Split: High-Value vs. Low-Value UseThis model effectively separates the market into two segments:Casual Consumers: They use free AI apps for daily chats, while their telecom carrier quietly counts the "free" tokens behind the scenes.Heavy Users (Prosumers/SMEs) : For those using AI for coding, video generation, or data analysis, free tiers run out quickly. The token packages offer a 10x to 20x cost reduction compared to pay-as-you-go rates. Insiders note that government entities and state-owned enterprises are currently the major customers of these packages -5.In short, the logic is utility payment. The telecom bill has replaced the SaaS subscription, turning AI compute into a fungible commodity anyone can buy with a text message.
+May 26, 20260Are Chinese models running out of compute capacity?No, Chinese AI models are not "running out" of compute capacity in a way that halts operations, but they are facing a severe and intensifying supply-demand imbalance. The core issue is that surging demand is outpacing the delivery of hardware infrastructure-2-6.Here is a breakdown of the situation:📈 The Core Issue: A Demand ExplosionThe explosive growth in Token consumption—up over 300% since early 2026—is driven by a shift from simple chatbots to complex AI Agents-1-3. These tasks require 5-8x more Tokens per request and have pushed GPU cluster utilization above 95%-1-3.🚧 The "Compute Gap": Bottlenecks in Hardware DeliveryThe problem is less about a complete lack of chips and more about delivery bottlenecks.Stalled High-End Imports: While the US has licensed H200 exports, regulatory delays have halted shipments. Beijing has also discouraged over-reliance on foreign chips to boost domestic production-4-7.Domestic Production Limits: National champions like Huawei face constraints in advanced process nodes, affecting supply-7.Supply Chain Crunch: Shortages have spread to critical components like optical modules and PCBs. Lead times for some AI accelerators have stretched to 14 months-1-2.⚙️ A Tale of Two Use Cases: Training vs. InferenceTraining (Building models): Remains the biggest bottleneck. Chinese labs still rely almost entirely on Nvidia GPUs to train flagship models, as domestic alternatives struggle with stability for massive training runs.Inference (Running models): This is the strategic battleground. To circumvent hardware limits, firms are using software to optimize for delivering AI services—an area where Chinese models are becoming highly competitive-7-8.💡 The Solutions: Software Innovation & Price SignalsChina is overcoming these gaps through aggressive software innovation.Software Optimization: Techniques like knowledge distillation allow smaller models to achieve 92% of the accuracy of massive ones while running 11x faster, easing inference pressure-1. New frameworks like MEPipe optimize memory usage on limited hardware.Market Reality: The scarcity is so real that major cloud providers (Tencent, Alibaba) have raised service prices, and GPU rental rates have surged-6-9.In short, Chinese AI labs aren't frozen for lack of compute, but they are operating in a high-pressure environment. Innovation is being driven as much by necessity (working around hardware constraints) as by ambition.I hope this clarifies the current situation! Would you like to dive deeper into the specific software techniques being used to get around these hardware limits?
+May 26, 20260How much demand for ai compute demand that Chinese hyperscalers internal demandChinese hyperscalers are currently operating at maximum capacity, and crucially, this demand is coming overwhelmingly from their own internal ecosystem, not external customers.Here is the breakdown of who is consuming this compute internally:🏢 The "Self-First" StrategyThe primary driver of demand is the companies' own product suites. They are deliberately prioritizing internal use due to compute scarcity.Tencent is the clearest example: In a strategic trade-off, Tencent has actively chosen not to rent out its GPU capacity on Tencent Cloud to external clients -8. Instead, it reserves its scarce supply for its own fleet of AI products including WeChat AI, Yuanbao, Hunyuan model training, and Gaming/AI Ads-3-8.The Financial Impact: This internal focus is expensive. In Q1 2026 alone, new AI products like Hunyuan, Yuanbao, and CodeBuddy dragged down Tencent’s profits by approximately RMB 8.8 billion-3.🔬 The "AI Factory" Model (Alibaba)Alibaba is taking a manufacturing approach, building "AI factories" (training and inference centers) to fuel its growth -2.Full Utilization: CEO Wu Yongming recently stated, "There is not a single idle card in Alibaba's servers" -2-3.Internal Revenue Driver: While Alibaba does serve external clients (AI revenue is set to hit RMB 10 billion ARR in mid-2026), the bulk of its MaaS (Model-as-a-Service) income is currently driven by its self-developed models (Tongyi Qianwen), not third-party open-source models -2.🎯 Internal Ecosystem Integration (Apps & Ads)Beyond "cloud" and "models," AI is being embedded into existing cash-cow products.Baidu: While Baidu Cloud is growing, AI is rapidly replacing search. AI business now accounts for 52% of Baidu’s core revenue, up from 43% last quarter. This is driven by integrating AI into their "Big Search" business and digital humans -4.Tencent Ads: AI-driven ad recommendations now power approximately 30% of Tencent’s ad spend, providing an immediate and high-margin return on their internal compute investment -3.📊 Quantifying the Internal DemandThe numbers explain why there is no capacity left for outsiders:Token Explosion: Daily token consumption in China is exploding. Between November and May alone, Alibaba's token demand grew 10x-2.Supply Constraints: The shortage is real. Tencent specifically noted they are facing "difficulties meeting external customer demand" for cloud services due to the internal GPU crunch -3.💡 When Will They Sell to Outsiders?The strict internal prioritization is expected to ease in the second half of 2026, primarily due to the arrival of domestic AI chips (ASICs) from Huawei and others -3-8.Short Term: Most hyperscalers are in "capacity acquisition" mode. They are spending tens of billions to build data centers just to keep up with their own needs.Long Term: Once domestic chip supply stabilizes, these companies plan to unlock significant capacity to sell to external enterprises -8.In short, the hyperscalers are currently their own biggest customers. Would you like to know which specific applications (like coding or advertising) are driving the highest return on this internal compute investment?
+May 26, 20260How are Chinese model able to serve such a huge token demand given they have limited hardwareChinese AI companies have successfully turned the challenge of limited hardware into a major competitive advantage through relentless architectural and software innovation. They serve enormous token demand by being radically more efficient, not by having more powerful chips-5.Here are the key strategies bridging the hardware gap:⚙️ 1. Architectural Efficiency: The "Mixture of Experts" (MoE) RevolutionInstead of activating a whole model for every task, MoE models only use the "expert" sub-networks needed for a specific query.How it works: A model like MiniMax M2.5 has 230 billion total parameters, but only activates 10 billion per token. This makes it as smart as a massive model but as agile as a small one-5.The Result: This allows for ultra-low pricing. For instance, MiniMax M2.5 can cost as little as $0.15 per million input tokens, which is roughly 1/30th the price of comparable Western models-5.📉 2. Model Distillation: Creating "Mini-Me" SpecialistsDistillation is the process of training a smaller, faster "student" model to replicate the complex reasoning of a much larger "teacher" model-3.The Strategy: Instead of running a massive, slow model for every request, companies create compact, specialized models (as small as 0.6B to 3B parameters) that run extremely fast on limited hardware-3-8.The Benefit: These smaller models can achieve over 44,000 tokens per second on a laptop, making real-time AI on phones and edge devices possible without cloud dependency-8. Crucially, new techniques (like TCOD) solve the problem of "distilled" models failing on complex multi-step tasks, making them viable for real work-7.🏭 3. Infrastructure & Hardware Innovation: "Good Enough" at ScaleChinese firms cannot rely on high-end Nvidia H100s, so they have redesigned the software stack to maximize the chips they can get (like the restricted H800 or domestic Huawei Ascend chips)-1-5.GPU Pooling (Alibaba's Aegaeon): Alibaba found that most models sit idle on dedicated GPUs. Their "pooling" system serves hundreds of models from a single pool, allowing them to reduce GPU usage by 82% (saving nearly 1,000 GPUs) while increasing output 9x-9.Inference Optimization (Huawei's UCM): AI inference involves a massive memory bottleneck (KV Cache). Huawei's Unified Cache Manager (UCM) cleverly shuttles data between different types of memory (HBM, DRAM, SSD), resulting in 2-22x improvements in tokens-per-second in long tasks-6.🏆 A Real-World Example: GLM-5Z.ai's GLM-5 is a flagship model that demonstrates the success of this approach. It was trained entirely on Huawei Ascend chips (a domestic, less powerful alternative to Nvidia) using the MindSpore framework, with zero dependency on Nvidia hardware-1.Despite the hardware constraint, GLM-5 outperforms GPT-5.2 on complex coding benchmarks (SWE-bench) and is priced 5-8x cheaper than comparable US models-1.In summary, the US focuses on brute-force single-chip power, while China excels at system-level efficiency and clever software. For many real-world applications, "good enough" hardware with world-class software is beating "best-in-class" hardware with standard software.Are you interested in a deeper look at how these specific strategies, like knowledge distillation or GPU pooling, could impact your own projects?
+May 26, 20260USA vs China gpu chips and how far behind is China from USA in chip innovationBased on data from 2025-2026, China is currently 2 to 5 years behind the US in cutting-edge AI chip design and manufacturing. While the US leads in raw performance and manufacturing technology, China is rapidly closing the gap through system architecture innovation and volume.💻 Single Chip Performance: The Specs GapFor a single chip, the US maintains a clear lead in raw compute, but Chinese chips are narrowing the gap in specific categories like AI inference.Feature / MetricUS (Nvidia)China (Huawei Ascend)GapFlagship AI ChipB200 (Blackwell) / H100Ascend 910C / 950Compute (BF16)~2,500 TFLOPS (B200)~780 TFLOPS (910C)~3.2x behind-3Memory Bandwidth8.0 TB/s (H100)3.2 TB/s (910C)~2.5x behind-3Gaming GPURTX 4090 / 5090Lisuan LX-7G100~30% slower than RTX 4060 -2Manufacturing Node4nm (TSMC N4)7nm (SMIC N+2)~2 generation gap-5Note: Nvidia’s flagship AI chips for 2025-2026 are built on TSMC’s 4nm-class process. Huawei’s 910C uses SMIC’s 7nm-class “N+2” process. The Lisuan gaming card runs at ~485butperformsworsethana485 but performs worse than a 485butperformsworsethana300 RTX 4060 -2.🏭 Manufacturing: The Lithography WallThe single biggest bottleneck is lithography equipment.The US/Allied Advantage: ASML (Netherlands) has a monopoly on Extreme Ultraviolet (EUV) lithography, essential for sub-7nm chips. Due to US pressure, China cannot buy these machines-9.China’s Reality: SMIC, China’s best foundry, is stuck at 7nm-class using older Deep Ultraviolet (DUV) tools. Each DUV machine can produce 5-10x fewer advanced chips than an EUV machine, making mass production inefficient -5-9.🧠 The "System" Solution: How China CompensatesUnable to make a single chip as powerful as Nvidia's, China is connecting many weaker chips together to compete at the system level.Recent benchmarks show Huawei’s Cloud Matrix 384 (connecting 384 Ascend 910Cs) actually outperforms Nvidia’s GB200 NVL72 in certain multi-chip tasks like memory capacity and scale-up bandwidth -3-8:System Metric (Per Rack)Nvidia GB200 NVL72Huawei Cloud Matrix 384ResultTotal Memory13.8 TB49.2 TB3.6x More-3Total Bandwidth576 TB/s1,229 TB/s2.1x Faster-3Scaling Limit72 GPUs384 GPUs5.3x Larger-3Power Efficiency0.81 W/TFLOP2.00 W/TFLOPUS is 2.5x Efficient-3The Trade-off: This brute-force linking consumes 4.1x more power than Nvidia's solution. For massive AI data centers, electricity costs are becoming the new battleground -1-3.🔄 Software & Innovation: The Race is ChangingRaw hardware is only half the battle; the software ecosystem is the other.Software Moat (CUDA): Nvidia's CUDA platform is the industry standard. China is playing catch-up, with Huawei’s ecosystem about 5-10 years behind in developer maturity -7.The Efficiency Workaround: Chinese firms (like DeepSeek) are using "distillation" and other algorithmic tricks to run high-end AI models on less advanced chips, effectively using software to mask hardware limits -1.Future Horizon: China is aggressively investing in 2nm chip design (Dishan Technology) and domestic lithography, aiming to break the EUV bottleneck by 2030-10.In summary, the US leads in premium, standalone chips, but China is closing fast in real-world AI deployment. Are you more interested in the financial performance of these Chinese chip giants or the specific details of the US export controls?
+May 26, 20260china vs us ai models revenueHere’s the English version of the revenue comparison between Chinese and US AI models:The revenue gap between Chinese and US AI models is far wider than the market share difference. In short: The US has achieved hundreds of billions in AI revenue, while China is still searching for profitable models.💰 Core Revenue Comparison (Q1 2026)US (OpenAI)Quarterly revenue: 57billion∗∗Annualizedrevenue:∗∗57 billion** Annualized revenue: **57billion∗∗Annualizedrevenue:∗∗250–300 billionNote: Still operating at a large loss, but user willingness to pay is high.US (Anthropic)Quarterly revenue: 48billion∗∗Annualizedrevenue:∗∗nearly48 billion** Annualized revenue: **nearly 48billion∗∗Annualizedrevenue:∗∗nearly450 billionGrowing rapidly, expected to turn profitable soon.China (Alibaba Cloud)Quarterly AI revenue: ¥9 billion (approx. $1.24 billion USD)Seeing triple-digit growth for 11 straight quarters, mainly driven by computing power and cloud infrastructure.China (Independent model companies)Combined annual revenue of leading firms (Zhipu, MiniMax, etc.) is under $500 million — orders of magnitude behind OpenAI.🆚 Why such a huge gap?US: Value-based pricing → Revenue comes from direct user subscriptions (ChatGPT has ~55 million paid users) and enterprise software services. Pure software value capture.China: Cost pass-through model → Most consumer services are free (e.g., Doubao has 345 million MAUs, all free). B2B revenue often comes from selling compute and cloud bundles — more like IT infrastructure than AI software.Extreme price war → DeepSeek’s API price is 1/20 to 1/50 of OpenAI’s. China has ~4x the API call volume of the US, but earns far less total revenue due to razor-thin margins.🔮 What’s next?China is entering a commercial validation phase — investors now demand real revenue growth, not just user numbers.DeepSeek is taking a different path: ultra-low pricing + open-source to build an ecosystem around domestic AI chips, aiming for long-term hardware+software dominance rather than short-term API revenue.Would you like a deeper breakdown of a specific company (like Alibaba or DeepSeek) or their B2B versus B2C revenue mix?