What happened?

According to Bloomberg, DeepSeek plans to deploy at least 160,000 of Huawei's next-generation Ascend-950DT chips in a large data centre in Inner Mongolia. If built, it would be the largest known Huawei chip cluster.

One detail sets an important limit: the chips would run inference only, not training. For the much heavier training workloads DeepSeek still uses Nvidia hardware.

Why does that distinction matter?

Inference and training are two ends of the same job. Training is the stage that builds a model from scratch, can run for weeks and puts the heaviest load on hardware; inference is running the finished model to produce answers for users. Inference asks less of the hardware — memory bandwidth and stability are enough, without the vast synchronised computation training requires.

So the headline "China is weaning itself off Nvidia" is premature here. The deployment also marks the boundary of what China can do with its own silicon: domestic hardware is enough to serve answers, not yet enough to raise the model.

There is a product side to the choice as well. DeepSeek publishes its models with open weights and prices its API far below rivals; that business model means a very large and continuous inference load. For a company selling cheap tokens, moving inference onto domestic hardware is a direct gain on the cost side — and it makes supply predictable under export controls.

The obstacle ahead: memory

Huawei looks unlikely to deliver the full order within a year. The reason is twofold: production capacity and a shortage of memory chips.

  • China's largest memory maker, CXMT, has begun producing HBM3E — the high-speed memory that powers most AI processors — in small batches for the first time.
  • But CXMT is still three to five years behind Samsung, SK Hynix and Micron, all three of which are already mass-producing the next generation, HBM4.
  • The bottleneck is therefore not the processor but the memory feeding it, and that gap has not closed.

In the wider picture

DeepSeek's order reads less as one company's preference than as part of the Chinese government's push to grow its own chip industry without falling behind in AI. The balance here is delicate: moving to domestic silicon reduces strategic dependence but means giving up performance in the short term.

The number itself conveys the scale. 160,000 accelerators sits near the upper bound of what a single facility can host — and all of it would be dedicated to producing answers alone. That is also an indicator of how far inference demand in China has grown.

What's next?

Neither DeepSeek nor Huawei has publicly confirmed the plan; the report rests on Bloomberg's sources, and the figure describes an intended order rather than an installed facility. No timetable has been announced, and Huawei's delivery capacity will be decisive. The number to watch is not the chip count but whether CXMT can move HBM production from small batches to mass manufacturing; that will determine when the cluster comes online.