CourionAI
EN
Newsletter
← All news
china 3 min read

Alibaba Wants a Ten Trillion Parameter Model, and Its Own Chips to Train It

At its Apsara conference CEO Eddie Wu said Alibaba is planning a Qwen successor of 5 to 10 trillion parameters and unveiled the Zhenwu V900 chip, which the company says triples its predecessor's performance.

A microchip seen from above rendered as a dense city grid with towers rising at its centre

At Alibaba Cloud’s annual Apsara conference in Hangzhou, chief executive Eddie Wu laid out a plan with two halves that only make sense together. The first is a model: a successor to Qwen in the range of 5 to 10 trillion parameters, against roughly 2.4 trillion for the current flagship, so two to four times larger. Qwen 4 is already in training, with Qwen 4.5 and Qwen 5 the ones aiming at that scale. The second half is the hardware to run it. Alibaba unveiled the Zhenwu V900, an AI accelerator from its in-house chip unit T-Head, which the company says delivers three times the processing power of its predecessor, the M890, and can be wired into clusters of up to 500,000 units. Mass production is set for the first quarter of 2027. Wu also committed Alibaba Cloud to 20 gigawatts of data centre capacity by 2032, extending a previously announced three-year AI spend of 53 billion dollars. Shares rose in Hong Kong.

What is behind this

Parameters are the adjustable numbers inside a model, the things training tunes, and more of them generally means more capacity to store and combine patterns. The relationship to actual usefulness is much looser than the headline number suggests, and most labs have spent the past two years getting more out of models rather than simply building bigger ones. So a jump to 10 trillion is best read as a statement of intent about compute, not a promise about quality.

The chip is the more consequential announcement. US export controls have restricted which Nvidia parts Chinese companies can buy, and the response across the Chinese industry has been to build domestic alternatives. A 500,000-card cluster target is the number that matters there, because training at frontier scale is less about any single chip’s speed than about whether you can make hundreds of thousands of them work as one machine without the whole thing falling over. Whether the V900 actually delivers is unknown: the performance claim is Alibaba’s own, there is no independent benchmarking, and the part does not enter mass production for another two quarters. The 20 gigawatt figure is the one to keep an eye on, because electricity, not silicon, is becoming the real constraint on this industry.

What this means for you: Nothing you will notice this year. The reason to pay attention is that the Qwen family is where a lot of the free, downloadable models come from, so Alibaba’s compute plans eventually shape what anyone can run without paying a subscription. There is also a quieter point in here about how these announcements work. A model that does not exist, a chip that is not in production and a build-out that finishes in 2032 are all real plans, and none of them is a product. It is worth holding roadmap news and shipped news in different mental folders.

Sources

Source: https://www.reuters.com/business/retail-consumer/alibaba-plans-ai-model-with-5-trillion-10-trillion-parameters-unveils-new-chip-2026-09-22/

Next story

A 125 Billion Parameter Model on One Ordinary Graphics Card

Tim Dettmers previewed his lab's open-source week: an inference framework running Qwen 3.8 Flash Next on a single 24 GB GPU, DeepSeek V4.1 on a MacBook, and a compaction technique he says halves agent costs.

An enormous sphere compressed and deforming to fit inside a small desktop box