Alibaba's new model beats its predecessor at one-ninth the training cost
Qwen3.8-Flash-Next has 125 billion parameters but activates only 6 billion per token. It outperforms a model three times its active size.
Qwen3.8-Flash-Next has 125 billion parameters but activates only 6 billion per token. It outperforms a model three times its active size.
Alibaba says it beats the larger previous model at roughly one-ninth the training cost. The release is framed as an architecture preview of Qwen4.
Apple is reported to have developed a large language model specific to the Chinese market together with Alibaba, and has registered the service with China's regulator.
Alibaba's Qwen team has published the 27-billion-parameter multimodal Qwen3.8-27B under the Apache 2.0 licence. The model handles a 262,000-token context natively.
Alibaba's open-weight Qwen model family surpassed 3 billion downloads globally in six months, outpacing Google and Meta. According to Hugging Face data, Qwen has become the most-derived model in the open-source ecosystem.