Alibaba's Qwen AI team launched Qwen3.8-Flash-Next. This new cost-efficient, open-weight model offers developers an early preview of the Qwen4 series architecture.

The model uses a Mixture-of-Experts (MoE) architecture. This design enhances efficiency by activating only a fraction of its total parameters per task.

Qwen3.8-Flash-Next reportedly features 125 billion parameters. Only 6 billion parameters are active per token. Alibaba aims for "ultimate cost-efficiency" with this design.

The company claims the new model significantly reduces training and inference costs. It requires about one-ninth of the training resources compared to predecessors. The model also delivers superior performance in coding and office-related tasks.

Developers can access the model's weights on platforms such as Hugging Face. This allows them to run and modify the model on their own servers.