Alibaba's Qwen AI team launched Qwen3.8-Flash-Next. This new cost-efficient, open-weight model offers developers an early preview of the Qwen4 series architecture.
The model uses a Mixture-of-Experts (MoE) architecture. This design enhances efficiency by activating only a fraction of its total parameters per task.
Qwen3.8-Flash-Next reportedly features 125 billion parameters. Only 6 billion parameters are active per token. Alibaba aims for "ultimate cost-efficiency" with this design.
The company claims the new model significantly reduces training and inference costs. It requires about one-ninth of the training resources compared to predecessors. The model also delivers superior performance in coding and office-related tasks.
Developers can access the model's weights on platforms such as Hugging Face. This allows them to run and modify the model on their own servers.