The Qwen3.8-Flash-Next model utilizes a multimodal Mixture-of-Experts architecture with 125B total tokens and 6B active parameters. This release serves as an early preview for the upcoming Qwen4. Early tests via Unsloth quantized versions show high reasoning capabilities. Practitioners can now benchmark this efficiency-focused architecture against existing open-weights models.