Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Rank
#139
Up 2% week over week
Tokens
11.3B
2026-08-14
Requests
13.5M
13,495,103
Tool calls
-
Input / M
$0.12
Output / M
$0.45
Cache / M
Free
Context
262K
262,144 tokens