The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios.
Try NowDense vision model for document understanding
Visual detection or spatial reasoning
Self-hosted multimodal with strong perception
131,072 tokens
32,768 tokens
$0.16 per 1M tokens
$0.64 per 1M tokens
$15 per 1K calls
$0.19 per 1K calls