Qwen 2 VL 72B Instruct
Qwen 2 VL 72B Instruct is Alibaba's 2024 multimodal model, pairing a 72B dense language model with integrated visual encoding for images, charts, diagrams, and extended video. It can analyze video segments longer than 20 minutes and answer questions about their content. It is built for visual workflows including document parsing, form extraction, and reasoning over text embedded in images and graphics, with support for spatial coordinate and bounding-box output in JSON. It fits applications that need visual document understanding or video-based analysis.
Key info
Input
Output
Features
Context window
33K
Max output
120K
Available routes
No routes currently available — Qwen 2 VL 72B Instruct isn't routed through the Opper gateway right now. It may return.
Contact us about this model →