Qwen 2 VL 72B Instruct

by Alibaba

Qwen 2 VL 72B Instruct is Alibaba's 2024 multimodal model, pairing a 72B dense language model with integrated visual encoding for images, charts, diagrams, and extended video. It can analyze video segments longer than 20 minutes and answer questions about their content. It is built for visual workflows including document parsing, form extraction, and reasoning over text embedded in images and graphics, with support for spatial coordinate and bounding-box output in JSON. It fits applications that need visual document understanding or video-based analysis.

Key info

Input
Output
Features
Context window
33K
Max output
120K

Available routes

No routes currently available — Qwen 2 VL 72B Instruct isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Alibaba

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
Qwen 2 VL 72B Instruct by Alibaba — not currently on Opper | Opper AI