Qwen2.5 VL 72B Instruct
Qwen
qwen/qwen2.5-vl-72b-instructQwen 2.5 VL 72B Instruct is Alibaba's large open-weight vision-language model, capable of image understanding, document/OCR parsing, and visual reasoning alongside text. It is designed for multimodal tasks that combine strong language ability with detailed visual comprehension.
Cost rate
0.3x
Context
128K
Released
Feb 1, 2025
Input
TextImage
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.