GLM 4.5V
Z.AI
z-ai/glm-4.5vGLM-4.5V is Z.ai's vision-language model built on the GLM-4.5-Air architecture (106B total, 12B active), designed for versatile multimodal reasoning. It handles complex STEM problem-solving, GUI agent tasks, and video understanding via a 3D convolutional vision encoder.
Cost rate
0.4x
Context
66K
Released
Aug 11, 2025
Input
TextImage
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
Image understanding
#97 · top 66%
Medicine & Healthcare
#177 · top 49%
Entertainment, Sports, & Media
#190 · top 49%
Software & IT Services
#191 · top 49%
Performance
Median latency and throughput measured across recent requests.
Throughput92 tok/s
Latency