GLM 4.6V
Z.AI
z-ai/glm-4.6vGLM-4.6V is Z.ai's vision-language model that treats images, video, and tools as first-class agent inputs, extending training context to 128K tokens with native multimodal function calling. It is designed for agentic multimodal workflows such as screenshot- and document-driven tool use.
Cost rate
0.2x
Context
131K
Released
Dec 8, 2025
Input
ImageTextVideo
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
Image understanding
#89 · top 61%
Mathematical
#90 · top 24%
Medicine & Healthcare
#95 · top 26%
Writing, Literature, & Language
#102 · top 26%
Performance
Median latency and throughput measured across recent requests.
Throughput73 tok/s
Latency