Provider

Z.AI

Every Z.AI model UltraGPT can reach, kept in sync with the live catalog — pricing, context window and capabilities for each one.

19 models from Z.AI Up to 1.31M tokens of context
Open the appBrowse every provider

19

Models tracked

1

Providers

19

Reasoning models

6

See images (vision)

1,311K

Largest context window

0

Free to run

Directory

All Z.AI models.

Search and filter this provider's lineup — pricing per million tokens, context window, modalities, reasoning effort and independent benchmark scores where available.

19 of 19 models

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output98.3K tok
ReleasedJul 25, 2025
Knowledge cutoffDec 31, 2024
TextText
Input$0.6 / 1M tok
Output$2.2 / 1M tok
Cache read$0.11 / 1M tok

Design Arena — ranked #48 across 7 categories

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

ReasoningTool callingStreaming
Context window131.1K tok
Max output98.3K tok
ReleasedJul 25, 2025
Knowledge cutoffDec 31, 2024
TextText
Input$0.13 / 1M tok
Output$0.85 / 1M tok
Cache read$0.025 / 1M tok

Design Arena — ranked #51 across 7 categories

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

ReasoningTool callingVisionStreamingJSON mode
Context window65.5K tok
Max output16.4K tok
ReleasedAug 11, 2025
Knowledge cutoffDec 31, 2024
Text + ImageText
Input$0.6 / 1M tok
Output$1.8 / 1M tok
Cache read$0.11 / 1M tok

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

ReasoningTool callingStreamingJSON mode
Context window204.8K tok
Max output16.4K tok
ReleasedSep 30, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.43 / 1M tok
Output$1.75 / 1M tok
Cache read$0.08 / 1M tok
Coding45.8

Design Arena — ranked #12 across 11 categories

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

ReasoningTool callingVisionStreamingJSON mode
Context window131.1K tok
Max output32.8K tok
ReleasedDec 8, 2025
Knowledge cutoff
Image + Text + VideoText
Input$0.3 / 1M tok
Output$0.9 / 1M tok
Cache read$0.055 / 1M tok

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

ReasoningTool callingStreamingJSON mode
Context window204.8K tok
Max output131.1K tok
ReleasedDec 22, 2025
Knowledge cutoff
TextText
Input$0.4 / 1M tok
Output$1.75 / 1M tok
Cache read$0.08 / 1M tok
Coding45.3

Design Arena — ranked #27 across 12 categories

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

ReasoningTool callingStreamingJSON mode
Context window200K tok
Max output118.0K tok
ReleasedJan 19, 2026
Knowledge cutoff
TextText
Input$0.061 / 1M tok
Output$0.4 / 1M tok

Design Arena — ranked #45 across 7 categories

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

ReasoningTool callingStreamingJSON mode
Context window204.8K tok
Max output128K tok
ReleasedFeb 11, 2026
Knowledge cutoff
TextText
Input$0.6 / 1M tok
Output$1.92 / 1M tok
Cache read$0.12 / 1M tok

Design Arena — ranked #15 across 13 categories

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

ReasoningTool callingStreamingJSON mode
Context window202.8K tok
Max output131.1K tok
ReleasedMar 15, 2026
Knowledge cutoff
TextText
Input$1.2 / 1M tok
Output$4 / 1M tok
Cache read$0.24 / 1M tok

Design Arena — ranked #18 across 8 categories

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

ReasoningTool callingStreamingJSON mode
Context window204.8K tok
Max output128K tok
ReleasedApr 7, 2026
Knowledge cutoff
TextText
Input$0.966 / 1M tok
Output$3.04 / 1M tok
Cache read$0.179 / 1M tok
Intelligence26.4
Coding55.8
Agentic25.2

Design Arena — ranked #1 across 21 categories

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output131.1K tok
ReleasedJun 16, 2026
Knowledge cutoff
TextText
Input$0.683 / 1M tok
Output$2.15 / 1M tok
Cache read$0.127 / 1M tok
Effort: xhigh, high
Coding68.8
Agentic39.4

Design Arena — ranked #10 across 16 categories

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedJun 16, 2026
Knowledge cutoff
TextText
Input$0.7 / 1M tok
Output$2.2 / 1M tok
Cache read$0.07 / 1M tok
Effort: xhigh, high
Coding68.8
Agentic39.4

Design Arena — ranked #10 across 16 categories

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

ReasoningTool callingStreamingJSON mode
Context window1.31M tok
Max output943.7K tok
ReleasedAug 18, 2026
Knowledge cutoff
TextText
Input$1.4 / 1M tok
Output$4.4 / 1M tok
Cache read$0.26 / 1M tok
Effort: max, high, low · always on
Intelligence44.9
Coding74.8
Agentic53.4

Design Arena — ranked #3 across 10 categories

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedAug 18, 2026
Knowledge cutoff
TextText
Input$0.7 / 1M tok
Output$2.2 / 1M tok
Cache read$0.13 / 1M tok
Effort: max, high, low · always on
Intelligence44.9
Coding74.8
Agentic53.4

Design Arena — ranked #3 across 10 categories

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

ReasoningTool callingVisionStreamingJSON mode
Context window1.31M tok
Max output131.1K tok
ReleasedAug 26, 2026
Knowledge cutoff
Text + Image + VideoText
Input$0.15 / 1M tok
Output$0.5 / 1M tok
Cache read$0.03 / 1M tok
Effort: max, high, low · always on
Intelligence41.9
Coding71.5
Agentic51.2

Design Arena — ranked #5 across 8 categories

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

ReasoningTool callingVisionStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedAug 26, 2026
Knowledge cutoff
Text + Image + VideoText
Input$0.075 / 1M tok
Output$0.25 / 1M tok
Cache read$0.015 / 1M tok
Effort: max, high, low · always on
Intelligence41.9
Coding71.5
Agentic51.2

Design Arena — ranked #5 across 8 categories

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

ReasoningTool callingVisionStreamingJSON mode
Context window202.8K tok
Max output131.1K tok
ReleasedApr 1, 2026
Knowledge cutoff
Image + Text + VideoText
Input$1.2 / 1M tok
Output$4 / 1M tok
Cache read$0.24 / 1M tok

Design Arena — ranked #3 across 21 categories

This model always redirects to the latest model in the GLM Flash family.

ReasoningTool callingVisionStreamingJSON mode
Context window1.31M tok
Max output131.1K tok
ReleasedAug 27, 2026
Knowledge cutoff
Text + Image + VideoText
Input$0.075 / 1M tok
Output$0.25 / 1M tok
Cache read$0.015 / 1M tok
Effort: max, high, low · always on

This model always redirects to the latest GLM model from Z.ai.

ReasoningTool callingStreamingJSON mode
Context window1.31M tok
Max output131.1K tok
ReleasedAug 19, 2026
Knowledge cutoff
TextText
Input$0.91 / 1M tok
Output$2.86 / 1M tok
Cache read$0.169 / 1M tok
Effort: max, high, low · always on

Every Z.AI model above, one subscription away.

UltraGPT puts this whole directory behind a single login and one flat monthly price — no separate bill per provider.