Provider

Qwen

Every Qwen model UltraGPT can reach, kept in sync with the live catalog — pricing, context window and capabilities for each one.

53 models from Qwen Up to 1.05M tokens of context
Open the appBrowse every provider

53

Models tracked

1

Providers

35

Reasoning models

26

See images (vision)

1,049K

Largest context window

0

Free to run

Directory

All Qwen models.

Search and filter this provider's lineup — pricing per million tokens, context window, modalities, reasoning effort and independent benchmark scores where available.

53 of 53 models

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Tool callingStreamingJSON mode
Context window32.8K tok
Max output16.4K tok
ReleasedSep 1, 2024
Knowledge cutoffJun 30, 2024
TextText
Input$0.36 / 1M tok
Output$0.4 / 1M tok

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**...

Streaming
Context window32.8K tok
Max output29.5K tok
ReleasedNov 11, 2024
Knowledge cutoffJun 30, 2024
TextText
Input$0.66 / 1M tok
Output$1 / 1M tok

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

Tool callingStreamingJSON mode
Context window1M tok
Max output32.8K tok
ReleasedSep 8, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.26 / 1M tok
Output$0.78 / 1M tok

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

ReasoningTool callingStreamingJSON mode
Context window1M tok
Max output32.8K tok
ReleasedJan 25, 2024
Knowledge cutoffMar 31, 2025
TextText
Input$0.26 / 1M tok
Output$0.78 / 1M tok
Cache read$0.052 / 1M tok
Cache write$0.325 / 1M tok

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Tool callingStreamingJSON mode
Context window32.8K tok
Max output29.5K tok
ReleasedSep 1, 2024
Knowledge cutoffJun 30, 2024
TextText
Input$0.1 / 1M tok
Output$0.2 / 1M tok

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Tool callingVisionStreamingJSON mode
Context window128K tok
Max output115.2K tok
ReleasedSep 1, 2024
Knowledge cutoffJun 30, 2024
Text + ImageText
Input$0.8 / 1M tok
Output$1 / 1M tok
Cache read$0.4 / 1M tok

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output16.4K tok
ReleasedApr 1, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.12 / 1M tok
Output$0.24 / 1M tok
Intelligence6.4
Coding13.8
Agentic0.9

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output8.2K tok
ReleasedApr 1, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.455 / 1M tok
Output$1.82 / 1M tok

Design Arena — ranked #105 across 6 categories

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedJul 21, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.087 / 1M tok
Output$0.35 / 1M tok
Cache read$0.018 / 1M tok

Design Arena — ranked #95 across 6 categories

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output118.0K tok
ReleasedJul 25, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.23 / 1M tok
Output$2.3 / 1M tok
Intelligence12.7
Coding22.1
Agentic1.3

Design Arena — ranked #100 across 6 categories

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output16.4K tok
ReleasedApr 28, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.12 / 1M tok
Output$0.5 / 1M tok

Design Arena — ranked #109 across 5 categories

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output32K tok
ReleasedJul 29, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.048 / 1M tok
Output$0.193 / 1M tok

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...

ReasoningTool callingStreamingJSON mode
Context window81.9K tok
Max output32.8K tok
ReleasedAug 28, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.2 / 1M tok
Output$2.4 / 1M tok
Intelligence9.8
Coding12.1
Agentic0.9

Design Arena — ranked #114 across 2 categories

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output16.4K tok
ReleasedApr 1, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.08 / 1M tok
Output$0.28 / 1M tok
Intelligence7.2
Coding15.3
Agentic0.9

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

ReasoningTool callingStreamingJSON mode
Context window131.1K tok
Max output8.2K tok
ReleasedApr 1, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.117 / 1M tok
Output$0.455 / 1M tok
Intelligence5.2
Coding9.0
Agentic0.8

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedApr 1, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.07 / 1M tok
Output$0.28 / 1M tok

Design Arena — ranked #93 across 3 categories

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output65.5K tok
ReleasedJul 23, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.3 / 1M tok
Output$1 / 1M tok
Cache read$0.1 / 1M tok

Design Arena — ranked #73 across 5 categories

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...

Tool callingStreamingJSON mode
Context window1M tok
Max output65.5K tok
ReleasedJul 28, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.195 / 1M tok
Output$0.975 / 1M tok
Cache read$0.039 / 1M tok
Cache write$0.244 / 1M tok

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedFeb 4, 2026
Knowledge cutoff
TextText
Input$0.12 / 1M tok
Output$0.8 / 1M tok
Cache read$0.07 / 1M tok
Intelligence10.1
Coding36.2
Agentic3.6

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

Tool callingStreamingJSON mode
Context window1M tok
Max output65.5K tok
ReleasedJul 23, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.65 / 1M tok
Output$3.25 / 1M tok
Cache read$0.13 / 1M tok
Cache write$0.813 / 1M tok

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output65.5K tok
ReleasedSep 23, 2025
Knowledge cutoffJun 30, 2025
TextText
Input$0.78 / 1M tok
Output$3.9 / 1M tok
Cache read$0.156 / 1M tok
Cache write$0.975 / 1M tok

Design Arena — ranked #48 across 8 categories

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output65.5K tok
ReleasedFeb 9, 2026
Knowledge cutoff
TextText
Input$0.78 / 1M tok
Output$3.9 / 1M tok

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

Tool callingStreamingJSON mode
Context window262.1K tok
Max output16.4K tok
ReleasedSep 1, 2025
Knowledge cutoffSep 30, 2025
TextText
Input$0.09 / 1M tok
Output$1.1 / 1M tok

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedSep 1, 2025
Knowledge cutoffSep 30, 2025
TextText
Input$0.15 / 1M tok
Output$1.2 / 1M tok
Coding17.4

Every Qwen model above, one subscription away.

UltraGPT puts this whole directory behind a single login and one flat monthly price — no separate bill per provider.