Provider

DeepSeek

Every DeepSeek model UltraGPT can reach, kept in sync with the live catalog — pricing, context window and capabilities for each one.

21 models from DeepSeek Up to 1.31M tokens of context
Open the appBrowse every provider

21

Models tracked

1

Providers

19

Reasoning models

5

See images (vision)

1,311K

Largest context window

0

Free to run

Directory

All DeepSeek models.

Search and filter this provider's lineup — pricing per million tokens, context window, modalities, reasoning effort and independent benchmark scores where available.

21 of 21 models

This model always redirects to the latest model in the DeepSeek Flash family.

ReasoningTool callingVisionStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedSep 14, 2026
Knowledge cutoff
Text + ImageText
Input$0.15 / 1M tok
Output$0.6 / 1M tok
Cache read$0.003 / 1M tok
Effort: max, high, low

This model always redirects to the latest model in the DeepSeek Pro family.

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output393.2K tok
ReleasedSep 14, 2026
Knowledge cutoff
TextText
Input$0.581 / 1M tok
Output$1.74 / 1M tok
Cache read$0.058 / 1M tok
Effort: max, high, low

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...

Tool callingStreamingJSON mode
Context window163.8K tok
Max output16K tok
ReleasedDec 26, 2024
Knowledge cutoffJul 31, 2024
TextText
Input$0.257 / 1M tok
Output$1.03 / 1M tok

Design Arena — ranked #79 across 7 categories

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

Tool callingStreamingJSON mode
Context window163.8K tok
Max output147.5K tok
ReleasedMar 24, 2025
Knowledge cutoffJul 31, 2024
TextText
Input$0.25 / 1M tok
Output$1 / 1M tok
Intelligence9.7
Coding21.2
Agentic0.8

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

ReasoningTool callingStreamingJSON mode
Context window163.8K tok
Max output32.8K tok
ReleasedAug 21, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.25 / 1M tok
Output$0.95 / 1M tok
Cache read$0.13 / 1M tok

Design Arena — ranked #84 across 7 categories

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

ReasoningTool callingStreamingJSON mode
Context window163.8K tok
Max output32.8K tok
ReleasedSep 22, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.27 / 1M tok
Output$1 / 1M tok
Cache read$0.135 / 1M tok
Intelligence15.4
Coding43.5
Agentic8.9

Design Arena — ranked #52 across 7 categories

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

ReasoningTool callingStreamingJSON mode
Context window163.8K tok
Max output65.5K tok
ReleasedDec 1, 2025
Knowledge cutoff
TextText
Input$0.269 / 1M tok
Output$0.4 / 1M tok
Cache read$0.134 / 1M tok
Coding44.2

Design Arena — ranked #59 across 8 categories

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

ReasoningTool callingStreamingJSON mode
Context window163.8K tok
Max output65.5K tok
ReleasedSep 29, 2025
Knowledge cutoffJul 31, 2025
TextText
Input$0.27 / 1M tok
Output$0.41 / 1M tok

Design Arena — ranked #56 across 7 categories

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

ReasoningTool callingVisionStreamingJSON mode
Context window1.05M tok
Max output384K tok
ReleasedSep 10, 2026
Knowledge cutoffMay 1, 2025
TextText
Input$0.084 / 1M tok
Output$0.168 / 1M tok
Cache read$0.017 / 1M tok
Effort: xhigh, high
Intelligence24.8
Coding52.0
Agentic27.9

Design Arena — ranked #36 across 8 categories

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

ReasoningTool callingStreamingJSON mode
Context window1.31M tok
Max output943.7K tok
ReleasedJul 31, 2026
Knowledge cutoff
TextText
Input$0.06 / 1M tok
Output$0.12 / 1M tok
Cache read$0.012 / 1M tok
Effort: max, high, low
Intelligence34.5
Coding69.1
Agentic41.7

Design Arena — ranked #24 across 8 categories

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedJul 31, 2026
Knowledge cutoff
TextText
Input$0.11 / 1M tok
Output$0.33 / 1M tok
Cache read$0.0035 / 1M tok
Effort: max, high, low
Intelligence34.5
Coding69.1
Agentic41.7

Design Arena — ranked #24 across 8 categories

This model always redirects to the latest model in the DeepSeek V4 Flash family.

ReasoningTool callingStreamingJSON mode
Context window1.31M tok
Max output393.2K tok
ReleasedAug 1, 2026
Knowledge cutoff
TextText
Input$0.04 / 1M tok
Output$0.1 / 1M tok
Cache read$0.01 / 1M tok
Effort: max, high, low

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

ReasoningTool callingVisionStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedSep 10, 2026
Knowledge cutoffMay 1, 2025
Text + ImageText
Input$0.22 / 1M tok
Output$0.66 / 1M tok
Cache read$0.007 / 1M tok
Effort: max, high, low

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

ReasoningTool callingVisionStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedAug 21, 2026
Knowledge cutoff
Text + ImageText
Input$0.11 / 1M tok
Output$0.33 / 1M tok
Cache read$0.0035 / 1M tok
Effort: max, high, low

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output393.2K tok
ReleasedAug 12, 2026
Knowledge cutoff
TextText
Input$1.6 / 1M tok
Output$3.2 / 1M tok
Cache read$0.135 / 1M tok
Effort: xhigh, high
Intelligence30.9
Coding59.4
Agentic27.7

Design Arena — ranked #22 across 11 categories

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output393.2K tok
ReleasedAug 12, 2026
Knowledge cutoff
TextText
Input$0.659 / 1M tok
Output$1.98 / 1M tok
Cache read$0.021 / 1M tok
Effort: max, high, low
Intelligence36.3
Coding68.8
Agentic42.3

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

ReasoningTool callingStreamingJSON mode
Context window1.05M tok
Max output943.7K tok
ReleasedAug 12, 2026
Knowledge cutoff
TextText
Input$0.66 / 1M tok
Output$1.98 / 1M tok
Cache read$0.022 / 1M tok
Effort: max, high, low
Intelligence36.3
Coding68.8
Agentic42.3

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

ReasoningTool callingVisionStreamingJSON mode
Context window1.05M tok
Max output384K tok
ReleasedSep 10, 2026
Knowledge cutoff
Text + ImageText
Input$0.15 / 1M tok
Output$0.6 / 1M tok
Cache read$0.003 / 1M tok
Effort: max, high, low
Intelligence39.5

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

ReasoningTool callingStreamingJSON mode
Context window64K tok
Max output16K tok
ReleasedJan 20, 2025
Knowledge cutoffJul 31, 2024
TextText
Input$0.7 / 1M tok
Output$2.5 / 1M tok
Intelligence11.4
Coding24.6
Agentic1.1

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

ReasoningTool callingStreamingJSON mode
Context window163.8K tok
Max output32.8K tok
ReleasedMay 28, 2025
Knowledge cutoffMar 31, 2025
TextText
Input$0.5 / 1M tok
Output$2.15 / 1M tok
Cache read$0.35 / 1M tok

Design Arena — ranked #53 across 7 categories

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

ReasoningStreaming
Context window8.2K tok
Max output7.4K tok
ReleasedJan 23, 2025
Knowledge cutoffJul 31, 2024
TextText
Input$0.8 / 1M tok
Output$0.8 / 1M tok

Every DeepSeek model above, one subscription away.

UltraGPT puts this whole directory behind a single login and one flat monthly price — no separate bill per provider.