Provider

NVIDIA

Every NVIDIA model UltraGPT can reach, kept in sync with the live catalog — pricing, context window and capabilities for each one.

10 models from NVIDIA Up to 1M tokens of context 5 free to run
Open the appBrowse every provider

10

Models tracked

1

Providers

10

Reasoning models

3

See images (vision)

1,000K

Largest context window

5

Free to run

Directory

All NVIDIA models.

Search and filter this provider's lineup — pricing per million tokens, context window, modalities, reasoning effort and independent benchmark scores where available.

10 of 10 models

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedDec 14, 2025
Knowledge cutoff
TextText
Input$0.05 / 1M tok
Output$0.2 / 1M tok
Cache read$0.03 / 1M tok
Intelligence8.9
Coding14.4
Agentic1.0

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

ReasoningTool callingVisionAudioStreaming
Context window256K tok
Max output65.5K tok
ReleasedApr 28, 2026
Knowledge cutoff
Text + Audio + Image + VideoText
InputFree
OutputFree
Coding13.8

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedMar 11, 2026
Knowledge cutoff
TextText
Input$0.08 / 1M tok
Output$0.45 / 1M tok
Effort: medium, low
Intelligence13.6
Coding37.7
Agentic4.1

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output235.9K tok
ReleasedMar 11, 2026
Knowledge cutoff
TextText
InputFree
OutputFree
Effort: medium, low
Intelligence13.6
Coding37.7
Agentic4.1

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output182.5K tok
ReleasedJun 4, 2026
Knowledge cutoff
TextText
Input$0.6 / 1M tok
Output$2.4 / 1M tok
Cache read$0.12 / 1M tok
Effort: high, medium
Intelligence23.4
Coding49.3
Agentic21.7

Design Arena — ranked #60 across 8 categories

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

ReasoningTool callingStreaming
Context window1M tok
Max output65.5K tok
ReleasedJun 4, 2026
Knowledge cutoff
TextText
InputFree
OutputFree
Effort: high, medium
Intelligence23.4
Coding49.3
Agentic21.7

Design Arena — ranked #60 across 8 categories

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

ReasoningVisionStreamingJSON mode
Context window131.1K tok
Max output118.0K tok
ReleasedJun 4, 2026
Knowledge cutoff
Text + ImageText
Input$0.2 / 1M tok
Output$0.2 / 1M tok

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

ReasoningVisionStreaming
Context window128K tok
Max output8.2K tok
ReleasedJun 4, 2026
Knowledge cutoff
Text + ImageText
InputFree
OutputFree

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

ReasoningTool callingStreamingJSON mode
Context window262.1K tok
Max output131.1K tok
ReleasedAug 11, 2026
Knowledge cutoff
TextText
Input$0.08 / 1M tok
Output$0.2 / 1M tok
Cache read$0.04 / 1M tok
Intelligence13.6
Coding26.8
Agentic6.1

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

ReasoningTool callingStreaming
Context window1M tok
Max output65.5K tok
ReleasedAug 11, 2026
Knowledge cutoff
TextText
InputFree
OutputFree
Intelligence13.6
Coding26.8
Agentic6.1

Every NVIDIA model above, one subscription away.

UltraGPT puts this whole directory behind a single login and one flat monthly price — no separate bill per provider.