NVIDIA

NVIDIA: Switchyard

Switchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use OpenRouter market data to select the most popular...

Streaming
1M token context Released Sep 21, 2026

Intelligence overview

1M tok

Context window

— tok

Max output

Capabilities

What NVIDIA: Switchyard can do.

Development

Streaming

Streams tokens as they're generated instead of waiting on the full response.

Pricing

Pay for what you use.

Every model in UltraGPT runs behind one flat subscription — this is the underlying per-token cost the catalog tracks.

INPUT

$-1000000 / 1M tok

OUTPUT

$-1000000 / 1M tok

Use cases

Built for serious work.

Research

Holds up to 1M tokens of source material in context without losing the thread.

Preview

Ask NVIDIA: Switchyard anything.

A preview of the real UltraGPT composer — the model you see here is the model you get.

NVIDIA: Switchyard Ready

Continues in the UltraGPT app — one login, every model.

Comparison

How NVIDIA: Switchyard stacks up.

The closest models in the catalog by provider, capability and score — real entries, pulled from the same directory.

Specification

Technical details.

Specification

Context window
1M tokens
Max output
— tokens
Input modalities
Text
Output modalities
Text
Release date
Sep 21, 2026
Streaming
Supported
Tool calling
Not supported
JSON mode
Not supported

Model identity

Provider
NVIDIA
Provider ID
nvidia

Model ID

Directory

More from NVIDIA, and everything else.

This model is one entry in the full UltraGPT catalog — browse the rest, compare pricing and context windows, or filter by capability.

NVIDIA: Switchyard, one subscription away.

UltraGPT puts this model — and every other one in the directory — behind a single login and one flat monthly price.