Inception: Mercury 2
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...
Intelligence overview
128K tok
Context window
50K tok
Max output
What Inception: Mercury 2 can do.
Reasoning
Extended reasoning
Works through problems step by step before answering, trading latency for accuracy on harder tasks.
Adjustable effort
Choose how much the model deliberates per request — trade speed for depth as the task demands.
Development
Tool calling
Calls external functions mid-response, then continues reasoning from the result.
Structured output
Returns responses constrained to a JSON schema, ready to parse without cleanup.
Streaming
Streams tokens as they're generated instead of waiting on the full response.
Performance
Design Arena — ranked #64 across 8 categories
Pay for what you use.
Every model in UltraGPT runs behind one flat subscription — this is the underlying per-token cost the catalog tracks.
$0.25 / 1M tok
$0.75 / 1M tok
Built for serious work.
Build
Generate and refactor code, with the ability to run and verify tool calls and return structured, parseable output.
Think
Works through hard problems with adjustable reasoning effort — high, medium, low, none.
Research
Holds up to 128K tokens of source material in context without losing the thread.
Automate
Chains tool calls across multi-step workflows, reasoning between each one.
Ask Inception: Mercury 2 anything.
A preview of the real UltraGPT composer — the model you see here is the model you get.
Continues in the UltraGPT app — one login, every model.
How Inception: Mercury 2 stacks up.
The closest models in the catalog by provider, capability and score — real entries, pulled from the same directory.
Technical details.
Specification
- Context window
- 128K tokens
- Max output
- 50K tokens
- Input modalities
- Text
- Output modalities
- Text
- Knowledge cutoff
- Jan 1, 2025
- Release date
- Feb 24, 2026
- Reasoning effort
- high, medium, low, none
- Streaming
- Supported
- Tool calling
- Supported
- JSON mode
- Supported
Reasoning
More from Inception, and everything else.
This model is one entry in the full UltraGPT catalog — browse the rest, compare pricing and context windows, or filter by capability.
