Nous: Hermes 4 405B
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...
Intelligence overview
131.1K tok
Context window
118.0K tok
Max output
What Nous: Hermes 4 405B can do.
Reasoning
Extended reasoning
Works through problems step by step before answering, trading latency for accuracy on harder tasks.
Development
Structured output
Returns responses constrained to a JSON schema, ready to parse without cleanup.
Streaming
Streams tokens as they're generated instead of waiting on the full response.
Pay for what you use.
Every model in UltraGPT runs behind one flat subscription — this is the underlying per-token cost the catalog tracks.
$1 / 1M tok
$3 / 1M tok
Built for serious work.
Build
Generate and refactor code, with the ability to return structured, parseable output.
Think
Works through multi-step problems with extended reasoning before it answers.
Research
Holds up to 131.1K tokens of source material in context without losing the thread.
Ask Nous: Hermes 4 405B anything.
A preview of the real UltraGPT composer — the model you see here is the model you get.
Continues in the UltraGPT app — one login, every model.
How Nous: Hermes 4 405B stacks up.
The closest models in the catalog by provider, capability and score — real entries, pulled from the same directory.
Technical details.
Specification
- Context window
- 131.1K tokens
- Max output
- 118.0K tokens
- Input modalities
- Text
- Output modalities
- Text
- Knowledge cutoff
- Aug 31, 2024
- Release date
- Aug 26, 2025
- Streaming
- Supported
- Tool calling
- Not supported
- JSON mode
- Supported
More from Nous Research, and everything else.
This model is one entry in the full UltraGPT catalog — browse the rest, compare pricing and context windows, or filter by capability.
