Inference.net: Schematron V2 Turbo
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...
Intelligence overview
128K tok
Context window
8.2K tok
Max output
What Inference.net: Schematron V2 Turbo can do.
Development
Structured output
Returns responses constrained to a JSON schema, ready to parse without cleanup.
Streaming
Streams tokens as they're generated instead of waiting on the full response.
Pay for what you use.
Every model in UltraGPT runs behind one flat subscription — this is the underlying per-token cost the catalog tracks.
$0.03 / 1M tok
$0.15 / 1M tok
Built for serious work.
Build
Generate and refactor code, with the ability to return structured, parseable output.
Research
Holds up to 128K tokens of source material in context without losing the thread.
Ask Inference.net: Schematron V2 Turbo anything.
A preview of the real UltraGPT composer — the model you see here is the model you get.
Continues in the UltraGPT app — one login, every model.
How Inference.net: Schematron V2 Turbo stacks up.
The closest models in the catalog by provider, capability and score — real entries, pulled from the same directory.
Technical details.
Specification
- Context window
- 128K tokens
- Max output
- 8.2K tokens
- Input modalities
- Text
- Output modalities
- Text
- Release date
- Sep 12, 2026
- Streaming
- Supported
- Tool calling
- Not supported
- JSON mode
- Supported
More from Inference Net, and everything else.
This model is one entry in the full UltraGPT catalog — browse the rest, compare pricing and context windows, or filter by capability.
