Nemotron 3 Ultra (free) API: toman pricing and code
nvidia/nemotron-3-ultra-550b-a55b:free
Free · shared capacitytoolsreasoningYou are billed for the usage the request actually reported. Prices follow the market.Pricing and top-ups
Nemotron 3 Ultra (free) example: a multi-step problem with reasoning
The example is picked from this model's own capabilities. Drop in your key and it runs as is.
from openai import OpenAI
client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")
resp = client.chat.completions.create(
model="nvidia/nemotron-3-ultra-550b-a55b:free",
messages=[{"role": "user", "content": (
"A shop has 3 warehouses. A ships 120 orders a day, B ships 85 and C ships 40. "
"If C closes and its load is split between A and B in proportion to their capacity, how many does each ship a day? "
"Work through it step by step and give just the two numbers at the end."
)}],
reasoning_effort="medium", # low | medium | high
)
print(resp.choices[0].message.content)
# reasoning tokens count as output tokens too:
print(resp.usage.completion_tokens_details)curl https://api.uttapen.ir/v1/chat/completions \
-H "Authorization: Bearer sk-up-…" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-3-ultra-550b-a55b:free",
"messages": [{"role": "user", "content": "120 and 85 daily orders split across two warehouses; work through it step by step."}],
"reasoning": {"effort": "medium"}
}'import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uttapen.ir/v1", apiKey: "sk-up-…" });
const resp = await client.chat.completions.create({
model: "nvidia/nemotron-3-ultra-550b-a55b:free",
messages: [{ role: "user", content: "120 and 85 daily orders are split between two warehouses; work through it step by step and give just the two numbers at the end." }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);
console.log(resp.usage.completion_tokens_details);What is Nemotron 3 Ultra (free) good for?
NVIDIA publishes this model; we expose it under the id "nvidia/nemotron-3-ultra-550b-a55b:free". Its context window is 1,000,000 tokens, roughly 750k English words in one request. A single response can run to 65,536 tokens. It is free: nothing leaves your wallet. The capacity is shared across all users and there is a daily cap per account, so treat it as a testing lane rather than a production dependency.
Capabilities available on this id: it supports tool calling, so it can invoke your own functions with valid arguments; it has a reasoning mode that pays off on multi-step problems, maths and debugging. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper. Keep in mind that reasoning tokens are output tokens and do appear on the bill.
For a back-of-envelope figure: about 0 toman per 1,000-word exchange (0 in, 0 out, per 1M tokens). You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.
The closest alternative with the same capabilities from a different provider is North Mini Code (free), and its context window is larger. Both run on the same key and the same code, so trying the other one is a single string change.
This id gets confused with "nvidia/nemotron-3-ultra-550b-a55b", because the underlying model is the same. The difference is the suffix: "free" against the standard variant. The context windows differ too: 1,000,000 against 262,144 tokens. Maximum answer length differs as well: 65,536 against 32,768 tokens. On parameters the other takes frequency_penalty, logit_bias, min_p, presence_penalty.
NVIDIA has 10 models in our catalogue; the cheapest is Nemotron 3 Nano 30B A3B at 94 toman per thousand words and the dearest Nemotron 3 Ultra at 1,414. Among the less common parameters it accepts reasoning_effort — all through the standard request body. It does not support response_format, structured_outputs, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Nova 2 Lite at 1,000,000 tokens.
What to use it for: proving an idea works before you spend anything, logic puzzles and code review, agents that reach out to APIs and databases, analysing a long document or codebase in one request.
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Three real jobs, priced on this model
Each figure is derived from the prices above and moves when they do.
| Job | Tokens | Cost |
|---|---|---|
| One chat turn with a medium history | 1,500 in + 400 out | 0 toman |
| Summarising a ten-page document | 4,000 in + 600 out | 0 toman |
| Classifying a thousand short rows | 120,000 in + 20,000 out | 0 toman |
Other variants of this model
Same core model, different execution terms and different price. This page is the free variant.
| Variant | Model id | Output / 1M | Context |
|---|---|---|---|
| standard | nvidia/nemotron-3-ultra-550b-a55b | 906,641 | 262,144 |
Put the variant's id verbatim in the model field; nothing else in your code changes.
Frequently asked
- How do I call Nemotron 3 Ultra (free) from Iran?
- Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "nvidia/nemotron-3-ultra-550b-a55b:free". Nothing else in your code changes.
- What does Nemotron 3 Ultra (free) cost in toman?
- 0 toman per 1M input tokens and 0 toman per 1M output tokens; a 1,000-word request is around 0 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
- How much input does Nemotron 3 Ultra (free) take?
- Up to 1,000,000 tokens per request, roughly 750k words. A single answer can reach 65,536 tokens.
- Does Nemotron 3 Ultra (free) support streaming and tool calling?
- Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape.
- What are the limits on a free model?
- Free models have a daily per-user cap and their capacity is shared with everyone else. When the shared pool is exhausted the request is refused with a clear error and your wallet is untouched.