uttapen

Nemotron 3 Super (free) API: toman pricing and code

nvidia/nemotron-3-super-120b-a12b:free

Free · shared capacitytoolsreasoningjson
Input · per 1M tokens
0 toman
Output · per 1M tokens
0 toman
One 1,000-word request ≈
0 toman

You are billed for the usage the request actually reported. Prices follow the market.Pricing and top-ups

Nemotron 3 Super (free) example: tool calling

The example is picked from this model's own capabilities. Drop in your key and it runs as is.

from openai import OpenAI
import json

client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")

tools = [{
    "type": "function",
    "function": {
        "name": "check_stock",
        "description": "Returns the stock level of a product",
        "parameters": {
            "type": "object",
            "properties": {"sku": {"type": "string"}},
            "required": ["sku"],
        },
    },
}]

messages = [{"role": "user", "content": "How many of the Nike NK-42 shoe are in stock?"}]
first = client.chat.completions.create(model="nvidia/nemotron-3-super-120b-a12b:free", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]

# you run the function yourself — the model never touches the database
args = json.loads(call.function.arguments)
result = {"sku": args["sku"], "qty": 7}

messages += [first.choices[0].message, {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}]
final = client.chat.completions.create(model="nvidia/nemotron-3-super-120b-a12b:free", messages=messages, tools=tools)
print(final.choices[0].message.content)

What is Nemotron 3 Super (free) good for?

Nemotron 3 Super (free) comes from NVIDIA; in uttapen you reach it with the model id "nvidia/nemotron-3-super-120b-a12b:free". It accepts up to 262,144 tokens of input per call, so a long document or several source files fit in a single request. A single response can run to 235,929 tokens. Using it costs nothing. In exchange the capacity is shared, so during busy hours a request can come back empty-handed.

Capabilities available on this id: it supports tool calling, so it can invoke your own functions with valid arguments; it returns schema-valid JSON through response_format, ready to hand to your code; it has a reasoning mode that pays off on multi-step problems, maths and debugging. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper. Keep in mind that reasoning tokens are output tokens and do appear on the bill.

For a back-of-envelope figure: about 0 toman per 1,000-word exchange (0 in, 0 out, per 1M tokens). You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.

The closest alternative with the same capabilities from a different provider is North Mini Code (free), and its context window is larger. Both run on the same key and the same code, so trying the other one is a single string change.

This id gets confused with "nvidia/nemotron-3-super-120b-a12b", because the underlying model is the same. The difference is the suffix: "free" against the standard variant. The context windows differ too: 262,144 against 1,000,000 tokens. Maximum answer length differs as well: 235,929 against 16,384 tokens. On parameters the other takes frequency_penalty, logit_bias, logprobs, min_p.

NVIDIA has 10 models in our catalogue; the cheapest is Nemotron 3 Nano 30B A3B at 94 toman per thousand words and the dearest Nemotron 3 Ultra at 1,414. Among the less common parameters it accepts reasoning_effort — all through the standard request body. By context size the nearest option from another provider is Trinity Large Thinking at 262,144 tokens.

What to use it for: proving an idea works before you spend anything, logic puzzles and code review, agents that reach out to APIs and databases, extracting data against a fixed schema, analysing a long document or codebase in one request.

Provider's own description

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Three real jobs, priced on this model

Each figure is derived from the prices above and moves when they do.

JobTokensCost
One chat turn with a medium history1,500 in + 400 out0 toman
Summarising a ten-page document4,000 in + 600 out0 toman
Classifying a thousand short rows120,000 in + 20,000 out0 toman

Other variants of this model

Same core model, different execution terms and different price. This page is the free variant.

VariantModel idOutput / 1MContext
standardnvidia/nemotron-3-super-120b-a12b116,0501,000,000

Put the variant's id verbatim in the model field; nothing else in your code changes.

Frequently asked

How do I call Nemotron 3 Super (free) from Iran?
Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "nvidia/nemotron-3-super-120b-a12b:free". Nothing else in your code changes.
What does Nemotron 3 Super (free) cost in toman?
0 toman per 1M input tokens and 0 toman per 1M output tokens; a 1,000-word request is around 0 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
How much input does Nemotron 3 Super (free) take?
Up to 262,144 tokens per request, roughly 197k words. A single answer can reach 235,929 tokens.
Does Nemotron 3 Super (free) support streaming and tool calling?
Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape. Structured output through response_format works too.
What are the limits on a free model?
Free models have a daily per-user cap and their capacity is shared with everyone else. When the shared pool is exhausted the request is refused with a clear error and your wallet is untouched.