Nemotron 3.5 Lightning API: toman pricing and code
nvidia/nemotron-3.5-lightning
toolsreasoningjsonYou are billed for the usage the request actually reported. Prices follow the market. Cached input: 11,605 toman / 1M.Pricing and top-ups
Nemotron 3.5 Lightning example: tool calling
The example is picked from this model's own capabilities. Drop in your key and it runs as is.
from openai import OpenAI
import json
client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")
tools = [{
"type": "function",
"function": {
"name": "check_stock",
"description": "Returns the stock level of a product",
"parameters": {
"type": "object",
"properties": {"sku": {"type": "string"}},
"required": ["sku"],
},
},
}]
messages = [{"role": "user", "content": "How many of the Nike NK-42 shoe are in stock?"}]
first = client.chat.completions.create(model="nvidia/nemotron-3.5-lightning", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
# you run the function yourself — the model never touches the database
args = json.loads(call.function.arguments)
result = {"sku": args["sku"], "qty": 7}
messages += [first.choices[0].message, {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}]
final = client.chat.completions.create(model="nvidia/nemotron-3.5-lightning", messages=messages, tools=tools)
print(final.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uttapen.ir/v1", apiKey: "sk-up-…" });
const tools = [{
type: "function",
function: {
name: "check_stock",
description: "Returns the stock level of a product",
parameters: { type: "object", properties: { sku: { type: "string" } }, required: ["sku"] },
},
}];
const messages = [{ role: "user", content: "How many of the Nike NK-42 shoe are in stock?" }];
const first = await client.chat.completions.create({ model: "nvidia/nemotron-3.5-lightning", messages, tools });
const call = first.choices[0].message.tool_calls[0];
const args = JSON.parse(call.function.arguments);
const result = { sku: args.sku, qty: 7 };
messages.push(first.choices[0].message, { role: "tool", tool_call_id: call.id, content: JSON.stringify(result) });
const final = await client.chat.completions.create({ model: "nvidia/nemotron-3.5-lightning", messages, tools });
console.log(final.choices[0].message.content);curl https://api.uttapen.ir/v1/chat/completions \
-H "Authorization: Bearer sk-up-…" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-3.5-lightning",
"messages": [{"role": "user", "content": "How many of the Nike NK-42 shoe are in stock?"}],
"tools": [{
"type": "function",
"function": {
"name": "check_stock",
"parameters": {"type": "object", "properties": {"sku": {"type": "string"}}, "required": ["sku"]}
}
}]
}'What is Nemotron 3.5 Lightning good for?
NVIDIA publishes this model; we expose it under the id "nvidia/nemotron-3.5-lightning". Its context window is 262,144 tokens, roughly 197k English words in one request. A single response can run to 131,072 tokens. On price it sits in the "very cheap" band — cheaper than 33 and dearer than 373 of the other paid models in the catalogue.
Capabilities available on this id: it supports tool calling, so it can invoke your own functions with valid arguments; it returns schema-valid JSON through response_format, ready to hand to your code; it has a reasoning mode that pays off on multi-step problems, maths and debugging. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper. Keep in mind that reasoning tokens are output tokens and do appear on the bill.
For a back-of-envelope figure: about 106 toman per 1,000-word exchange (23,210 in, 58,025 out, per 1M tokens). It supports cached input: repeated context is billed at 11,605 toman per 1M, which matters a lot if your system prompt is long. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.
The closest alternative with the same capabilities from a different provider is Qwen2.5 7B Instruct: Nemotron 3.5 Lightning works out roughly 1.1× cheaper, and its context window is larger. Both run on the same key and the same code, so trying the other one is a single string change.
To make the figure concrete: 100,000 toman of credit buys roughly 943 thousand-word requests on Nemotron 3.5 Lightning, and every 1,000 toman is about 9,434 words of round trip. A job with one million input tokens and one million output tokens comes to 81,235 toman in total. Filling this model's 262,144-token window costs 6,084 toman on the input side alone, which is the real reason to keep conversation history short.
This id gets confused with "nvidia/nemotron-3.5-lightning:free", because the underlying model is the same. The difference is the suffix: no suffix (the standard variant) against "free". The context windows differ too: 262,144 against 1,000,000 tokens. Maximum answer length differs as well: 131,072 against 65,536 tokens. On parameters, this one takes frequency_penalty, logit_bias, logprobs, min_p.
NVIDIA has 10 models in our catalogue; the cheapest is Nemotron 3 Nano 30B A3B at 94 toman per thousand words and the dearest Nemotron 3 Ultra at 1,414. Among the less common parameters it accepts logit_bias, logprobs, min_p, repetition_penalty, top_k, top_logprobs — all through the standard request body. By context size the nearest option from another provider is Trinity Large Thinking at 262,144 tokens.
What to use it for: high-volume work such as classification, tagging and bulk summarising, logic puzzles and code review, agents that reach out to APIs and databases, extracting data against a fixed schema, analysing a long document or codebase in one request.
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Three real jobs, priced on this model
Each figure is derived from the prices above and moves when they do.
| Job | Tokens | Cost |
|---|---|---|
| One chat turn with a medium history | 1,500 in + 400 out | 58 toman |
| Summarising a ten-page document | 4,000 in + 600 out | 128 toman |
| Classifying a thousand short rows | 120,000 in + 20,000 out | 3,946 toman |
Other variants of this model
Same core model, different execution terms and different price. This page is the standard variant.
| Variant | Model id | Output / 1M | Context |
|---|---|---|---|
| free | nvidia/nemotron-3.5-lightning:free | 0 | 1,000,000 |
Put the variant's id verbatim in the model field; nothing else in your code changes.
Frequently asked
- How do I call Nemotron 3.5 Lightning from Iran?
- Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "nvidia/nemotron-3.5-lightning". Nothing else in your code changes.
- What does Nemotron 3.5 Lightning cost in toman?
- 23,210 toman per 1M input tokens and 58,025 toman per 1M output tokens; a 1,000-word request is around 106 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
- How much input does Nemotron 3.5 Lightning take?
- Up to 262,144 tokens per request, roughly 197k words. A single answer can reach 131,072 tokens.
- Does Nemotron 3.5 Lightning support streaming and tool calling?
- Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape. Structured output through response_format works too.