Ministral 3 3B 2512 API: toman pricing and code
mistralai/ministral-3b-2512
visiontoolsjsonYou are billed for the usage the request actually reported. Prices follow the market. Cached input: 2,901 toman / 1M.Pricing and top-ups
Ministral 3 3B 2512 example: reading an image into JSON
The example is picked from this model's own capabilities. Drop in your key and it runs as is.
# tip: a data URI works too — base64 the file and prefix it with data:image/jpeg;base64,
curl https://api.uttapen.ir/v1/chat/completions \
-H "Authorization: Bearer sk-up-…" \
-H "Content-Type: application/json" \
-d '{
"model": "mistralai/ministral-3b-2512",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Return only the invoice number and the total, as JSON."},
{"type": "image_url", "image_url": {"url": "https://example.com/factor.jpg"}}
]
}]
}'import base64, json
from openai import OpenAI
client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")
img = base64.b64encode(open("factor.jpg", "rb").read()).decode()
resp = client.chat.completions.create(
model="mistralai/ministral-3b-2512",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Return the invoice number, the date and the total as JSON."},
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img}"}},
],
}],
response_format={"type": "json_object"},
)
print(json.loads(resp.choices[0].message.content))import OpenAI from "openai";
import { readFileSync } from "node:fs";
const client = new OpenAI({ baseURL: "https://api.uttapen.ir/v1", apiKey: "sk-up-…" });
const img = readFileSync("factor.jpg").toString("base64");
const resp = await client.chat.completions.create({
model: "mistralai/ministral-3b-2512",
messages: [{
role: "user",
content: [
{ type: "text", text: "Return the invoice number, the date and the total as JSON." },
{ type: "image_url", image_url: { url: `data:image/jpeg;base64,${img}` } },
],
}],
response_format: { type: "json_object" },
});
console.log(JSON.parse(resp.choices[0].message.content));What is Ministral 3 3B 2512 good for?
The id for Ministral 3 3B 2512 in our API is "mistralai/ministral-3b-2512", served from Mistral. Context is 131,072 tokens; past that you have to summarise the history yourself. A single response can run to 104,857 tokens. On price it sits in the "very cheap" band — cheaper than 8 and dearer than 398 of the other paid models in the catalogue.
What it can do beyond plain text: it reads images directly, which makes it a real option for invoices, forms and screenshots; it supports tool calling, so it can invoke your own functions with valid arguments; it returns schema-valid JSON through response_format, ready to hand to your code. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper.
Pricing is 29,013 toman per 1M input tokens and 29,013 per 1M output tokens. A 1,000-word round trip on Ministral 3 3B 2512 lands near 75 toman. It supports cached input: repeated context is billed at 2,901 toman per 1M, which matters a lot if your system prompt is long. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.
The closest alternative with the same capabilities from a different provider is Nex-N2-Mini: Ministral 3 3B 2512 works out roughly 1.6× more expensive, and its context window is smaller. Both run on the same key and the same code, so trying the other one is a single string change.
To make the figure concrete: 100,000 toman of credit buys roughly 1,333 thousand-word requests on Ministral 3 3B 2512, and every 1,000 toman is about 13,333 words of round trip. A job with one million input tokens and one million output tokens comes to 58,025 toman in total. Filling this model's 131,072-token window costs 3,803 toman on the input side alone, which is the real reason to keep conversation history short.
Its nearest relative in the catalogue is "mistralai/mistral-small-3.2-24b-instruct", and choosing between those two is where most people hesitate. On a thousand-word request this variant works out 1.4× cheaper (75 against 104 toman). Maximum answer length differs as well: 104,857 against 16,384 tokens. On parameters the other takes logit_bias, logprobs, min_p, repetition_penalty.
Mistral has 20 models in our catalogue; the cheapest is Mistral Nemo at 18 toman per thousand words and the dearest Mistral Medium 3.5 at 3,394. It does not support include_reasoning, reasoning, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Aion-2.0 at 131,072 tokens.
Good fits: high-volume work such as classification, tagging and bulk summarising, pulling text and fields out of images, agents that reach out to APIs and databases, extracting data against a fixed schema.
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
Three real jobs, priced on this model
Each figure is derived from the prices above and moves when they do.
| Job | Tokens | Cost |
|---|---|---|
| One chat turn with a medium history | 1,500 in + 400 out | 55 toman |
| Summarising a ten-page document | 4,000 in + 600 out | 133 toman |
| Classifying a thousand short rows | 120,000 in + 20,000 out | 4,062 toman |
Frequently asked
- How do I call Ministral 3 3B 2512 from Iran?
- Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "mistralai/ministral-3b-2512". Nothing else in your code changes.
- What does Ministral 3 3B 2512 cost in toman?
- 29,013 toman per 1M input tokens and 29,013 toman per 1M output tokens; a 1,000-word request is around 75 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
- How much input does Ministral 3 3B 2512 take?
- Up to 131,072 tokens per request, roughly 98k words. A single answer can reach 104,857 tokens.
- Does Ministral 3 3B 2512 support streaming and tool calling?
- Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape. Image input is accepted through image_url, as a data URI or a public URL. Structured output through response_format works too.