uttapen

Codestral 2508 API: toman pricing and code

mistralai/codestral-2508

toolsjsonfiles
Input · per 1M tokens
87,038 toman
Output · per 1M tokens
261,113 toman
One 1,000-word request ≈
453 toman

You are billed for the usage the request actually reported. Prices follow the market. Cached input: 8,704 toman / 1M.Pricing and top-ups

Codestral 2508 example: JSON output against a schema

The example is picked from this model's own capabilities. Drop in your key and it runs as is.

from openai import OpenAI
import json

client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")

schema = {
    "name": "ticket",
    "schema": {
        "type": "object",
        "properties": {
            "category": {"type": "string", "enum": ["fani", "mali", "forush"]},
            "priority": {"type": "integer", "minimum": 1, "maximum": 5},
            "summary": {"type": "string"},
        },
        "required": ["category", "priority", "summary"],
        "additionalProperties": False,
    },
    "strict": True,
}

resp = client.chat.completions.create(
    model="mistralai/codestral-2508",
    messages=[{"role": "user", "content": "Ticket: "For two days I cannot download my invoice and I was charged twice.""}],
    response_format={"type": "json_schema", "json_schema": schema},
)
print(json.loads(resp.choices[0].message.content))

What is Codestral 2508 good for?

Mistral publishes this model; we expose it under the id "mistralai/codestral-2508". Its context window is 256,000 tokens, roughly 192k English words in one request. A single response can run to 204,800 tokens. On price it sits in the "cheap" band — cheaper than 129 and dearer than 277 of the other paid models in the catalogue.

Capabilities available on this id: it takes files such as PDFs as input; it supports tool calling, so it can invoke your own functions with valid arguments; it returns schema-valid JSON through response_format, ready to hand to your code. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper.

For a back-of-envelope figure: about 453 toman per 1,000-word exchange (87,038 in, 261,113 out, per 1M tokens). It supports cached input: repeated context is billed at 8,704 toman per 1M, which matters a lot if your system prompt is long. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.

The closest alternative with the same capabilities from a different provider is DeepSeek V3: Codestral 2508 works out roughly 1× cheaper, and its context window is larger. Both run on the same key and the same code, so trying the other one is a single string change.

To make the figure concrete: 100,000 toman of credit buys roughly 221 thousand-word requests on Codestral 2508, and every 1,000 toman is about 2,208 words of round trip. A job with one million input tokens and one million output tokens comes to 348,150 toman in total. Filling this model's 256,000-token window costs 22,282 toman on the input side alone, which is the real reason to keep conversation history short.

Its nearest relative in the catalogue is "mistralai/devstral-2512", and choosing between those two is where most people hesitate. On a thousand-word request this variant works out 2× cheaper (453 against 905 toman). The context windows differ too: 256,000 against 262,144 tokens. Maximum answer length differs as well: 204,800 against 209,715 tokens. On parameters, this one takes prediction.

Mistral has 20 models in our catalogue; the cheapest is Mistral Nemo at 18 toman per thousand words and the dearest Mistral Medium 3.5 at 3,394. Among the less common parameters it accepts prediction — all through the standard request body. It does not support include_reasoning, reasoning, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Command A at 256,000 tokens.

What to use it for: high-volume work such as classification, tagging and bulk summarising, agents that reach out to APIs and databases, extracting data against a fixed schema, analysing a long document or codebase in one request.

Provider's own description

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

Three real jobs, priced on this model

Each figure is derived from the prices above and moves when they do.

JobTokensCost
One chat turn with a medium history1,500 in + 400 out235 toman
Summarising a ten-page document4,000 in + 600 out505 toman
Classifying a thousand short rows120,000 in + 20,000 out15,667 toman

Frequently asked

How do I call Codestral 2508 from Iran?
Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "mistralai/codestral-2508". Nothing else in your code changes.
What does Codestral 2508 cost in toman?
87,038 toman per 1M input tokens and 261,113 toman per 1M output tokens; a 1,000-word request is around 453 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
How much input does Codestral 2508 take?
Up to 256,000 tokens per request, roughly 192k words. A single answer can reach 204,800 tokens.
Does Codestral 2508 support streaming and tool calling?
Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape. Structured output through response_format works too.