uttapen

Magnum v4 72B API: toman pricing and code

anthracite-org/magnum-v4-72b

json
Input · per 1M tokens
725,313 toman
Output · per 1M tokens
1,450,625 toman
One 1,000-word request ≈
2,829 toman

You are billed for the usage the request actually reported. Prices follow the market.Pricing and top-ups

Magnum v4 72B example: JSON output against a schema

The example is picked from this model's own capabilities. Drop in your key and it runs as is.

from openai import OpenAI
import json

client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")

schema = {
    "name": "ticket",
    "schema": {
        "type": "object",
        "properties": {
            "category": {"type": "string", "enum": ["fani", "mali", "forush"]},
            "priority": {"type": "integer", "minimum": 1, "maximum": 5},
            "summary": {"type": "string"},
        },
        "required": ["category", "priority", "summary"],
        "additionalProperties": False,
    },
    "strict": True,
}

resp = client.chat.completions.create(
    model="anthracite-org/magnum-v4-72b",
    messages=[{"role": "user", "content": "Ticket: "For two days I cannot download my invoice and I was charged twice.""}],
    response_format={"type": "json_schema", "json_schema": schema},
)
print(json.loads(resp.choices[0].message.content))

What is Magnum v4 72B good for?

The id for Magnum v4 72B in our API is "anthracite-org/magnum-v4-72b", served from Anthracite Org. Context is 32,768 tokens; past that you have to summarise the history yourself. A single response can run to 4,096 tokens. On price it sits in the "mid-range" band — cheaper than 287 and dearer than 119 of the other paid models in the catalogue.

What it can do beyond plain text: it returns schema-valid JSON through response_format, ready to hand to your code. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper.

Pricing is 725,313 toman per 1M input tokens and 1,450,625 per 1M output tokens. A 1,000-word round trip on Magnum v4 72B lands near 2,829 toman. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.

The closest alternative with the same capabilities from a different provider is Palmyra X5: Magnum v4 72B works out roughly 1.1× more expensive, and its context window is smaller. Both run on the same key and the same code, so trying the other one is a single string change.

To make the figure concrete: 100,000 toman of credit buys roughly 35 thousand-word requests on Magnum v4 72B, and every 1,000 toman is about 353 words of round trip. A job with one million input tokens and one million output tokens comes to 2,175,938 toman in total. Filling this model's 32,768-token window costs 23,767 toman on the input side alone, which is the real reason to keep conversation history short.

Among the less common parameters it accepts logit_bias, logprobs, min_p, repetition_penalty, top_a, top_k — all through the standard request body. It does not support include_reasoning, reasoning, tool_choice, tools, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Aion-RP 1.0 (8B) at 32,768 tokens.

Good fits: product chatbots and internal assistants, where cost and quality have to balance, extracting data against a fixed schema.

Provider's own description

This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct).

Three real jobs, priced on this model

Each figure is derived from the prices above and moves when they do.

JobTokensCost
One chat turn with a medium history1,500 in + 400 out1,668 toman
Summarising a ten-page document4,000 in + 600 out3,772 toman
Classifying a thousand short rows120,000 in + 20,000 out116,050 toman

Frequently asked

How do I call Magnum v4 72B from Iran?
Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "anthracite-org/magnum-v4-72b". Nothing else in your code changes.
What does Magnum v4 72B cost in toman?
725,313 toman per 1M input tokens and 1,450,625 toman per 1M output tokens; a 1,000-word request is around 2,829 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
How much input does Magnum v4 72B take?
Up to 32,768 tokens per request, roughly 25k words. A single answer can reach 4,096 tokens.
Can I stream Magnum v4 72B's output?
Yes — with stream=true you get SSE events as the tokens are produced. This model has no tool calling and no image input, so pick a different one if you need either.