Llama 3.1 Euryale 70B v2.2 API: toman pricing and code
sao10k/l3.1-euryale-70b
toolsjsonYou are billed for the usage the request actually reported. Prices follow the market.Pricing and top-ups
Llama 3.1 Euryale 70B v2.2 example: JSON output against a schema
The example is picked from this model's own capabilities. Drop in your key and it runs as is.
from openai import OpenAI
import json
client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")
schema = {
"name": "ticket",
"schema": {
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["fani", "mali", "forush"]},
"priority": {"type": "integer", "minimum": 1, "maximum": 5},
"summary": {"type": "string"},
},
"required": ["category", "priority", "summary"],
"additionalProperties": False,
},
"strict": True,
}
resp = client.chat.completions.create(
model="sao10k/l3.1-euryale-70b",
messages=[{"role": "user", "content": "Ticket: "For two days I cannot download my invoice and I was charged twice.""}],
response_format={"type": "json_schema", "json_schema": schema},
)
print(json.loads(resp.choices[0].message.content))import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uttapen.ir/v1", apiKey: "sk-up-…" });
const resp = await client.chat.completions.create({
model: "sao10k/l3.1-euryale-70b",
messages: [{ role: "user", content: "Ticket: "For two days I cannot download my invoice and I was charged twice."" }],
response_format: {
type: "json_schema",
json_schema: {
name: "ticket",
strict: true,
schema: {
type: "object",
properties: {
category: { type: "string", enum: ["fani", "mali", "forush"] },
priority: { type: "integer", minimum: 1, maximum: 5 },
summary: { type: "string" },
},
required: ["category", "priority", "summary"],
additionalProperties: false,
},
},
},
});
console.log(JSON.parse(resp.choices[0].message.content));curl https://api.uttapen.ir/v1/chat/completions \
-H "Authorization: Bearer sk-up-…" \
-H "Content-Type: application/json" \
-d '{
"model": "sao10k/l3.1-euryale-70b",
"messages": [{"role": "user", "content": "Classify the ticket and give it a priority from 1 to 5."}],
"response_format": {"type": "json_object"}
}'What is Llama 3.1 Euryale 70B v2.2 good for?
Sao10k publishes this model; we expose it under the id "sao10k/l3.1-euryale-70b". Its context window is 131,072 tokens, roughly 98k English words in one request. A single response can run to 16,384 tokens. On price it sits in the "cheap" band — cheaper than 125 and dearer than 281 of the other paid models in the catalogue.
What you get on top of text in, text out: it supports tool calling, so it can invoke your own functions with valid arguments; it returns schema-valid JSON through response_format, ready to hand to your code. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper.
Input runs at 246,606 toman per 1M tokens and output at 246,606 — output costs 1× input, so trimming the answer saves more than trimming the prompt. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.
The closest alternative with the same capabilities from a different provider is GLM 4.5 Air: Llama 3.1 Euryale 70B v2.2 works out roughly 1.7× more expensive, and its context window is the same size. Both run on the same key and the same code, so trying the other one is a single string change.
To make the figure concrete: 100,000 toman of credit buys roughly 156 thousand-word requests on Llama 3.1 Euryale 70B v2.2, and every 1,000 toman is about 1,560 words of round trip. A job with one million input tokens and one million output tokens comes to 493,213 toman in total. Filling this model's 131,072-token window costs 32,323 toman on the input side alone, which is the real reason to keep conversation history short.
Sao10k has 3 models in our catalogue; the cheapest is Llama 3 8B Lunaris at 34 toman per thousand words and the dearest Llama 3.1 Euryale 70B v2.2 at 641. Among the less common parameters it accepts logit_bias, min_p, repetition_penalty, top_k — all through the standard request body. It does not support include_reasoning, reasoning, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Aion-2.0 at 131,072 tokens.
Where it makes sense: high-volume work such as classification, tagging and bulk summarising, agents that reach out to APIs and databases, extracting data against a fixed schema.
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
Three real jobs, priced on this model
Each figure is derived from the prices above and moves when they do.
| Job | Tokens | Cost |
|---|---|---|
| One chat turn with a medium history | 1,500 in + 400 out | 469 toman |
| Summarising a ten-page document | 4,000 in + 600 out | 1,134 toman |
| Classifying a thousand short rows | 120,000 in + 20,000 out | 34,525 toman |
Frequently asked
- How do I call Llama 3.1 Euryale 70B v2.2 from Iran?
- Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "sao10k/l3.1-euryale-70b". Nothing else in your code changes.
- What does Llama 3.1 Euryale 70B v2.2 cost in toman?
- 246,606 toman per 1M input tokens and 246,606 toman per 1M output tokens; a 1,000-word request is around 641 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
- How much input does Llama 3.1 Euryale 70B v2.2 take?
- Up to 131,072 tokens per request, roughly 98k words. A single answer can reach 16,384 tokens.
- Does Llama 3.1 Euryale 70B v2.2 support streaming and tool calling?
- Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape. Structured output through response_format works too.