uttapen

GLM 4.5 Air API: toman pricing and code

z-ai/glm-4.5-air

toolsreasoning
Input · per 1M tokens
37,716 toman
Output · per 1M tokens
246,606 toman
One 1,000-word request ≈
370 toman

You are billed for the usage the request actually reported. Prices follow the market. Cached input: 7,253 toman / 1M.Pricing and top-ups

GLM 4.5 Air example: tool calling

The example is picked from this model's own capabilities. Drop in your key and it runs as is.

from openai import OpenAI
import json

client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")

tools = [{
    "type": "function",
    "function": {
        "name": "check_stock",
        "description": "Returns the stock level of a product",
        "parameters": {
            "type": "object",
            "properties": {"sku": {"type": "string"}},
            "required": ["sku"],
        },
    },
}]

messages = [{"role": "user", "content": "How many of the Nike NK-42 shoe are in stock?"}]
first = client.chat.completions.create(model="z-ai/glm-4.5-air", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]

# you run the function yourself — the model never touches the database
args = json.loads(call.function.arguments)
result = {"sku": args["sku"], "qty": 7}

messages += [first.choices[0].message, {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}]
final = client.chat.completions.create(model="z-ai/glm-4.5-air", messages=messages, tools=tools)
print(final.choices[0].message.content)

What is GLM 4.5 Air good for?

The id for GLM 4.5 Air in our API is "z-ai/glm-4.5-air", served from Z.ai (GLM). Context is 131,072 tokens; past that you have to summarise the history yourself. A single response can run to 98,304 tokens. On price it sits in the "cheap" band — cheaper than 125 and dearer than 281 of the other paid models in the catalogue.

Capabilities available on this id: it supports tool calling, so it can invoke your own functions with valid arguments; it has a reasoning mode that pays off on multi-step problems, maths and debugging. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper. Keep in mind that reasoning tokens are output tokens and do appear on the bill.

For a back-of-envelope figure: about 370 toman per 1,000-word exchange (37,716 in, 246,606 out, per 1M tokens). It supports cached input: repeated context is billed at 7,253 toman per 1M, which matters a lot if your system prompt is long. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.

The closest alternative with the same capabilities from a different provider is Llama 3.1 Euryale 70B v2.2: GLM 4.5 Air works out roughly 1.7× cheaper, and its context window is the same size. Both run on the same key and the same code, so trying the other one is a single string change.

To make the figure concrete: 100,000 toman of credit buys roughly 270 thousand-word requests on GLM 4.5 Air, and every 1,000 toman is about 2,703 words of round trip. A job with one million input tokens and one million output tokens comes to 284,323 toman in total. Filling this model's 131,072-token window costs 4,944 toman on the input side alone, which is the real reason to keep conversation history short.

Z.ai (GLM) has 17 models in our catalogue; the cheapest is GLM Flash Latest at 116 toman per thousand words and the dearest GLM 5.3 at 2,188. Among the less common parameters it accepts repetition_penalty, top_k — all through the standard request body. It does not support response_format, structured_outputs, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Aion-2.0 at 131,072 tokens.

What to use it for: high-volume work such as classification, tagging and bulk summarising, logic puzzles and code review, agents that reach out to APIs and databases.

Provider's own description

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

Three real jobs, priced on this model

Each figure is derived from the prices above and moves when they do.

JobTokensCost
One chat turn with a medium history1,500 in + 400 out155 toman
Summarising a ten-page document4,000 in + 600 out299 toman
Classifying a thousand short rows120,000 in + 20,000 out9,458 toman

Frequently asked

How do I call GLM 4.5 Air from Iran?
Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "z-ai/glm-4.5-air". Nothing else in your code changes.
What does GLM 4.5 Air cost in toman?
37,716 toman per 1M input tokens and 246,606 toman per 1M output tokens; a 1,000-word request is around 370 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
How much input does GLM 4.5 Air take?
Up to 131,072 tokens per request, roughly 98k words. A single answer can reach 98,304 tokens.
Does GLM 4.5 Air support streaming and tool calling?
Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape.