uttapen

GPT-3.5 Turbo 16k API: toman pricing and code

openai/gpt-3.5-turbo-16k

toolsjson
Input · per 1M tokens
870,375 toman
Output · per 1M tokens
1,160,500 toman
One 1,000-word request ≈
2,640 toman

You are billed for the usage the request actually reported. Prices follow the market.Pricing and top-ups

GPT-3.5 Turbo 16k example: tool calling

The example is picked from this model's own capabilities. Drop in your key and it runs as is.

from openai import OpenAI
import json

client = OpenAI(base_url="https://api.uttapen.ir/v1", api_key="sk-up-…")

tools = [{
    "type": "function",
    "function": {
        "name": "check_stock",
        "description": "Returns the stock level of a product",
        "parameters": {
            "type": "object",
            "properties": {"sku": {"type": "string"}},
            "required": ["sku"],
        },
    },
}]

messages = [{"role": "user", "content": "How many of the Nike NK-42 shoe are in stock?"}]
first = client.chat.completions.create(model="openai/gpt-3.5-turbo-16k", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]

# you run the function yourself — the model never touches the database
args = json.loads(call.function.arguments)
result = {"sku": args["sku"], "qty": 7}

messages += [first.choices[0].message, {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}]
final = client.chat.completions.create(model="openai/gpt-3.5-turbo-16k", messages=messages, tools=tools)
print(final.choices[0].message.content)

What is GPT-3.5 Turbo 16k good for?

GPT-3.5 Turbo 16k comes from OpenAI; in uttapen you reach it with the model id "openai/gpt-3.5-turbo-16k". It accepts up to 16,385 tokens of input per call, so a long document or several source files fit in a single request. A single response can run to 4,096 tokens. On price it sits in the "mid-range" band — cheaper than 265 and dearer than 141 of the other paid models in the catalogue.

What it can do beyond plain text: it supports tool calling, so it can invoke your own functions with valid arguments; it returns schema-valid JSON through response_format, ready to hand to your code. All of it works through the standard parameters of the official OpenAI SDK — no custom client, no wrapper.

Pricing is 870,375 toman per 1M input tokens and 1,160,500 per 1M output tokens. A 1,000-word round trip on GPT-3.5 Turbo 16k lands near 2,640 toman. You are always charged for the usage the request actually reported, never for the estimate, and a failed request costs nothing.

The closest alternative with the same capabilities from a different provider is GLM 5 Turbo: GPT-3.5 Turbo 16k works out roughly 1.3× more expensive, and its context window is smaller. Both run on the same key and the same code, so trying the other one is a single string change.

To make the figure concrete: 100,000 toman of credit buys roughly 38 thousand-word requests on GPT-3.5 Turbo 16k, and every 1,000 toman is about 379 words of round trip. A job with one million input tokens and one million output tokens comes to 2,030,875 toman in total. Filling this model's 16,385-token window costs 14,261 toman on the input side alone, which is the real reason to keep conversation history short.

Its nearest relative in the catalogue is "openai/gpt-3.5-turbo", and choosing between those two is where most people hesitate. On a thousand-word request this variant works out 3.5× dearer (2,640 against 754 toman). On parameters, this one takes max_completion_tokens.

OpenAI has 95 models in our catalogue; the cheapest is gpt-oss-20b at 60 toman per thousand words and the dearest o1-pro at 282,872. Among the less common parameters it accepts logit_bias, logprobs, max_completion_tokens, top_logprobs — all through the standard request body. It does not support include_reasoning, reasoning, which most models here do, so test before switching if your code relies on them. By context size the nearest option from another provider is Phi 4 at 16,384 tokens.

Good fits: product chatbots and internal assistants, where cost and quality have to balance, agents that reach out to APIs and databases, extracting data against a fixed schema.

Provider's own description

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up...

Three real jobs, priced on this model

Each figure is derived from the prices above and moves when they do.

JobTokensCost
One chat turn with a medium history1,500 in + 400 out1,770 toman
Summarising a ten-page document4,000 in + 600 out4,178 toman
Classifying a thousand short rows120,000 in + 20,000 out127,655 toman

Frequently asked

How do I call GPT-3.5 Turbo 16k from Iran?
Sign up with your mobile number, top the wallet up in toman, create an API key, then in the official OpenAI SDK point base_url at https://api.uttapen.ir/v1 and set model to "openai/gpt-3.5-turbo-16k". Nothing else in your code changes.
What does GPT-3.5 Turbo 16k cost in toman?
870,375 toman per 1M input tokens and 1,160,500 toman per 1M output tokens; a 1,000-word request is around 2,640 toman. You pay for the usage that request actually reported, and a failed request costs nothing.
How much input does GPT-3.5 Turbo 16k take?
Up to 16,385 tokens per request, roughly 12k words. A single answer can reach 4,096 tokens.
Does GPT-3.5 Turbo 16k support streaming and tool calling?
Streaming (stream=true) works on every model here. This one supports tool calling in the standard OpenAI shape. Structured output through response_format works too.