uttapen

Batch models, ranked

If no user is staring at a spinner, you are paying for speed you do not need. A batch variant runs the same model behind the same endpoint at a lower rate, with latency measured in minutes rather than seconds. Good for overnight jobs, dataset labelling and bulk content; wrong for anything interactive.

Best value in this category

Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.

#ModelContext≈ 1,000 wordsScore
1Gemini 2.5 Flash Lite (batch)1,048,5769487.4
2GPT-5 Nano (batch)400,0008585.2
3GPT-4.1 Nano (batch)1,047,5769481.4
4GPT-5.6 Luna Pro (batch)1,050,00026476.2
5GPT-5.6 Luna (batch)1,050,00026476.2
6Gemini 3.1 Flash Lite (batch)1,048,57633073.8
7GPT-5.4 Nano (batch)400,00027372.5
8Qwen3.5-9B (batch)262,14415871
9GLM 5.3 Flash (batch)1,048,57524571
10DeepSeek V4 Flash 0731 (batch)1,048,57615869.8
11GPT-4o-mini (batch)128,00014169.7
12Gemini 2.5 Flash (batch)1,048,57652868.7
13Gemini 3.5 Flash Lite (batch)1,048,57652868.7
14GPT-5 Mini (batch)400,00042467.7
15Gemini 3 Flash Preview (batch)1,048,57666066.3

Full list of batch variants (cheaper, slower) with prices →