Batch models, ranked
If no user is staring at a spinner, you are paying for speed you do not need. A batch variant runs the same model behind the same endpoint at a lower rate, with latency measured in minutes rather than seconds. Good for overnight jobs, dataset labelling and bulk content; wrong for anything interactive.
Best value in this category
Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.
| # | Model | Context | ≈ 1,000 words | Score |
|---|---|---|---|---|
| 1 | Gemini 2.5 Flash Lite (batch) | 1,048,576 | 94 | 87.4 |
| 2 | GPT-5 Nano (batch) | 400,000 | 85 | 85.2 |
| 3 | GPT-4.1 Nano (batch) | 1,047,576 | 94 | 81.4 |
| 4 | GPT-5.6 Luna Pro (batch) | 1,050,000 | 264 | 76.2 |
| 5 | GPT-5.6 Luna (batch) | 1,050,000 | 264 | 76.2 |
| 6 | Gemini 3.1 Flash Lite (batch) | 1,048,576 | 330 | 73.8 |
| 7 | GPT-5.4 Nano (batch) | 400,000 | 273 | 72.5 |
| 8 | Qwen3.5-9B (batch) | 262,144 | 158 | 71 |
| 9 | GLM 5.3 Flash (batch) | 1,048,575 | 245 | 71 |
| 10 | DeepSeek V4 Flash 0731 (batch) | 1,048,576 | 158 | 69.8 |
| 11 | GPT-4o-mini (batch) | 128,000 | 141 | 69.7 |
| 12 | Gemini 2.5 Flash (batch) | 1,048,576 | 528 | 68.7 |
| 13 | Gemini 3.5 Flash Lite (batch) | 1,048,576 | 528 | 68.7 |
| 14 | GPT-5 Mini (batch) | 400,000 | 424 | 67.7 |
| 15 | Gemini 3 Flash Preview (batch) | 1,048,576 | 660 | 66.3 |