uttapen

Reasoning models, ranked

A reasoning model thinks in tokens before it answers, and those tokens are output tokens on your bill. The quality gain on multi-step problems is real, and so is the multiplier: we reserve roughly three times the usual output when a request goes to one of these, then charge only what the model actually used.

Best value in this category

Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.

#ModelContext≈ 1,000 wordsScore
1Gemini 2.5 Flash Lite (batch)1,048,5769486.2
2Qwen3.7 Flash1,000,0006084.9
3Muse Spark 1.2 Contributor1,048,57611384.2
4Muse Spark 1.3 Contributor1,048,57611384.2
5GPT-5 Nano (batch)400,0008583.9
6Nex-N2-Mini262,1444782.9
7Ling 3.0 Flash262,1443281.1
8GLM Flash Latest1,310,72011678.6
9Gemini 2.5 Flash Lite1,048,57618978.6
10GLM 5.3 Flash1,310,72012378
11Solar Pro 4524,2885777.2
12Qwen3.5-Flash1,000,00012377.1
13DeepSeek V4 Flash Latest1,310,7207976.8
14GPT-5 Nano400,00017076.4
15Qwen3.5-9B262,1449475.4

Full list of reasoning models, and what they cost with prices →