uttapen

Long context models, ranked

A long context buys you the option to skip retrieval altogether and hand the model the whole document. The cost grows linearly with what you send, so if most of the prompt repeats between calls, pick a model with cheap cached input — that one column changes the monthly bill more than the model choice does.

Best value in this category

Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.

#ModelContext≈ 1,000 wordsScore
1Gemini 2.5 Flash Lite (batch)1,048,5769486.9
2Qwen3.7 Flash1,000,0006085.6
3Muse Spark 1.2 Contributor1,048,57611384.9
4Muse Spark 1.3 Contributor1,048,57611384.9
5GPT-5 Nano (batch)400,0008584.7
6Nex-N2-Mini262,1444783.6
7Ling 3.0 Flash262,1443281.8
8GPT-4.1 Nano (batch)1,047,5769480.9
9GLM Flash Latest1,310,72011679.4
10Gemini 2.5 Flash Lite1,048,57618979.4
11GLM 5.3 Flash1,310,72012378.8
12Solar Pro 4524,2885778
13Qwen3.5-Flash1,000,00012377.9
14DeepSeek V4 Flash Latest1,310,7207977.6
15GPT-5 Nano400,00017077.2

Full list of long-context models for documents and codebases with prices →