Long context models, ranked
A long context buys you the option to skip retrieval altogether and hand the model the whole document. The cost grows linearly with what you send, so if most of the prompt repeats between calls, pick a model with cheap cached input — that one column changes the monthly bill more than the model choice does.
Best value in this category
Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.
| # | Model | Context | ≈ 1,000 words | Score |
|---|---|---|---|---|
| 1 | Gemini 2.5 Flash Lite (batch) | 1,048,576 | 94 | 86.9 |
| 2 | Qwen3.7 Flash | 1,000,000 | 60 | 85.6 |
| 3 | Muse Spark 1.2 Contributor | 1,048,576 | 113 | 84.9 |
| 4 | Muse Spark 1.3 Contributor | 1,048,576 | 113 | 84.9 |
| 5 | GPT-5 Nano (batch) | 400,000 | 85 | 84.7 |
| 6 | Nex-N2-Mini | 262,144 | 47 | 83.6 |
| 7 | Ling 3.0 Flash | 262,144 | 32 | 81.8 |
| 8 | GPT-4.1 Nano (batch) | 1,047,576 | 94 | 80.9 |
| 9 | GLM Flash Latest | 1,310,720 | 116 | 79.4 |
| 10 | Gemini 2.5 Flash Lite | 1,048,576 | 189 | 79.4 |
| 11 | GLM 5.3 Flash | 1,310,720 | 123 | 78.8 |
| 12 | Solar Pro 4 | 524,288 | 57 | 78 |
| 13 | Qwen3.5-Flash | 1,000,000 | 123 | 77.9 |
| 14 | DeepSeek V4 Flash Latest | 1,310,720 | 79 | 77.6 |
| 15 | GPT-5 Nano | 400,000 | 170 | 77.2 |
Full list of long-context models for documents and codebases with prices →