uttapen

Vision models, ranked

Every model here accepts image input. If your documents are in Persian, resolution decides the outcome more than the model does: send a high-DPI scan and skip aggressive JPEG compression. Pair it with a model that also supports structured output and the fields come back as JSON instead of prose.

Most used in this category

  1. 1GPT-4o-mini118 tok174,075

Best value in this category

Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.

#ModelContext≈ 1,000 wordsScore
1Gemini 2.5 Flash Lite (batch)1,048,5769486.9
2Qwen3.7 Flash1,000,0006085.6
3Muse Spark 1.2 Contributor1,048,57611384.9
4Muse Spark 1.3 Contributor1,048,57611384.9
5GPT-5 Nano (batch)400,0008584.7
6Nex-N2-Mini262,1444783.6
7GPT-4.1 Nano (batch)1,047,5769480.9
8GLM Flash Latest1,310,72011679.4
9Gemini 2.5 Flash Lite1,048,57618979.4
10GLM 5.3 Flash1,310,72012378.8
11Qwen3.5-Flash1,000,00012377.9
12GPT-5 Nano400,00017077.2
13Qwen3.5-9B262,1449476.1
14GPT-5.6 Luna Pro (batch)1,050,00026475.7
15GPT-5.6 Luna (batch)1,050,00026475.7

Full list of models that read images with prices →