uttapen

Tool calling models, ranked

An agent lives or dies on two things: calling the right function with valid arguments, and returning JSON your code can parse without a retry. Everything here does both. In a multi-step loop, a cheaper model with a low error rate beats an expensive one with the same error rate, because every retry pays twice.

Most used in this category

  1. 1GPT-4o-mini118 tok174,075

Best value in this category

Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.

#ModelContext≈ 1,000 wordsScore
1Gemini 2.5 Flash Lite (batch)1,048,5769485.8
2Qwen3.7 Flash1,000,0006084.5
3Muse Spark 1.2 Contributor1,048,57611383.8
4Muse Spark 1.3 Contributor1,048,57611383.8
5GPT-5 Nano (batch)400,0008583.5
6Nex-N2-Mini262,1444782.5
7Ling 3.0 Flash262,1443280.7
8GPT-4.1 Nano (batch)1,047,5769479.8
9GLM Flash Latest1,310,72011678.3
10Gemini 2.5 Flash Lite1,048,57618978.2
11GLM 5.3 Flash1,310,72012377.6
12Solar Pro 4524,2885776.8
13Qwen3.5-Flash1,000,00012376.7
14DeepSeek V4 Flash Latest1,310,7207976.4
15GPT-5 Nano400,00017076

Full list of models that call your tools reliably with prices →