uttapen

Hermes 3 70B Instruct vs Llama 4 Maverick

Both sit behind the same key and the same code on uttapen — the only thing that changes is the model string, so you can run each of them against your own workload without touching anything else. On cost, Llama 4 Maverick comes out about 1.6× cheaper.

FeatureHermes 3 70B InstructLlama 4 Maverick
ProviderNous ResearchMeta
Model idnousresearch/hermes-3-llama-3.1-70bmeta-llama/llama-4-maverick
Context131,072 tokens1,048,576 tokens
Max output16,384 tokens115,200 tokens
Input / 1M tokens203,088 toman58,025 toman
Output / 1M tokens203,088 toman201,927 toman
≈ one 1,000-word request528 toman338 toman
Input cachingNoNo
Image inputNoYes
File inputNoNo
Tool callingNoYes
JSON outputYesYes
Reasoning modeNoNo

Which one for what?

Hermes 3 70B Instruct

  • Chat, summarising, and everyday text generation
528 toman per 1,000 words · Model page

Llama 4 Maverick

  • Whole documents or a codebase in a single request
  • Reading screenshots, invoices, and scanned forms
  • Agents that call your own APIs and database
  • High-volume work where cost per call decides
338 toman per 1,000 words · Model page

Run both with the same code

Swap the model value between the two ids and send the same request twice; what each answer cost comes back in the X-Uttapen-Cost-Toman header.

curl https://api.uttapen.ir/v1/chat/completions \
  -H "Authorization: Bearer sk-up-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nousresearch/hermes-3-llama-3.1-70b",
    "messages": [{"role": "user", "content": "Introduce yourself in one sentence."}],
    "stream": true
  }'

model = "meta-llama/llama-4-maverick"

More comparisons: all pairs · model rankings · full catalogue

base_url = https://api.uttapen.ir/v1