Embeddings models, ranked
For non-Latin text the tokenizer matters as much as the vector quality, so prefer a multilingual model and test it on your own documents. Embeddings are billed on input tokens only, which makes the one-off indexing pass cheap; the recurring cost sits in the answering step, not here.
Best value in this category
Score out of 100: 50% cheapness, 30% capabilities, 20% context size — not a quality benchmark.
| # | Model | Context | ≈ 1,000 words | Score |
|---|---|---|---|---|
| No model in this category carries a price yet. | ||||
Full list of embedding models for semantic search with prices →