Topic
Embedding Model Throughput Benchmarks
12 posts on embedding model throughput benchmarks — part of benchmarks & performance on the n4n AI blog.
Voyage AI embeddings: throughput at large batch sizes
Analyze how batch size affects Voyage AI embeddings throughput, where gains plateau, and how to tune batching and concurrency for production embedding pipelines.
Self-hosted embedding models: throughput vs API options
Practical comparison of self-hosted embedding model throughput vs API options across cost, latency, and ops, with a verdict for engineering teams.
Parallel embedding requests: throughput scaling benchmarks
Analyze parallel embedding request throughput scaling under concurrency, batch limits, and rate caps, with practical async client patterns for engineers.
Nomic Embed vs text-embedding-3-small: throughput benchmarks
Head-to-head Nomic Embed vs text-embedding-3-small throughput: capabilities, cost, latency, ergonomics, and which embedding model to pick for your workload.
Jina Embeddings v3 throughput across batch sizes
A practical analysis of Jina Embeddings v3 throughput benchmark across batch sizes, covering memory bandwidth limits, late interaction overhead, and optimal batch sizing.
GTE-large vs BGE-large: embedding throughput compared
Practical head-to-head of GTE-large vs BGE-large throughput: capabilities, cost, latency, ergonomics, and which embedding model to pick for your workload.
Embedding 1 million documents: a throughput benchmark
A practical analysis of embedding 1 million documents: how batching, concurrency, and fallback shape your embedding large document set throughput benchmark.
Dimension size and its effect on embedding throughput
How embedding dimension size drives throughput tradeoffs in vector pipelines. A practical analysis of latency, batching, and cost for engineers.
text-embedding-3-large vs Cohere embed-v3 throughput
Head-to-head comparison of text-embedding-3-large vs Cohere embed-v3 throughput across capabilities, cost, latency, ergonomics, and limits for engineers.
Choosing an embedding model for high-throughput pipelines
A practical guide to selecting the best embedding model for throughput in production pipelines, covering benchmark methodology, tradeoffs, and code.
BGE-M3 vs OpenAI embeddings: tokens per second benchmarked
A head-to-head engineering comparison of BGE-M3 and OpenAI embeddings, focusing on real throughput, cost, and operational tradeoffs for production RAG.
Batch size and its effect on embedding throughput
Analyze how batch size drives embedding throughput on transformer encoders, where gains peak, and how to size batches for production embedding pipelines.
More topics in benchmarks & performance
- Agentic Workflow Performance Benchmarks14
- Benchmark Methodology and Measurement14
- Code Generation Latency for Dev Tools14
- Flagship Model Speed Showdown14
- Llama 4 Inference Speed by Provider14
- Price-Performance Rankings14
- Provider Uptime and Reliability Benchmarks14
- Reasoning Model Latency Overhead14
- Customer Support Chatbot Latency13
- DeepSeek Performance Benchmarks13
- GPU Inference Benchmarks13
- Long-Context Latency Benchmarks13