Topic
Self-Hosted vs API Performance
13 posts on self-hosted vs api performance — part of benchmarks & performance on the n4n AI blog.
vLLM on A100 vs hosted API: time to first token
A head-to-head look at vLLM A100 vs API time to first token across cost, latency, and ops, with a verdict for self-hosted versus hosted LLM inference performance.
TGI vs API gateways: measuring end-to-end latency
A head-to-head comparison of TGI vs API gateway latency across capabilities, cost, throughput, and ergonomics, with a verdict for self-hosted vs API use.
Self-hosting Mixtral 8x7B: latency vs a pay-per-token API
Practical head-to-head: self-hosted Mixtral 8x7B latency vs API across cost, throughput, ergonomics, and scaling limits, with a use-case verdict.
Self-hosted Qwen2.5 72B throughput vs API-based inference
A head-to-head comparison of Qwen2.5 72B self-hosted throughput vs API inference across cost, latency, ergonomics, and limits, plus a verdict by use case.
Self-hosted Llama 3.3 70B: throughput at 50 concurrent users
Analysis of Llama 3.3 70B self-hosted concurrency throughput at 50 users: hardware, batching, tradeoffs vs API, and a decisive takeaway for engineers.
Self-hosted inference on 4x A6000 vs hosted API latency
A practitioner's comparison of 4x A6000 self-hosted inference vs API latency: cost model, throughput, ops burden, and when to choose each for production LLMs.
Self-hosted DeepSeek-V3 vs API: cost per token vs speed
Analyze DeepSeek-V3 self-hosted vs API cost and speed tradeoffs with concrete deployment examples to decide when owning GPUs beats paying per token.
Running Llama 3.1 8B locally vs calling an API endpoint
Head-to-head: Llama 3.1 8B local vs API latency, cost, and ergonomics compared for engineers deciding between self-hosted and managed inference.
Ollama on a single GPU vs API latency benchmarks
Compare Ollama on a single GPU vs API latency across capabilities, cost, throughput, and ergonomics with a head-to-head table and verdict.
llama.cpp on consumer GPUs vs API round-trip time
Compare llama.cpp consumer GPU vs API latency across cost, throughput, ergonomics, and limits, with a verdict for self-hosted versus API inference.
Is self-hosting Mistral Large actually faster than an API?
Self-hosted Mistral Large latency vs API: we break down real tradeoffs in time-to-first-token, throughput, and ops to show when self-hosting actually wins.
When self-hosting beats an API: a latency breakeven analysis
Engineer's guide to the self-hosted LLM vs API latency breakeven: benchmark APIs, model GPU throughput, and run shadow tests to decide what to run locally.
Self-hosted Llama 3.3 70B vs API latency, benchmarked
Practical latency and cost comparison of self-hosted Llama 3.3 70B versus API inference, with real deployment tradeoffs for engineers.
More topics in benchmarks & performance
- Agentic Workflow Performance Benchmarks14
- Benchmark Methodology and Measurement14
- Code Generation Latency for Dev Tools14
- Flagship Model Speed Showdown14
- Llama 4 Inference Speed by Provider14
- Price-Performance Rankings14
- Provider Uptime and Reliability Benchmarks14
- Reasoning Model Latency Overhead14
- Customer Support Chatbot Latency13
- DeepSeek Performance Benchmarks13
- GPU Inference Benchmarks13
- Long-Context Latency Benchmarks13