Topic
Tokens-per-Second Throughput Rankings
13 posts on tokens-per-second throughput rankings — part of benchmarks & performance on the n4n AI blog.
Why tokens per second varies by provider and region
Why tokens per second varies by provider and region: a systems-level analysis of GPU SKUs, batching, and capacity allocation for LLM inference engineers.
What tokens per second actually means for your app
Tokens per second measures LLM output speed. Learn what does tokens per second mean for app latency, cost, and UX, plus how to measure it correctly.
Tokens per second benchmark: single request vs concurrent
Tokens per second single vs concurrent requests: how concurrency changes throughput, honest benchmark method, and which metric matters for LLM workloads.
Tokens per second benchmark: quantized vs full-precision
Benchmarking tokens per second quantized vs full precision: a head-to-head on capabilities, cost, latency, and ergonomics to guide LLM inference choices.
Tokens per second benchmark: GPT-5 vs DeepSeek V3
A head-to-head engineering comparison of tokens per second GPT-5 vs DeepSeek V3 across throughput, cost, latency, and ergonomics, with a use-case verdict.
Tokens per second benchmark for streaming chat apps
A practical analysis of tokens per second streaming chat benchmarks: why raw throughput misleads, what to measure, and how to test realistically for UX.
Tokens per second benchmark for coding-focused models
Analyze tokens per second coding models across tiers, separating latency from throughput, with a reproducible harness and a decision framework for engineers.
Qwen 3 235B tokens per second across five providers
We measured Qwen 3 235B tokens per second across five LLM providers. Analysis of throughput, concurrency, and tradeoffs for engineers shipping with large models.
Llama 4 Maverick tokens per second across providers
Analyze Llama 4 Maverick tokens per second across providers: why raw benchmarks mislead, which variables dominate throughput, and how to measure for production.
Highest throughput LLMs ranked by tokens per second
Engineer-focused ranking of the highest throughput LLMs ranked by tokens per second, covering serving stacks, measured generation speeds, and how to benchmark them.
Batch size and tokens per second: what changes at scale
Practical guide to scaling LLM inference: how batch size tokens per second interact with KV cache, latency, and continuous batching, with code and load tests.
Tokens per second benchmark: top 20 LLMs ranked
Ranked tokens per second LLM benchmark rankings for 20 models from 1B edge weights to GPT-4o, with throughput tiers, caveats, and a Python measurement snippet.
Tokens per second benchmark on Groq, Cerebras, and SambaNova
Practical head-to-head comparison of tokens per second on Groq, Cerebras, and SambaNova across cost, latency, ergonomics, and limits for engineers.
More topics in benchmarks & performance
- Agentic Workflow Performance Benchmarks14
- Benchmark Methodology and Measurement14
- Code Generation Latency for Dev Tools14
- Flagship Model Speed Showdown14
- Llama 4 Inference Speed by Provider14
- Price-Performance Rankings14
- Provider Uptime and Reliability Benchmarks14
- Reasoning Model Latency Overhead14
- Customer Support Chatbot Latency13
- DeepSeek Performance Benchmarks13
- GPU Inference Benchmarks13
- Long-Context Latency Benchmarks13