Topic
GPU Inference Benchmarks
13 posts on gpu inference benchmarks — part of benchmarks & performance on the n4n AI blog.
Multi-GPU scaling: H100 clusters and inference throughput
Analysis of H100 cluster multi-GPU inference throughput: why scaling isn't linear, where bottlenecks shift, and how to architect GPU clusters for LLM serving.
Mixtral 8x22B inference: H100 vs A100 throughput
Practical comparison of Mixtral 8x22B H100 vs A100 throughput, covering hardware, cost, latency, and which GPU to pick for production inference.
MI300X vs H100: inference benchmarks for Llama 3
MI300X vs H100 inference benchmark for Llama 3: head-to-head on throughput, latency, cost, and ergonomics to guide GPU selection for teams.
Llama 3.1 405B inference: H100 vs H200 throughput
A head-to-head engineering comparison of Llama 3.1 405B H100 vs H200 throughput, cost, and ergonomics to guide GPU selection for production inference.
H200's extra memory bandwidth: does it cut latency?
Does the H200's extra memory bandwidth cut inference latency? We analyze H200 memory bandwidth inference latency versus H100 for LLM decode, prefill, and batch tradeoffs.
H100 SXM vs PCIe: inference latency differences
Practical comparison of H100 SXM vs PCIe inference latency differences: raw compute, NVLink, thermals, and cost tradeoffs for LLM serving.
GPU memory bandwidth and its effect on inference latency
Analyzes how GPU memory bandwidth drives inference latency for LLMs, with roofline math, batch tradeoffs, and quantization strategies for engineers.
B200 NVLink and its impact on large model inference
Analyzing how B200 NVLink bandwidth reshapes large model inference: intra-node tensor parallelism, latency tradeoffs, and when the hardware actually pays off.
H100 vs H200 vs B200: inference latency compared
Practical comparison of H100 vs H200 vs B200 inference latency across memory, cost, throughput, and ops, with a verdict for production LLM serving.
Choosing a GPU for LLM inference: H100 vs H200 vs B200
Practical guide to selecting the best GPU for LLM inference H100 H200 B200: compare VRAM, bandwidth, and real-world tradeoffs for production.
B200 vs H100: tokens per second on 70B models
B200 vs H100 tokens per second 70B: head-to-head comparison of throughput, memory, cost, and ergonomics for serving 70B LLMs in production.
B200 inference benchmarks for Llama 3.3 70B
A practical analysis of B200 Llama 3.3 70B inference benchmark results: real throughput, latency tradeoffs, and when Blackwell beats Hopper.
A100 vs H100: price-to-performance for LLM inference
Practical head-to-head comparison of A100 vs H100 price to performance inference for LLM serving: capabilities, cost, latency, and which GPU to pick per workload.
More topics in benchmarks & performance
- Agentic Workflow Performance Benchmarks14
- Benchmark Methodology and Measurement14
- Code Generation Latency for Dev Tools14
- Flagship Model Speed Showdown14
- Llama 4 Inference Speed by Provider14
- Price-Performance Rankings14
- Provider Uptime and Reliability Benchmarks14
- Reasoning Model Latency Overhead14
- Customer Support Chatbot Latency13
- DeepSeek Performance Benchmarks13
- Long-Context Latency Benchmarks13
- Model Size vs Inference Speed Tradeoffs13