n4nAI

Topic

GPU Inference Benchmarks

13 posts on gpu inference benchmarks — part of benchmarks & performance on the n4n AI blog.

Benchmarks & performanceAnalysis

Multi-GPU scaling: H100 clusters and inference throughput

Analysis of H100 cluster multi-GPU inference throughput: why scaling isn't linear, where bottlenecks shift, and how to architect GPU clusters for LLM serving.

4 min read
Benchmarks & performanceComparison

Mixtral 8x22B inference: H100 vs A100 throughput

Practical comparison of Mixtral 8x22B H100 vs A100 throughput, covering hardware, cost, latency, and which GPU to pick for production inference.

4 min read
Benchmarks & performanceComparison

MI300X vs H100: inference benchmarks for Llama 3

MI300X vs H100 inference benchmark for Llama 3: head-to-head on throughput, latency, cost, and ergonomics to guide GPU selection for teams.

5 min read
Benchmarks & performanceComparison

Llama 3.1 405B inference: H100 vs H200 throughput

A head-to-head engineering comparison of Llama 3.1 405B H100 vs H200 throughput, cost, and ergonomics to guide GPU selection for production inference.

4 min read
Benchmarks & performanceAnalysis

H200's extra memory bandwidth: does it cut latency?

Does the H200's extra memory bandwidth cut inference latency? We analyze H200 memory bandwidth inference latency versus H100 for LLM decode, prefill, and batch tradeoffs.

4 min read
Benchmarks & performanceComparison

H100 SXM vs PCIe: inference latency differences

Practical comparison of H100 SXM vs PCIe inference latency differences: raw compute, NVLink, thermals, and cost tradeoffs for LLM serving.

5 min read
Benchmarks & performanceAnalysis

GPU memory bandwidth and its effect on inference latency

Analyzes how GPU memory bandwidth drives inference latency for LLMs, with roofline math, batch tradeoffs, and quantization strategies for engineers.

3 min read
Benchmarks & performanceAnalysis

B200 NVLink and its impact on large model inference

Analyzing how B200 NVLink bandwidth reshapes large model inference: intra-node tensor parallelism, latency tradeoffs, and when the hardware actually pays off.

4 min read
Benchmarks & performanceComparison

H100 vs H200 vs B200: inference latency compared

Practical comparison of H100 vs H200 vs B200 inference latency across memory, cost, throughput, and ops, with a verdict for production LLM serving.

4 min read
Benchmarks & performanceGuide

Choosing a GPU for LLM inference: H100 vs H200 vs B200

Practical guide to selecting the best GPU for LLM inference H100 H200 B200: compare VRAM, bandwidth, and real-world tradeoffs for production.

4 min read
Benchmarks & performanceComparison

B200 vs H100: tokens per second on 70B models

B200 vs H100 tokens per second 70B: head-to-head comparison of throughput, memory, cost, and ergonomics for serving 70B LLMs in production.

3 min read
Benchmarks & performanceAnalysis

B200 inference benchmarks for Llama 3.3 70B

A practical analysis of B200 Llama 3.3 70B inference benchmark results: real throughput, latency tradeoffs, and when Blackwell beats Hopper.

4 min read
Benchmarks & performanceComparison

A100 vs H100: price-to-performance for LLM inference

Practical head-to-head comparison of A100 vs H100 price to performance inference for LLM serving: capabilities, cost, latency, and which GPU to pick per workload.

4 min read