LLM Inference Bench
M3 Max · 64 GB · Z8 G4 · 2× A6000 · 452 configs · 822 rows · updated 2026-08-15
822 of 822 rows
Decode = per-user
generation t/s ·
Throughput = total
system output t/s (all concurrent users combined) ·
TTFT = time to first
token (warm) ·
Cold = first-request TTFT ·
Prefill = prompt eval t/s ·
RSS = peak process RAM (mmap/tiered backends like omlx, vllm-mlx, llama-server may underreport — model not fully resident in RSS) ·
MTP = Multi-Token Prediction
draft length ·
MTP Accept = draft
acceptance rate ·
TP = Tensor Parallel GPUs ·
AA = Artificial Analysis Index.
M3 Max · 64 GB:
Apple M3 Max, 64 GB unified memory, macOS. Median of 3 runs.
Z8 G4 · 2× A6000:
HP Z8 G4, 2× NVIDIA RTX A6000 48 GB
(Ampere GA102, 768 GB/s HBM) on
PCIe 3.0 ×16
(~16 GB/s per direction) — no NVLink.
Intel Xeon Gold 6226R (Cascade Lake),
256 GB DDR4. Two workstations connected via
1 GbE only;
benchmarks are per-workstation. k3s + CUDA 13.