AH
Type to search...

LLM Inference Bench

M3 Max · 64 GB · Z8 G4 · 2× A6000  ·  452 configs · 822 rows  ·  updated 2026-08-15

822 of 822 rows

Decode = per-user generation t/s  ·  Throughput = total system output t/s (all concurrent users combined)  ·  TTFT = time to first token (warm)  ·  Cold = first-request TTFT  ·  Prefill = prompt eval t/s  ·  RSS = peak process RAM (mmap/tiered backends like omlx, vllm-mlx, llama-server may underreport — model not fully resident in RSS)  ·  MTP = Multi-Token Prediction draft length  ·  MTP Accept = draft acceptance rate  ·  TP = Tensor Parallel GPUs  ·  AA = Artificial Analysis Index. M3 Max · 64 GB: Apple M3 Max, 64 GB unified memory, macOS. Median of 3 runs. Z8 G4 · 2× A6000: HP Z8 G4, 2× NVIDIA RTX A6000 48 GB (Ampere GA102, 768 GB/s HBM) on PCIe 3.0 ×16 (~16 GB/s per direction) — no NVLink. Intel Xeon Gold 6226R (Cascade Lake), 256 GB DDR4. Two workstations connected via 1 GbE only; benchmarks are per-workstation. k3s + CUDA 13.