AH
Type to search...

LLM Inference Bench

Apple M3 Max · 64 GB  ·  248 configs · 679 rows · median of 3 runs  ·  updated 2026-07-07

679 of 679 rows

Arena ELO = LMSYS Chatbot Arena (model-level; higher = smarter)  ·  AA Index = Artificial Analysis Intelligence Index v4 (0–60, non-reasoning)  ·  TTFT = time to first token (warm, ms)  ·  Cold = first-request TTFT  ·  Decode = generation t/s  ·  Prefill = prompt eval t/s  ·  Total = wall-clock seconds  ·  Peak RSS = max process RAM. Median of 3 runs (Peak RSS = max). Q4 ≈ FP16 quality within ~1–3%.