LLM Inference Bench
Apple M3 Max · 64 GB · 248 configs · 679 rows · median of 3 runs · updated 2026-07-07
679 of 679 rows
Arena ELO = LMSYS
Chatbot Arena (model-level; higher = smarter) ·
AA Index = Artificial
Analysis Intelligence Index v4 (0–60, non-reasoning) ·
TTFT = time to
first token (warm, ms) ·
Cold = first-request
TTFT ·
Decode = generation
t/s ·
Prefill = prompt
eval t/s ·
Total = wall-clock
seconds ·
Peak RSS = max
process RAM. Median of 3 runs (Peak RSS = max). Q4 ≈ FP16 quality within ~1–3%.