CALIBRATION · 13 ROWS · VERIFIED 2026-07-13 → 2026-07-17
What the community has measured.
13 model × hardware combinations cross-verified against published benchmark data, republished here as a single dated audit log — exact runner, quant, Flash Attention setting, context length, and per-row verification dates.
The editorial verdicts on the planner have to land within shouting distance of measured reality. This is the audit log. Every row below was cross-verified between 2026-07-13 and 2026-07-17 against its primary source — linked under every row, so you can check us — using the runner and quant stated. Single-stream throughput, not batched. PP = prompt-processing tok/s; TG = generation tok/s. Where a number is reported as a band, it’s because the published figures varied meaningfully across reproductions — and saying so is honest.
Rebuilt July 2026. The previous table benchmarked Llama 3.1 and Qwen3 — models the planner had stopped recommending. It measured things we don’t pick and skipped things we do, which makes it decoration rather than evidence. Every row is now a current pick, and every row carries its own source link and the date that source ran the test. Two dates matter and we show both: when the benchmark was measured, and when we last verified it against the source.
What is deliberately not here matters too. There are no rows for the Mac Studio M3 Ultra 96 GB, the Mac mini M4, or the Intel Arc B580, because no credible current-model benchmark exists for them — and we would rather show a gap than an estimate. Several sites will happily give you a number for those combinations; at least one generates them from a formula and says so in its own footer. A missing row beats an invented one.
Popular companion pages: the Mac Studio M3 Ultra 96 GB workstation, the AMD ROCm guide, and the find-by-model hardware lookup.
References — community benchmark sources
Every row above links its own primary source directly, so you can check any single number without trusting this list. These are the sources we consider credible enough to draw from in the first place — they publish a methodology, state their runner and build, and report results that survive arithmetic.
What we refuse to cite, and why you should care. A large share of the pages ranking for “[model] on [GPU] tok/s” are not measurements. Some are calculators — WillItRunAI states in its own footer that “all estimates are approximations based on mathematical models and public specifications”, and it returns the same tok/s for a 96 GB and a 256 GB machine because a formula cannot tell them apart. Others are simply broken: we found sites reporting prompt processing an order of magnitude below token generation, which is not physically possible on these runners, and one claiming an RTX 3090 beats an RTX 4090 by 15× on the same workload. A third group reprints Hardware Corner’s figures verbatim and reads like independent corroboration when it is nothing of the sort. We cited an estimate engine ourselves in planner copy until July 2026 — this list exists so that does not happen twice.
- LocalScore — community-submitted llama.cpp benchmark database with per-accelerator pages. Primary source for RTX 5090 / 5060 Ti / 5070 Ti / RTX 3090 figures.
- Hardware Corner — independent local-LLM hardware benchmarks. Primary source for Mac Apple Silicon (M3/M4/M5) figures across the Mac product line.
- llama.cpp GitHub Discussions + Issues — runner-author + maintainer benchmarks and the canonical place to find AMD ROCm vs Vulkan back-and-forth.
- NVIDIA Developer Forums — vendor-published DGX Spark throughput figures + community reproductions on driver versions (e.g. v2.1 patches for the 122B-A10B benchmark).
- Ollama blog — runner-author throughput claims for new backends (Metal, MLX, ROCm) at release.
- HuggingFace model card discussions on the canonical model repos (e.g. Qwen3.5-35B-A3B discussions, ubergarm GGUF quants for community llama-bench results).
- r/LocalLLaMA benchmark threads — community lived experience, especially useful for catching regressions in newer driver/runner builds.
- Databasemart benchmarks — server-hosted multi-GPU figures (dual 5090, dual A6000).
Want a row updated, added, or corrected? Send a reproducible benchmark — model, quant, runner, hardware, prompt, measured PP and TG — and we’ll cross-verify against the existing sources and either update the row or add a new one.