A little understanding. A better build.
You don’t need to be an engineer to choose your hardware. Start with these essentials.
Why hardware behaves so differently — for this model
Qwen3 8B · Q4 (Q4_K_M) · 8K context. Click a row for the full calculation.
1 · Memory bandwidth sets the speed limit for generation
To generate one token, the processor must read every weight it uses from memory. For Qwen3 8B at Q4 (Q4_K_M), that is about 4.3 GiB per token. So:
These are ceilings at 100% efficiency. Real software reaches roughly 60–85% of them, which the simulator models per architecture and backend. TFLOPS barely matter for single-user generation — memory does.
2 · VRAM vs unified memory: separate pools
64 GB RAM + 24 GB VRAM is not 88 GB of VRAM. Anything that spills into system RAM is read at system-RAM speed (and by the CPU), often making that part 10× slower. Unified memory lets the GPU use (most of) the whole pool directly.
3 · Quantization trades bits for size and speed
Halving bits per weight halves memory AND halves bytes read per token, so it roughly doubles generation speed. Q4–Q6 is the usual sweet spot; below Q4 quality drops noticeably.
4 · The KV cache grows with context
The model remembers each token as keys and values for every attention layer. It sits next to the weights in memory and is read for every new token, so long contexts cost both memory and speed. Grouped-query attention, sliding windows, MLA and hybrid linear attention all shrink it. KV-cache quantization (Q8) halves it.
5 · Mixture-of-Experts: stored vs used parameters
Qwen3 8B is dense: all 8.2B parameters are read for every token. An MoE model of the same size (e.g. Qwen3-30B-A3B: 30B stored, 3.3B active) needs the same memory but decodes many times faster. Select an MoE model to see its split.
This is why 128 GB unified-memory boxes with modest bandwidth are attractive for big MoE models such as gpt-oss-120b, GLM-4.5-Air or Qwen3-235B.
6 · TFLOPS and tensor cores matter for reading the prompt
Prompt processing (prefill) handles hundreds of tokens per weight read, so it is compute-bound. That is where tensor cores (NVIDIA), and to a lesser degree Apple's GPU, matter. A 8K prompt costs about 0.13 PFLOP before the first token appears.
Rule of thumb: generation speed ≈ memory bandwidth; time-to-first-token on long prompts ≈ compute. Macs and Strix Halo are much weaker here than NVIDIA GPUs.
7 · PCIe vs NVLink vs memory
Links between devices are 10–100× slower than VRAM. That is why models are split so that only small activations cross them — and why CPU offload over PCIe hurts prompt processing badly.
8 · Multi-GPU: more capacity, not more single-user speed
With llama.cpp's default layer split, GPU 1 runs layers 1–40, then GPU 2 runs 41–80. They take turns, so one user gets roughly the speed of one GPU reading the whole model — capacity adds up, speed does not. Tensor parallelism (vLLM) splits every layer, so GPUs read in parallel, but they must synchronise twice per layer, which is costly over PCIe.
Each GPU also needs its own runtime context, and layers can't be split across GPUs in layer mode, so you lose a little capacity per GPU.
9 · CPU offloading
If the model is too big for VRAM, llama.cpp can keep some layers in system RAM. Those layers are computed by the CPU at RAM speed (≈50–100 GB/s on desktops), and they dominate the time per token. For MoE models the smart move is to offload only the experts: attention stays on the GPU and only the few active experts are read from RAM each token.
Formulas & assumptions
- Memory = weights + KV cache × users + runtime/compute buffers (per GPU) + safety margin.
- Weights = parameters × bits/weight ÷ 8 (or measured file size).
- KV/token = Σ layers 2 × KV heads × head dim × bytes (MLA: latent + RoPE dims; sliding layers capped at window).
- Decode step = max(bytes ÷ (BW × η), FLOPs ÷ (FP32 × η)) + per-layer overhead + offload + interconnect.
- η = architecture × backend × quant-kernel × model-type factors from data/calibration.json, then corrected with measured benchmarks.
- Ranges: HIGH ±10%, MEDIUM −20/+15%, LOW −35/+15% around the central estimate.
Full derivation: docs/FORMULAS.md in the repository.