llm4funAI HARDWARE, EXPLAINED.
THE WORKBENCH
llm4fun / THE WORKBENCH

A little understanding. A better build.

You don’t need to be an engineer to choose your hardware. Start with these essentials.

SIMULATINGQwen3 8B
Q4 8K context 30 tok/s target

Why hardware behaves so differently — for this model

Qwen3 8B · Q4 (Q4_K_M) · 8K context. Click a row for the full calculation.

1 · Memory bandwidth sets the speed limit for generation

To generate one token, the processor must read every weight it uses from memory. For Qwen3 8B at Q4 (Q4_K_M), that is about 4.3 GiB per token. So:

tokens/s ≤ memory bandwidth ÷ bytes read per token
Strix Halo · 256 GB/s
56 tok/s
DGX Spark · 273 GB/s
60 tok/s
M2 Ultra · 800 GB/s
180 tok/s
RTX 3090 · 936 GB/s
200 tok/s
RTX 5090 · 1792 GB/s
390 tok/s
tok/s ceiling

These are ceilings at 100% efficiency. Real software reaches roughly 60–85% of them, which the simulator models per architecture and backend. TFLOPS barely matter for single-user generation — memory does.

2 · VRAM vs unified memory: separate pools

Gaming / workstation PC
System RAM · 64 GB DDR5
~80–96 GB/s · CPU only
↕ PCIe 4.0 x16 ≈ 32 GB/s
GPU VRAM · 24 GB GDDR6X
936 GB/s · GPU
Mac / DGX Spark / Strix Halo
128 GB unified memory
273–819 GB/s · CPU + GPU

64 GB RAM + 24 GB VRAM is not 88 GB of VRAM. Anything that spills into system RAM is read at system-RAM speed (and by the CPU), often making that part 10× slower. Unified memory lets the GPU use (most of) the whole pool directly.

3 · Quantization trades bits for size and speed

FP16 / BF16 · 16 bits
15 GiB
FP8 · 8.03 bits
7.7 GiB
Q8 (Q8_0) · 8.5 bits
8.1 GiB
Q6 (Q6_K) · 6.57 bits
6.3 GiB
Q5 (Q5_K_M) · 5.67 bits
5.4 GiB
Q4 (Q4_K_M) · 4.83 bits
4.6 GiB
Q3 (Q3_K_M) · 3.9 bits
3.7 GiB
GiB

Halving bits per weight halves memory AND halves bytes read per token, so it roughly doubles generation speed. Q4–Q6 is the usual sweet spot; below Q4 quality drops noticeably.

4 · The KV cache grows with context

4K tokens
0.56 GiB
16K tokens
2.3 GiB
32K tokens
4.5 GiB
64K tokens
9 GiB
128K tokens
18 GiB
256K tokens
36 GiB
GiB

The model remembers each token as keys and values for every attention layer. It sits next to the weights in memory and is read for every new token, so long contexts cost both memory and speed. Grouped-query attention, sliding windows, MLA and hybrid linear attention all shrink it. KV-cache quantization (Q8) halves it.

5 · Mixture-of-Experts: stored vs used parameters

Qwen3 8B is dense: all 8.2B parameters are read for every token. An MoE model of the same size (e.g. Qwen3-30B-A3B: 30B stored, 3.3B active) needs the same memory but decodes many times faster. Select an MoE model to see its split.

This is why 128 GB unified-memory boxes with modest bandwidth are attractive for big MoE models such as gpt-oss-120b, GLM-4.5-Air or Qwen3-235B.

6 · TFLOPS and tensor cores matter for reading the prompt

Prompt processing (prefill) handles hundreds of tokens per weight read, so it is compute-bound. That is where tensor cores (NVIDIA), and to a lesser degree Apple's GPU, matter. A 8K prompt costs about 0.13 PFLOP before the first token appears.

Rule of thumb: generation speed ≈ memory bandwidth; time-to-first-token on long prompts ≈ compute. Macs and Strix Halo are much weaker here than NVIDIA GPUs.

7 · PCIe vs NVLink vs memory

PCIe 4.0 x4
7.9 GB/s
PCIe 4.0 x16
32 GB/s
PCIe 5.0 x16
63 GB/s
DDR5 dual-channel
96 GB/s
NVLink (RTX 3090 bridge)
110 GB/s
RTX 3090 VRAM
940 GB/s
GB/s

Links between devices are 10–100× slower than VRAM. That is why models are split so that only small activations cross them — and why CPU offload over PCIe hurts prompt processing badly.

8 · Multi-GPU: more capacity, not more single-user speed

With llama.cpp's default layer split, GPU 1 runs layers 1–40, then GPU 2 runs 41–80. They take turns, so one user gets roughly the speed of one GPU reading the whole model — capacity adds up, speed does not. Tensor parallelism (vLLM) splits every layer, so GPUs read in parallel, but they must synchronise twice per layer, which is costly over PCIe.

Each GPU also needs its own runtime context, and layers can't be split across GPUs in layer mode, so you lose a little capacity per GPU.

9 · CPU offloading

If the model is too big for VRAM, llama.cpp can keep some layers in system RAM. Those layers are computed by the CPU at RAM speed (≈50–100 GB/s on desktops), and they dominate the time per token. For MoE models the smart move is to offload only the experts: attention stays on the GPU and only the few active experts are read from RAM each token.

time/token ≈ GPU bytes ÷ VRAM BW + CPU bytes ÷ RAM BW

Formulas & assumptions

  • Memory = weights + KV cache × users + runtime/compute buffers (per GPU) + safety margin.
  • Weights = parameters × bits/weight ÷ 8 (or measured file size).
  • KV/token = Σ layers 2 × KV heads × head dim × bytes (MLA: latent + RoPE dims; sliding layers capped at window).
  • Decode step = max(bytes ÷ (BW × η), FLOPs ÷ (FP32 × η)) + per-layer overhead + offload + interconnect.
  • η = architecture × backend × quant-kernel × model-type factors from data/calibration.json, then corrected with measured benchmarks.
  • Ranges: HIGH ±10%, MEDIUM −20/+15%, LOW −35/+15% around the central estimate.

Full derivation: docs/FORMULAS.md in the repository.