llm4funAI HARDWARE, EXPLAINED.
THE WORKBENCH
MODEL EXPLORER · ONE MACHINE

Your machine. Every model.

Hold your hardware and context steady. Compare what each model could feel like on this machine.

39 model profilesSame system · same scenario
THE SPEED LANDSCAPE

What does each model feel like on this machine?

Each point is one model on My Current PC (i7-11700K · 64GB · RTX 3080 Ti). Whiskers show the estimated speed range.

11,4 GiB usable accelerator memory · 64 GiB system RAM (assumed host; use Build a system to customize)

Required memory · GiBGeneration speed · tokens/sec per person
Fits in memoryUses system RAM offloadAlternate quantizationTarget speed

7 models fit entirely in accelerator memory and meet your 30 tok/s target across the estimated range.

SELECTED MODELQwen3 8B
78–110 tok/sFits · target met

Every model in the catalog

29 plotted · 10 without a comparable estimate

ModelMemory neededEstimated speedFit on this system
8,9 GiBUnavailable · not comparableContext exceeds native limitRequested context exceeds Llama 2 7B's native context (4,096 tokens); this estimate is excluded from comparison.
3,8 GiB150–220 tok/sFits · target met
6,7 GiB81–120 tok/sFits · target met
45,3 GiB0.56–0.98 tok/sRuns with RAM offload
244,8 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
67 GiB4.7–8.4 tok/sRuns with RAM offload
238,8 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
22,4 GiB1.7–3 tok/sRuns with RAM offload
46,6 GiB0.54–0.95 tok/sRuns with RAM offload
7 GiB78–110 tok/sFits · target met
11 GiB47–68 tok/sFits · target met
22,4 GiB1.7–3 tok/sRuns with RAM offload
19,8 GiB26–46 tok/sRuns with RAM offload
141,3 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
19,8 GiB26–46 tok/sRuns with RAM offload
286,5 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
49,2 GiB21–37 tok/sRuns with RAM offload
45,3 GiB0.56–0.98 tok/sRuns with RAM offload
22,4 GiB1.7–3 tok/sRuns with RAM offload
397,7 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
14,7 GiB32–57 tok/sRuns with RAM offloadRequested Q4_K_M; estimate uses MXFP4 (native). This is not an exact quantization match.
65,2 GiB10–18 tok/sRuns with RAM offloadRequested Q4_K_M; estimate uses MXFP4 (native). This is not an exact quantization match.
4,1 GiB150–210 tok/sFits · target met
9,3 GiB55–79 tok/sFits · target met
18,5 GiB2.5–4.4 tok/sRuns with RAM offload
16,4 GiB3.5–6.2 tok/sRuns with RAM offload
16,4 GiB3.5–6.2 tok/sRuns with RAM offload
29,4 GiB3.9–7 tok/sRuns with RAM offload
65,1 GiB5.1–9 tok/sRuns with RAM offload
213,6 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
592 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
11,2 GiB46–66 tok/sFits · target met
138,9 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.
18,3 GiB2.6–4.7 tok/sRuns with RAM offload
22,6 GiB36–65 tok/sRuns with RAM offload
20,8 GiB1.9–3.4 tok/sRuns with RAM offload
16,4 GiB38–66 tok/sRuns with RAM offload
19,8 GiB27–48 tok/sRuns with RAM offload
168,7 GiBUnavailable · not comparableDoes not fitPerformance data is unavailable; this model is excluded from comparison.

These are simulated ranges, not benchmark guarantees. Both chart axes are logarithmic. The accelerator memory line is a capacity reference; runtime needs and memory placement still determine fit. Unsupported contexts and missing estimates stay out of the plot. A larger model does not necessarily mean a better answer.

Start with your computer

Check which local AI models fit your computer

Choose a computer, then compare catalog models under one fixed workload. Required memory changes with model weights, quantization and context length; decode speed also depends on the hardware and its memory bandwidth. Results distinguish full fit, RAM offload, alternate-quant fallback and cases that cannot be compared fairly.

Questions people ask

How can I tell whether a model fits my computer?

Start with the selected model and context, then check the required-memory estimate against the computer's usable accelerator memory. The estimate accounts for model weights, context cache and runtime overhead. A result may require RAM offload or a supported alternate quantization, and context beyond a model's stated native limit is flagged. Real software use can still vary, so leave practical headroom.

Does a larger model always give better answers?

No. Parameter count alone does not measure answer quality. Training data, architecture, tuning, prompt, quantization and the task all matter, and this tool does not score model quality. Model size is useful here as one clue about memory needs; compare published evaluations or try the models on your own tasks when answer quality is the deciding factor.

Why do context length and quantization change speed?

Quantization stores model weights with fewer bits, which can reduce memory use and the bytes read during generation, with possible quality trade-offs. Longer context increases the key-value cache and can consume more memory and generation bandwidth. The calculator holds your selected settings steady while comparing catalog models; unsupported contexts or quantizations are called out rather than shown as ordinary speed results.