Your machine. Every model.
Hold your hardware and context steady. Compare what each model could feel like on this machine.
What does each model feel like on this machine?
Each point is one model on My Current PC (i7-11700K · 64GB · RTX 3080 Ti). Whiskers show the estimated speed range.
11,4 GiB usable accelerator memory · 64 GiB system RAM (assumed host; use Build a system to customize)
7 models fit entirely in accelerator memory and meet your 30 tok/s target across the estimated range.
Every model in the catalog
29 plotted · 10 without a comparable estimate
| Model | Memory needed | Estimated speed | Fit on this system |
|---|---|---|---|
| 8,9 GiB | Unavailable · not comparable | Context exceeds native limitRequested context exceeds Llama 2 7B's native context (4,096 tokens); this estimate is excluded from comparison. | |
| 3,8 GiB | 150–220 tok/s | Fits · target met | |
| 6,7 GiB | 81–120 tok/s | Fits · target met | |
| 45,3 GiB | 0.56–0.98 tok/s | Runs with RAM offload | |
| 244,8 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 67 GiB | 4.7–8.4 tok/s | Runs with RAM offload | |
| 238,8 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 22,4 GiB | 1.7–3 tok/s | Runs with RAM offload | |
| 46,6 GiB | 0.54–0.95 tok/s | Runs with RAM offload | |
| 7 GiB | 78–110 tok/s | Fits · target met | |
| 11 GiB | 47–68 tok/s | Fits · target met | |
| 22,4 GiB | 1.7–3 tok/s | Runs with RAM offload | |
| 19,8 GiB | 26–46 tok/s | Runs with RAM offload | |
| 141,3 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 19,8 GiB | 26–46 tok/s | Runs with RAM offload | |
| 286,5 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 49,2 GiB | 21–37 tok/s | Runs with RAM offload | |
| 45,3 GiB | 0.56–0.98 tok/s | Runs with RAM offload | |
| 22,4 GiB | 1.7–3 tok/s | Runs with RAM offload | |
| 397,7 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 14,7 GiB | 32–57 tok/s | Runs with RAM offloadRequested Q4_K_M; estimate uses MXFP4 (native). This is not an exact quantization match. | |
| 65,2 GiB | 10–18 tok/s | Runs with RAM offloadRequested Q4_K_M; estimate uses MXFP4 (native). This is not an exact quantization match. | |
| 4,1 GiB | 150–210 tok/s | Fits · target met | |
| 9,3 GiB | 55–79 tok/s | Fits · target met | |
| 18,5 GiB | 2.5–4.4 tok/s | Runs with RAM offload | |
| 16,4 GiB | 3.5–6.2 tok/s | Runs with RAM offload | |
| 16,4 GiB | 3.5–6.2 tok/s | Runs with RAM offload | |
| 29,4 GiB | 3.9–7 tok/s | Runs with RAM offload | |
| 65,1 GiB | 5.1–9 tok/s | Runs with RAM offload | |
| 213,6 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 592 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 11,2 GiB | 46–66 tok/s | Fits · target met | |
| 138,9 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. | |
| 18,3 GiB | 2.6–4.7 tok/s | Runs with RAM offload | |
| 22,6 GiB | 36–65 tok/s | Runs with RAM offload | |
| 20,8 GiB | 1.9–3.4 tok/s | Runs with RAM offload | |
| 16,4 GiB | 38–66 tok/s | Runs with RAM offload | |
| 19,8 GiB | 27–48 tok/s | Runs with RAM offload | |
| 168,7 GiB | Unavailable · not comparable | Does not fitPerformance data is unavailable; this model is excluded from comparison. |
These are simulated ranges, not benchmark guarantees. Both chart axes are logarithmic. The accelerator memory line is a capacity reference; runtime needs and memory placement still determine fit. Unsupported contexts and missing estimates stay out of the plot. A larger model does not necessarily mean a better answer.
Check which local AI models fit your computer
Choose a computer, then compare catalog models under one fixed workload. Required memory changes with model weights, quantization and context length; decode speed also depends on the hardware and its memory bandwidth. Results distinguish full fit, RAM offload, alternate-quant fallback and cases that cannot be compared fairly.
Questions people ask
How can I tell whether a model fits my computer?
Start with the selected model and context, then check the required-memory estimate against the computer's usable accelerator memory. The estimate accounts for model weights, context cache and runtime overhead. A result may require RAM offload or a supported alternate quantization, and context beyond a model's stated native limit is flagged. Real software use can still vary, so leave practical headroom.
Does a larger model always give better answers?
No. Parameter count alone does not measure answer quality. Training data, architecture, tuning, prompt, quantization and the task all matter, and this tool does not score model quality. Model size is useful here as one clue about memory needs; compare published evaluations or try the models on your own tasks when answer quality is the deciding factor.
Why do context length and quantization change speed?
Quantization stores model weights with fewer bits, which can reduce memory use and the bytes read during generation, with possible quality trade-offs. Longer context increases the key-value cache and can consume more memory and generation bandwidth. The calculator holds your selected settings steady while comparing catalog models; unsupported contexts or quantizations are called out rather than shown as ordinary speed results.