llm4funAI HARDWARE, EXPLAINED.
THE WORKBENCH
THE LOCAL AI SIMULATOR

See what your budget can do.

Explore the trade-offs. Adjust your budget and speed target to find your sweet spot.

85 hardware profilesOne workload. The whole picture.

The performance landscape

Every point is a setup. Select one to explore its cost, speed and memory.

SIMULATED
Generation speed · tok/s per userFull context · central estimates
NVIDIA consumerNVIDIA ProNVIDIA datacenterAppleAMD / IntelDGX Spark / ThorStrix HaloMulti-GPU / custom

↑ Faster · ← Lower cost

Central estimates, not guaranteed speeds. Select a point for its range.

● Single GPU■ Unified memory▲ Multi-GPU○ CPU offloadLarger = more AI memoryFaded = over budget
Plot coverage · 77 shown / 21 omitted

21 systems have missing or unplottable values for these axes.

Use Browse every setup below to explore the full hardware library.

Show
Speed and target matches are estimates; ranges show uncertainty, not a guarantee.

Over-budget setups stay visible so you can explore the next step. GPU prices include the configured host PC. Used prices where available; otherwise new.

YOUR HARDWARE SHORTLIST

Different strengths. A good fit.

Every pick fits your model, meets 30 tokens/s per user, and stays within budget.

The budget pickFits your plan
GPU SETUP

Single GPU · Host included

RTX 3060 12GB

10 700SEK

Used purchase estimate

Estimated generation41–50tokens/s
Inference power210W
Memory fit7.0 GiB / 11 GiB usable

4.4 GiB of spare memory

The lowest priced match. Enough memory and estimated speed for this workload, with more of your budget left over.

The power-conscious pickFits your plan
UNIFIED SYSTEM

Complete computer · Unified memory

Mac Studio M1 Max (32-core GPU) 64GB

16 700SEK

Used purchase estimate

Estimated generation33–48tokens/s
Inference power81W
Memory fit6.7 GiB / 48 GiB usable

41.3 GiB of spare memory

An estimated 81 W during inference. About 15 SEK/month at 4 hours a day. A useful option when ongoing energy use matters.

The performance pickFits your plan
GPU SETUP

Single GPU · Host included

RTX 4090 24GB

32 500SEK

Used purchase estimate

Estimated generation110–130tokens/s
Inference power410W
Memory fit7.0 GiB / 23 GiB usable

16.4 GiB of spare memory

Around 4.1× your speed target at the central estimate. A performance-focused option for more responsive generation.

A FEEL FOR THE SPEED

What does fast actually feel like?

Watch the same answer arrive at your target speed and on a system that fits your workload.

CURRENT WORKLOADQwen3 8B1 person · 8,192 context
YOUR PACE

Your target

30 tok/s

About 4.2 s for this ~125-token response.

Prompt

Write a short passage about a neighborhood garden.

Press play to see this response take shape.

Requested pace 30 tok/s

ELIGIBLE HARDWARE

RTX 4090 24GB

110–130 tok/s

About 1.0 s for this ~125-token response · 32 500 SEK

Prompt

Write a short passage about a neighborhood garden.

Press play to see this response take shape.

Central estimate 122.2 tok/s

Illustrative preview at the central estimate. About 125 tokens; excludes prompt processing. Actual tokenization and speed vary.

KEEP EXPLORING

There’s more under the hood.

A clearer starting point

Plan a local AI setup around your workload

Compare local language-model hardware by memory fit, estimated generation speed and price. GPU VRAM, Apple-style unified memory and CPU offload behave differently, so the same model and settings can produce very different trade-offs. Adjust the workload and budget to explore a practical shortlist.

Questions people ask

What does this local AI hardware calculator estimate?

It estimates whether a selected model, quantization and context can run on each listed system, then compares a range of per-user generation speeds. It also shows hardware prices where available. These are planning estimates based on hardware and software assumptions, not measured benchmarks or a promise that a particular setup will perform the same in your environment.

Is GPU VRAM the same as unified memory or system RAM?

No. A discrete GPU usually has its own VRAM, while unified-memory systems let the processor and GPU share one memory pool. Ordinary system RAM can also hold model data, but moving it to a GPU over PCIe can limit speed. The simulator treats those memory paths differently; installed RAM plus VRAM should not be read as one equally fast pool.

Why compare speed ranges instead of one tokens-per-second number?

Generation speed varies with backend, drivers, batch size, model implementation and other workload details. A range communicates that uncertainty more honestly than a single precise-looking number. The estimate describes decoding tokens per second per person; prompt processing and time to first token are separate parts of an end-to-end response. Check the assumptions and available measurements before buying hardware.