Will it run?
Fits · 25 to 38 tok/sQwen3 32B Q4_K_M · 8K · 22.0 GB of 24.0 GB
Back to calculator
Speed shown as a range, with the math openSpeed as a range, math open GPU, Mac, Strix Halo, DGX Spark Share your config by link

Local LLM Hardware: What Your GPU or Mac Can Actually Run, and How Fast

Pick your hardware and a model. See if it fits, how many tokens per second to expect, and what that feels like in use.

Will it run?
24 GB of VRAM at 936 GB/s. The value pick on the used market.
Quantization
Q4_K_M, about 4.85 bits per weight: the usual sweet spot.
Context length
Users at once
KV cache
Reserve for other local AI apps
Anything else that loads a model shares the same memory, for example on-device dictation such as Contextli's local transcription.

RTX 3090 (used) · Qwen3 32B Q4_K_M · 8K context
Fits, tight: 2.0 GB to spare
25 to 38tokens / second

Fast: feels instant for chat and code completion.

Estimate · confidence mediumceiling 45 tok/s
Weights 18.5 GB KV cache 2.0 GB Overhead 1.5 GB
How we calculate (open the math)

Weights. Parameters in billions times bits per weight, divided by 8. Q4_K_M averages about 4.85 bits, so Qwen3 32B (32.8B) needs about 18.5 GB. The exact figure is the size of the GGUF file you download.

KV cache. 2 (keys and values) times layers times KV heads times head size times 2 bytes, per token. Qwen3 32B: 2 x 64 x 8 x 128 x 2 = 256 KB per token, so 8K context is 2.0 GB. Layers and heads come from each model's config file; custom sizes are scaled from an 8B model's 128 KB per token.

Overhead. 1.5 GB for the runtime and compute buffers. Macs give the GPU about two thirds of memory up to 36 GB and three quarters above by default (community-reported; check yours with sysctl iogpu.wired_limit_mb). Sizes are binary GB; bandwidth is decimal GB/s, as vendors publish it.

Speed. Each token reads the active weights and the KV cache once, so the ceiling is bandwidth divided by those bytes. The range is 55 to 85% of the ceiling for dense models and 40 to 75% for mixture-of-experts models. On a graphics card, layers that do not fit spill to system RAM, taken as 64 GB of dual-channel DDR5 at 89.6 GB/s.

How much VRAM do I need to run a local LLM?

Short answer

Multiply the model's parameters in billions by about 0.6 GB at 4-bit quantization, 1 GB at 8-bit, or 2 GB at 16-bit, then add room for context (the KV cache) and 1-2 GB overhead. An 8B model fits in 8 GB; a 70B model at 4-bit needs about 40-48 GB.

Three numbers decide the fit: the weights, the KV cache that grows with every token of context, and a little runtime overhead. Quantization shrinks the weights; it is the first lever to pull. Below Q4 most models lose noticeable quality, so treat Q3 and Q2 as a way to try a model, not to live with it.

QuantBits / weight8B32B70BQuality
FP1616.0015.061.1131.5Reference quality
Q8_08.507.932.569.9Indistinguishable from FP16
Q6_K6.566.125.053.9Near lossless
Q5_K_M5.705.321.846.8Very good
Q4_K_M4.854.518.539.9The usual sweet spot
Q3_K_M3.903.614.932.1Noticeable losses
Q2_K3.353.112.827.5For trying a model only
Weights only, in GB. Add the KV cache and 1.5 GB overhead.

Context costs memory too

The KV cache stores keys and values for every token in the window. On Qwen3 32B that is 256 KB per token: 2.0 GB at 8K, 32.0 GB at 128K. Quantizing the cache to 8-bit roughly halves it, which is often the difference between fitting and spilling.

A rough rule of thumb is 1GB of RAM per billion parameters
Crowdfense · Home-made LLM recipe
I can run with 256k context and only uses ~117G
tarruda · Hacker News · Sep 2026

Why does memory bandwidth decide speed?

Short answer

Generating one token means reading every active weight from memory once. So the speed limit is memory bandwidth divided by the bytes read per token. A used RTX 3090 moves 936 GB/s, a Mac mini M4 about 120 GB/s, which is why the same 32B model runs several times faster on the old card.

RTX 5090
11 ms · 90 tok/s
RTX 3090
21 ms · 47 tok/s
M3 Ultra 256 GB
24 ms · 41 tok/s
M4 Max 64 GB
36 ms · 27 tok/s
DGX Spark
73 ms · 14 tok/s
Strix Halo 128 GB
78 ms · 13 tok/s
Mac mini 32 GB
166 ms · 6 tok/s
Time to read Qwen3 32B at Q4_K_M (18.5 GB) once, which is one token at best. Shorter is faster. Bandwidth from vendor spec sheets.
Load into the calculator
RTX 5090RTX 509032.0 GB1,792 GB/s90 tok/s
RTX Pro 6000RTX Pro 600096.0 GB1,792 GB/s90 tok/s
RTX 4090RTX 409024.0 GB1,008 GB/s51 tok/s
RTX 3090 (used)RTX 309024.0 GB936 GB/s47 tok/s
2x RTX 30902x RTX 309048.0 GB936 GB/s47 tok/s
Mac Studio M3 Ultra, 256 GBM3 Ultra 256 GB192 GB819 GB/s41 tok/s
Mac Studio M1 Ultra, 64 GBM1 Ultra 64 GB48.0 GB800 GB/s40 tok/s
MacBook Pro M4 Max, 64 GBM4 Max 64 GB48.0 GB546 GB/s27 tok/s
RTX 5060 Ti 16 GB5060 Ti 16 GB16.0 GB448 GB/s23 tok/s
MacBook Pro M4 Pro, 48 GBM4 Pro 48 GB36.0 GB273 GB/s14 tok/s
Mac mini M4 Pro, 64 GBM4 Pro 64 GB48.0 GB273 GB/s14 tok/s
DGX Spark, 128 GBDGX Spark115 GB273 GB/s14 tok/s
Strix Halo mini PC, 128 GBStrix Halo 128 GB96.0 GB256 GB/s13 tok/s
Mac mini M4, 32 GBMac mini 32 GB21.3 GB120 GB/s6 tok/s
CPU only, 64 GB DDR5CPU, 64 GB51.2 GB90 GB/s5 tok/s
Sorted by bandwidth. Ceiling = bandwidth / 19.9 billion bytes of weights.Download the hardware sheet (CSV, CC BY 4.0)JSON
you need to read the entire model and KV cache for every token
ein0p · Hacker News · Feb 2025
Even the M3 Max seems to be slower than my 3090 for LLMs that fit onto the 3090
coder543 · Hacker News · Jan 2024

MoE models flip the math. A 120B mixture-of-experts model reads only its 5B active parameters per token, so a big-memory, low-bandwidth box like Strix Halo or a Mac runs it at chat speed while a dense 70B crawls.

Is it worth buying hardware to run LLMs locally?

Short answer

Only if privacy, offline use or heavy daily volume matter to you. A 16 GB card runs useful small models but feels far behind a $20 subscription. Rent a cloud GPU for a weekend first, measure the model you actually want, then buy for that model's memory and bandwidth.

Some readers should close this tab and keep paying for a subscription. The ones who should buy have a clear reason: data that cannot leave the building, work on a plane, or enough volume that API bills hurt.

it seems really far off from the kind of experience even a basic $20/month subscription gets me
jasode · Hacker News · Aug 2026
the most cost-effective solution is probably to rent a GPU in the cloud
loudmax · Hacker News · Dec 2024

Best local LLM hardware at each budget

Approximate street prices, as owners report them. Each card loads its setup into the calculator so you can check the model you care about.

About $600

One 16 GB card

14B dense models at full speed, and 30B MoE models with a few layers in system RAM.

Qwen3 30B-A3B Q4_K_M45 to 78tok/s, estimate

"First level worth trying, Qwen 3.6 35B A3B with a 16 GB VRAM" jononor, Hacker News

About $1,500 to $3,000

Two used RTX 3090s, 48 GB

70B at 4-bit with short context, or 32B with long context. Plan for two 350 W cards.

Llama 3.3 70B Q4_K_M12 to 18tok/s, estimate

"To get 48GB of VRAM you can get 2x 3090s but that is $3k. A single 5090 is $4k but has 32GB" jborak, Hacker News

About $2,000

Strix Halo mini PC, 128 GB

120B MoE models at chat speed. Dense 70B fits but runs slowly.

gpt-oss-120b Q4_K_M32 to 59tok/s, estimate

"I got a Beelink GTR 9 Pro for $1980" anonym29, Hacker News

About $4,000

RTX 5090, or a DGX Spark

32B at 6-bit, fast, on the 5090. The DGX Spark trades speed for 128 GB.

Qwen3 32B Q6_K35 to 54tok/s, estimate

"A single 5090 is $4k but has 32GB" jborak, Hacker News

About $6,000 and up

Mac Studio, 128 GB or more

120B+ MoE models at 8-bit with long context, quietly. Dense 70B works at reading speed.

gpt-oss-120b Q8_059 to 110tok/s, estimate

"the M3 Ultra has 819GB/s memory bandwidth" jmyeet, Hacker News

Is a Mac good for local LLMs?

Short answer

Yes for capacity: unified memory lets a 64-128 GB Mac load models no single consumer GPU can hold. It is slower per token than a high-end NVIDIA card when the model fits in VRAM, because bandwidth is lower. Pick a Mac for big models, a GPU for speed.

Pick a Mac when

  • The model is bigger than 32 GB
  • You want silence and about 120 W under load
  • MoE models are your daily driver

Pick an NVIDIA GPU when

  • The model fits in 24 to 32 GB
  • Speed matters more than size
  • You want vLLM, CUDA tools or fine-tuning

macOS lets the GPU use only part of unified memory by default. You can raise the limit until the next reboot; leave 8 GB or more for the system.

macOS terminal · 64 GB Mac, give the GPU 56 GB
$ sudo sysctl iogpu.wired_limit_mb=57344
I get about 45-55 tokens per second with this setup
nozzlegear, serving from an M1 Ultra Mac Studio to a MacBook Air · Hacker News · Aug 2026

Worked examples

A 32B model on a used RTX 3090

The default above. It fits with 2.0 GB to spare at 8K context. At 16K the KV cache doubles to 4.0 GB, the total passes the card's 24.0 GB and layers spill into system RAM, unless you quantize the KV cache to 8-bit (22.1 GB in total).

Our estimate: 18.5 GB weights + 2.0 GB KV + 1.5 GB overhead = 22.0 GB of 24.0 GB. 25 to 38 tok/s.

70B on a 64 GB Mac

Llama 3.3 70B at Q4 needs about 44 GB, close to the 48 GB default GPU share of a 64 GB MacBook Pro. It runs, at reading speed rather than chat speed.

Our estimate: 39.9 GB weights + 2.5 GB KV + 1.5 GB overhead = 43.9 GB of 48.0 GB. 7 to 11 tok/s.

120B MoE on Strix Halo

gpt-oss-120b reads only 5.1B active parameters per token, so a 256 GB/s mini PC keeps up. Dense 70B on the same box would crawl at about 3.2 to 4.9 tokens per second.

Our estimate: 65.9 GB weights + 0.3 GB KV + 1.5 GB overhead = 67.7 GB of 96.0 GB. 32 to 59 tok/s.

A 30B MoE on a 16 GB card

It does not fit in VRAM, and that is fine: with only 3.3B active, the 3.5 GB that spills to system RAM costs far less than it would on a dense model.

Our estimate: 17.2 GB weights + 0.8 GB KV + 1.5 GB overhead = 19.5 GB of 16.0 GB. 45 to 78 tok/s.

Calibration: our estimate against owner reports

Owners have published their own speeds for these machines. Where their exact model is not one of our presets, we estimate with a preset of the same size and layout, named under each setup.

SetupOur estimateReported by owner
Mac mini M4 Pro 64 GB, a 27B dense modelEstimated with Gemma 3 27B Q4_K_M9 to 1412 to 15donmcronald, Hacker News
Mac Studio M2 Ultra 128 GB, Llama 2 70B Q6_K, llama.cppEstimated with Llama 3.3 70B Q6_K7 to 117.7vczf, Hacker News
Generation speed, tokens per second, single user, short context. Owner figures as reported in their own comments.
16GB for the weights at Q4
redox99, on a 32 GB Mac running a 27B model with 256K context · Hacker News · Aug 2026

Setups people actually run

Real rigs from their owners' own write-ups and comments. Load one to see what our math predicts for a typical model on that machine.

Under $4k rig replacing API spend
every pass of agentic cleanup over the OpenStreetMap extract I use for my geocoder cost me around $400 in API spend
Jonathon Ready · jonready.com
Qwen3 32B: 18 to 28 tok/s est.
M1 Ultra Mac Studio as a home server
I get about 45-55 tokens per second with this setup
nozzlegear · Hacker News
Qwen3 30B-A3B: 82 to 153 tok/s est.
Strix Halo on Linux
Strix Halo you can get at least 120 GB to the GPU
the_pwner224 · Hacker News
gpt-oss-120b: 32 to 59 tok/s est.
Used RTX 3090
you can get then for about $700, and they're dang good
swalsh · Hacker News
Qwen3 32B: 25 to 38 tok/s est.
128 GB box, long context
makes 128 GiB of memory feel incredibly tight
lambda · Hacker News
gpt-oss-120b: does not fit
Mac mini M4 Pro, 64 GB
it's borderline with 3.5-27B
donmcronald · Hacker News
Gemma 3 27B: 9 to 14 tok/s est.

What does it cost to run?

Short answer

Electricity is rarely the problem: a 350 W card answering 1,000 questions spends cents to tens of cents. The real cost is the hardware. Divide what you paid by what you save per answer against an API, and you get how many answers it takes to break even.

Running cost, using the speed from the calculator: 25 to 38 tok/s
Electricity per 1,000 answers $0.26
Generation time4.4 h
Energy1.55 kWh
Same tokens through the API$5.00
Break-even168,931 answers
The power consumption spikes by about 120 watts over idle
a_conservative, on an M4 Max under load · Hacker News · May 2025

Software that changes the numbers

The same model and machine can fit or fail depending on the runtime's defaults. What each one does with memory:

RuntimeWhat changes the memory math
OllamaIts default context window can be shorter than the model allows, and longer prompts get cut. Set num_ctx (or OLLAMA_CONTEXT_LENGTH) to what you need.
llama.cpp-ngl sets how many layers sit on the GPU; -ctk q8_0 -ctv q8_0 halves the KV cache.
LM StudioShows the GPU offload and context sliders per model before loading.
MLXApple's native runtime; often faster than GGUF on the same Mac.
vLLMBuilt for many users at once; reserves 90% of VRAM up front by default (--gpu-memory-utilization).
shell · llama.cpp, all layers on GPU, 8-bit KV cache, 32K context
$ llama-server -m qwen3-32b-q4_k_m.gguf -ngl 99 -c 32768 -ctk q8_0 -ctv q8_0
Terminal window where ollama run llama3 answers a question, with the model's reply printed as a numbered list and a prompt waiting for the next message
What a local model looks like in a terminal: ollama run and a chat prompt. Screenshot: MGeog2022 on Wikimedia Commons, CC BY 4.0.

Use your local model in your tools

Most runtimes serve an OpenAI-compatible API on localhost. Some editors only call cloud URLs, so they need a tunnel to reach your machine. Claude Code talks to an Anthropic-compatible endpoint, so a local model needs a translating proxy in between; install Claude Code first, then point ANTHROPIC_BASE_URL at the proxy.

setting up local llm in cursor is pain because they don't support local host
frainfreeze · Hacker News · May 2025

Questions people ask

Can I run a local LLM without a GPU?

Yes, from system RAM. Dual-channel DDR5-5600 moves about 90 GB/s, so an 8B model at Q4 runs at roughly 7 to 13 tokens per second by our estimate: readable for chat, slow for code.

How much RAM do I need for Ollama?

The model file plus the context plus 1-2 GB. An 8B model at Q4 is about a 4.5 GB file, so 8 GB of RAM is the floor and 16 GB is comfortable.

Is a used RTX 3090 still worth it?

For most people starting out, yes: 24 GB at 936 GB/s is the cheapest way to run 32B models fast. Check the fans and thermal pads before you buy.

Mac Studio or 2x RTX 3090?

Two 3090s are faster on anything that fits in 48 GB. A Mac Studio holds far bigger models, quietly, at lower speed. Decide by the largest model you need.

Does context length use VRAM?

Yes. The KV cache grows with every token. Qwen3 32B uses 256 KB per token: 2.0 GB at 8K, 32.0 GB at 128K.

What is the difference between total and active parameters?

A mixture-of-experts model stores all its parameters but uses only a few experts per token. Total parameters set the memory you need; active parameters set the speed.

Can I combine two GPUs?

Yes. Splitting layers across cards adds capacity, not speed. Tensor parallel runtimes such as vLLM can add speed, at the cost of more setup.

How fast is fast enough for coding?

About 10 tokens per second reads comfortably in chat. Coding agents write long files and loop many times, so aim for 30 or more.

Sources

  1. Vendor spec sheets for every GPU, Mac and mini PC listed (memory size and bandwidth), for example Apple M4 Pro and M4 Max and NVIDIA RTX 5060 Ti
  2. Model architectures (layers, KV heads, head size) from each model's config file on Hugging Face, for example Qwen3 32B
  3. GGUF format, Hugging Face Hub docs, and the Ollama FAQ on memory and context length
  4. Bits per weight from the llama.cpp quantize table
  5. llama.cpp performance on Apple Silicon, GitHub discussion
  6. Air-gapped LLM hardware, Pinchy
  7. Local LLM rig, Jonathon Ready
  8. Hacker News threads credited beside each quote

Changelog

  • 4 Oct: first version with presets for 15 machines and 7 models; all speeds shown as estimated ranges, checked against 2 owner reports.