If you want to know how fast a local large language model will run on your PC, the short answer is this: generation speed tracks memory bandwidth almost one-to-one. In published llama.cpp results, an RTX 5090 produces roughly 150 to 200 tokens per second on small and mixture-of-experts models, a 16GB mid-range card lands between 40 and 75 tokens per second on a 14B model, unified-memory boxes like the DGX Spark and Strix Halo sit around 50 to 60 tokens per second on gpt-oss-120b, and a desktop CPU on its own manages about 11 to 12 tokens per second on an 8B model. The rest of this page breaks those figures down by hardware tier, explains where each number comes from, and shows you how to read benchmarks without getting fooled.
Top 3 picks at a glance
Verdict at a glance
- Fastest single card for models that fit: GeForce RTX 5090 (32GB, 1,792 GB/s). Roughly 1.5 to 1.8 times an RTX 4090 in single-user generation.
- Best tokens-per-second for the money: RTX 5070 Ti (16GB, 896 GB/s). It runs 14B-class models at 58 to 73 tokens per second in the sources below.
- Best value for 24GB: a used RTX 3090. Slower than newer cards on prompt processing, but its 24GB fits 27B to 32B models at 4-bit.
- Best for 100B+ models on one box: DGX Spark, Strix Halo mini PCs or a Mac Studio. Slower per token than a big GPU, but they can hold models a 32GB card cannot.
- Fastest unified memory right now: Mac Studio with M5 Ultra, thanks to 1.2 TB/s of bandwidth.
- Skip: CPU-only inference for anything bigger than about 14B unless you have a Threadripper-class memory system.
Why tokens per second depends on bandwidth
A local LLM does two different jobs. First it reads your prompt, which is called prefill or prompt processing (pp). Then it writes the answer one token at a time, which is called decode or token generation (tg). These two phases stress hardware in different ways.
Prefill is compute-heavy. The model processes many tokens in parallel, so tensor cores, matrix units and raw TFLOPS matter. This is why an NVIDIA card with modern tensor cores reads a long prompt several times faster than an Apple chip or an AMD integrated GPU with similar memory bandwidth.
Decode is memory-bound. For every new token, the hardware has to stream the active model weights out of memory. If a 4-bit 8B model weighs about 5GB, generating one token means reading roughly 5GB. A card with 1,000 GB/s of bandwidth can, in theory, do that about 200 times a second. Real software reaches a fraction of that ceiling, but the ranking holds. That is why the RTX 5090’s bandwidth advantage over the RTX 4090, about 1.78 times, lines up so neatly with its measured advantage in generation speed.
Mixture-of-experts (MoE) models change the math in your favor. A model like gpt-oss-120b or a Qwen 30B-A3B has a large total parameter count but only activates a small slice per token. You still need enough memory to hold the whole model, but each token only reads the active experts, so generation is much faster than a dense model of the same size.
Benchmark table: tokens per second by hardware tier
The table below collects published figures. Numbers come from different llama.cpp builds, backends and settings, so treat them as a guide to the tier, not a lab-grade head-to-head. Each row names its source.
| Tier | Hardware | Memory / bandwidth | Model and quant | Generation (tok/s) | Source |
|---|---|---|---|---|---|
| Flagship GPU | RTX 5090 | 32GB / 1,792 GB/s | 8B Q8_0, tg128 | 159 | OpenBenchmarking.org llama.cpp results |
| Flagship GPU | RTX 5090 | 32GB / 1,792 GB/s | Qwen3.5-35B-A3B Q4_K_XL (MoE) | 194 | llama.cpp GitHub discussion #19890 |
| Previous flagship | RTX 4090 | 24GB / 1,008 GB/s | 8B Q8_0, tg128 | 101 | OpenBenchmarking.org |
| Used 24GB | RTX 3090 | 24GB / 936 GB/s | 8B Q8_0, tg128 | 91 | OpenBenchmarking.org |
| Pro 32GB | Radeon AI PRO R9700 | 32GB / 640 GB/s | Qwen3.5-35B-A3B Q4_K_XL (MoE) | 127 | llama.cpp GitHub discussion #19890 |
| Mid-range 16GB | RTX 5070 Ti | 16GB / 896 GB/s | Qwen2.5 14B Q4_K_M | 73.2 | ComputingForGeeks |
| Mid-range 16GB | RTX 5070 Ti | 16GB / 896 GB/s | Qwen3 14B, 16K context | 58.0 | Hardware Corner GPU ranking |
| Mid-range 16GB | RX 9070 XT | 16GB / 640 GB/s | Qwen3 14B Q4_K_M (HIP) | 54.3 | llama.cpp GitHub issue tracker |
| Entry 16GB | RTX 5060 Ti 16GB | 16GB / 448 GB/s | Qwen2.5 14B Q4_K_M | 42.3 | ComputingForGeeks |
| Entry 16GB | RTX 5060 Ti 16GB | 16GB / 448 GB/s | Qwen3 14B, 16K context | 32.9 | Hardware Corner GPU ranking |
| Unified memory | DGX Spark (GB10) | 128GB / 273 GB/s | gpt-oss-120b MXFP4 | 58.7 | llama.cpp GitHub discussion #16578 (Feb 2026 build) |
| Unified memory | Ryzen AI Max+ 395 (Strix Halo) | 128GB / 256 GB/s | gpt-oss-120b | 49 to 54.5 | Community llama.cpp results (ROCm and Vulkan) compiled by AIMultiple |
| CPU only | Ryzen 9 9950X | Dual-channel DDR5 | 8B Q4_K_M | 11 to 12 | LocalScore results cited by PopularAI |
Two patterns jump out. First, cards with similar bandwidth land close together even when their compute differs a lot. Second, the unified-memory machines run a 120B-parameter MoE model at speeds a 16GB card reaches only on a 14B dense model. That is the trade-off in one line: capacity versus speed.
Tier 1: Flagship GPUs (RTX 5090)
The RTX 5090 is the fastest consumer card you can put in a desktop for local inference. Its 32GB of GDDR7 and 1,792 GB/s of bandwidth put it in a class of its own for anything that fits. In the community llama-bench comparison in llama.cpp discussion #19890, it generated 194 tokens per second on a 35B MoE model at 4-bit and read prompts at about 7,000 tokens per second. In LMSYS’s DGX Spark review, the 5090 ran gpt-oss-20b in Ollama at 8,519 tokens per second prefill and 205 tokens per second decode.
The rule of thumb from the sources above: for single-user llama.cpp, the 5090 is about 1.5 to 1.8 times faster than the 4090. The extra 8GB over a 24GB card matters most for 32B models at higher-quality quants, long context windows, and 70B models squeezed to 2-bit or 3-bit.
Reasons to pick: fastest generation and prefill of any consumer card; 32GB fits 32B dense models with room for context; mature CUDA support in every inference engine.
Reasons to skip: 575W board power; street pricing has run far above launch MSRP during the 2026 memory shortage; a 70B model at 4-bit (about 40GB) still will not fit.
The catch: once a model spills out of VRAM into system RAM, speed falls off a cliff, so 32GB is a hard ceiling, not a soft one.
Tier 2: 24GB cards (RTX 4090 and used RTX 3090)
For years, 24GB has been the sweet spot for hobbyists, and it still is if you can find the cards. OpenBenchmarking’s controlled llama.cpp run put the RTX 4090 at 101 tokens per second and the RTX 3090 at 91 on an 8B model at Q8_0. Hardware Corner notes that the newer 4090 and 5090 are about 2.6 to 2.7 times faster than the 3090 at prompt processing, so the 3090’s weakness shows up most with long documents and coding agents that send big prompts.
At 24GB you can run Qwen-class 27B to 32B dense models at 4-bit with modest context. A 70B model at 4-bit does not fit on a single 24GB card; you either drop to 2-bit or offload layers to the CPU, which costs a lot of speed.
Reasons to pick: 24GB at a lower tier than a 5090; the used 3090 remains the cheapest route to that capacity; two 3090s give 48GB for 70B models.
Reasons to skip: the 4090 is out of production and hard to find; used 3090s carry wear and warranty risk; high power draw for the speed you get.
The catch: a used 3090’s slower prefill can make long-context work feel sluggish even when its generation speed looks fine on paper.
Tier 3: 16GB mid-range cards (RTX 5070 Ti, RX 9070 XT, RTX 5060 Ti 16GB)
This is where most gaming PCs sit, and it is perfectly usable for 7B to 14B models and for gpt-oss-20b, which OpenAI designed to run within 16GB. The spread inside this tier is large, and it follows bandwidth.
ComputingForGeeks measured Qwen2.5 14B at Q4_K_M at 73.2 tokens per second on the RTX 5070 Ti versus 42.3 on the RTX 5060 Ti 16GB, a 1.73 times gap, and found the 5070 Ti generally 1.66 to 1.85 times faster across models. Hardware Corner’s ranking, run at a 16K context, shows the same order with lower absolute numbers: 58 versus 33 tokens per second on Qwen3 14B.
AMD’s RX 9070 XT sits between them. A llama.cpp GitHub issue records 54.3 tokens per second on Qwen3 14B at Q4_K_M using the HIP backend, and Vache Sarkissian’s comparison measured 60.1 tokens per second with Vulkan on a similar 14B coder model. The same GitHub issue documents a Vulkan slowdown on Windows with some model shapes, so backend choice matters a lot on Radeon.
Reasons to pick: you already own one for gaming; 14B models run at very comfortable speeds; the 5070 Ti delivers more tokens per second per unit of cost than the 5060 Ti in ComputingForGeeks’ math.
Reasons to skip: 16GB rules out 27B-plus dense models at good quality; long contexts eat VRAM fast.
The catch: the 5060 Ti 16GB has the capacity of the 5070 Ti but half the bandwidth, so it runs the same models at roughly 55 to 60 percent of the speed.
Tier 4: 32GB pro cards (Radeon AI PRO R9700)
AMD’s Radeon AI PRO R9700 is the budget way into 32GB. It uses the same Navi 48 silicon family as the RX 9070 XT with doubled memory and a 300W limit. In llama.cpp discussion #19890 it generated 127.4 tokens per second on Qwen3.5-35B-A3B versus 194.0 for the 5090, so the 5090 was 1.52 times faster in generation. Prefill is where the gap widens: 2.6 times at short context and 3.4 times at 32K, according to the same post.
The R9700’s real trick is pairing. Two cards give 64GB, which holds models that simply do not fit on one 5090. One owner’s write-up describes large gains after building llama.cpp specifically for the gfx1201 target and splitting the model across both cards, which shows how much the ROCm and Vulkan setup still matters.
Reasons to pick: 32GB at a much lower tier than a 5090; low power; two cards reach 64GB.
Reasons to skip: slower prefill; software tuning is more involved than CUDA.
The catch: the tester in that discussion pointed out that 127 tokens per second is far above the 30 to 40 tokens per second where most people stop noticing, so the 5090’s lead may not matter for chat.
Tier 5: Unified-memory systems (DGX Spark, Strix Halo, Mac Studio)
These machines trade bandwidth for capacity. Instead of fast GDDR7 on a card, they share a large pool of LPDDR5X between CPU and GPU. That lets them hold 100B-plus parameter models that would need several discrete cards.
DGX Spark. NVIDIA’s GB10 box has 128GB of LPDDR5X at 273 GB/s. In llama.cpp maintainer results (discussion #16578), its gpt-oss-120b performance improved from 38.6 tokens per second at launch to 58.7 tokens per second on a February 2026 build, with prefill rising to about 2,444 tokens per second. At 32K tokens of context, generation dropped to about 42.8 tokens per second. Its prefill is the strongest of this group because of its Blackwell tensor cores.
Strix Halo. The Ryzen AI Max+ 395 pairs 16 Zen 5 cores with a 40-CU Radeon 8060S and up to 128GB of LPDDR5X-8000 on a 256-bit bus (256 GB/s). Community results put gpt-oss-120b generation at roughly 49 tokens per second on ROCm and 54.5 on Vulkan, close to the Spark, but prefill is several times slower. AMD’s 2026 Ryzen AI Max 400 refresh, nicknamed Gorgon Halo, raises the ceiling to 192GB.
Mac Studio. Apple’s new Mac Studio, announced in August 2026, offers the M5 Max at up to 614 GB/s and the M5 Ultra at 1.2 TB/s with up to 512GB of memory. MacStories reported a median of 108 tokens per second on a MoE model with MLX on the M5 Ultra, 54 percent more than the M3 Ultra, with prompt processing between about 2,000 and 2,800 tokens per second. Ars Technica reported just over 50 tokens per second on a 27B dense model at 4-bit in LM Studio.
Reasons to pick: run 70B to 400B-class models on one quiet box; low power next to multi-GPU rigs.
Reasons to skip: dense 70B models still crawl on 256 to 273 GB/s systems; LMSYS measured just 2.7 tokens per second decode for Llama 3.1 70B at FP8 on the Spark with SGLang.
The catch: these boxes shine on MoE models; on big dense models, bandwidth, not capacity, decides the experience.
Tier 6: CPU-only inference
Every PC can run a small model on the CPU alone, but expectations need to be low. LocalScore results cited by PopularAI put the Ryzen 9 9950X at about 11 to 12 tokens per second on an 8B model. A dual-channel DDR5-6400 desktop has a theoretical ceiling of about 102 GB/s, roughly a tenth of an RTX 4090. RAM speed helps: public results show Mistral 7B rising from 9.66 to 11.34 tokens per second going from DDR5-4800 to DDR5-6000.
Threadripper 9000 doubles the channel count to four, for about 205 GB/s theoretical. Puget Systems found the Threadripper 9980X fastest at both prompt processing and generation among the chips in its review, while Phoronix showed it topping larger-model llama.cpp runs.
The catch: CPU-only is fine for a 3B to 8B assistant, but anything larger feels slow unless you pair the CPU with a GPU and use MoE expert offloading.
How to read local LLM benchmarks
Benchmark charts for local AI are messier than gaming charts. Before trusting any number, check these points:
- Which phase is measured? A “tokens per second” figure might be prefill (hundreds to thousands) or generation (tens to low hundreds). Mixing them up makes a slow box look fast.
- Which model and quantization? An 8B model at Q4 is a tiny workload next to a 32B model at Q8. MoE models generate far faster than dense models of the same total size.
- How much context? Speed drops as the context fills. The DGX Spark lost about 27 percent of its gpt-oss-120b generation speed going from zero to 32K tokens of context in the llama.cpp results.
- Which backend and build? CUDA, ROCm, Vulkan, Metal and MLX can differ by double-digit percentages on the same hardware. The Spark’s own numbers improved by about half in four months of software updates.
- Single user or batched? vLLM and SGLang throughput figures that add up many parallel requests are not comparable with single-user llama.cpp speed.
- Measured or estimated? Several sites publish figures calculated from bandwidth rather than measured. They are useful as a ceiling, not as results.
Run your own benchmark in five minutes
The fairest number is the one you produce on your own machine with your own models. llama.cpp ships a tool called llama-bench for exactly this.
- Install a recent llama.cpp release for your backend (CUDA for NVIDIA, ROCm or Vulkan for AMD, Metal on Mac).
- Download a GGUF model you actually plan to use, for example an 8B or 14B model at Q4_K_M.
- Run llama-bench with the model path. By default it reports pp512 (prefill on a 512-token prompt) and tg128 (generating 128 tokens).
- Add a longer prompt test, for example 4,096 or 16,384 tokens, to see how your setup behaves with documents or code.
- Repeat after driver or llama.cpp updates. Gains of 10 to 50 percent from software alone are common on newer hardware.
How fast is fast enough?
Average reading speed works out to roughly 5 to 8 tokens per second. Most people find chat comfortable from about 15 to 20 tokens per second and stop noticing improvements somewhere around 30 to 40, which is the threshold the R9700 tester referenced in llama.cpp discussion #19890. Coding agents and document summaries are different: they send large prompts, so prefill speed matters more than generation speed. If you mainly chat, prioritize capacity and a decent generation rate. If you run agents over large codebases, prioritize prefill, which favors NVIDIA cards and the DGX Spark.
Common mistakes
- Buying for the biggest model you have heard of. Most people end up using 8B to 32B models daily. Buy for those, then decide whether 70B-plus is worth a second machine.
- Ignoring context memory. The KV cache grows with context. A model that fits at 4K context may not fit at 32K.
- Letting the model spill into system RAM by accident. One layer too many offloaded to the CPU can cut speed several times. Watch the loader output.
- Judging Radeon cards by Windows Vulkan results alone. Try ROCm and Vulkan on Linux before writing off an AMD card.
- Comparing launch-day reviews with current numbers. ServeTheHome’s launch review recorded 14.5 tokens per second on gpt-oss-120b on the Spark; later llama.cpp builds reached about four times that.
Products Mentioned in This Guide
These are specific listings for the hardware tiers benchmarked above, from the flagship card down to a unified-memory mini PC.
ASUS TUF Gaming GeForce RTX 5090 32GB OC Edition
As an Amazon Associate we earn from qualifying purchases at no extra cost to you.
Tier 1: the fastest single card for models that fit, with 32GB and 1,792 GB/s of bandwidth, about 1.5 to 1.8 times an RTX 4090 in single-user generation.
PNY GeForce RTX 5070 Ti OC Triple-Fan
Our best tokens-per-second for the money: 16GB at 896 GB/s, running 14B-class models at 58 to 73 tokens per second in the cited results.
ASRock Radeon RX 9070 XT Steel Legend 16GB
The AMD card in the 16GB tier, recorded at 54.3 tokens per second on Qwen3 14B with HIP. Compare ROCm and Vulkan builds before judging it.
ASUS Dual GeForce RTX 5060 Ti 16GB OC Edition
The entry 16GB option: the same capacity as the 5070 Ti but half the bandwidth, so it runs the same models at roughly 55 to 60 percent of the speed.
ASRock Radeon AI PRO R9700 Creator 32GB
Tier 4: the budget way into 32GB, generating 127 tokens per second on a 35B MoE model, with two cards reaching 64GB.
BOSGAME M5 AI Mini PC, Ryzen AI Max+ 395, 128GB LPDDR5X
A Strix Halo unified-memory box from Tier 5, which community results put at roughly 49 to 54.5 tokens per second on gpt-oss-120b, a model no 32GB card can hold.
How we compared
We compared these using manufacturer specifications, independent lab measurements and community benchmark results from the sources named above, and owner reports. We did not bench-test these units ourselves. Where sources disagreed, we reported a range and named both. Where a site appeared to publish estimates instead of measurements, we left the figure out.
Sources
- OpenBenchmarking.org, llama.cpp test profile results
- llama.cpp GitHub discussion #16578, performance of llama.cpp on NVIDIA DGX Spark
- llama.cpp GitHub discussion #19890, RTX 5090 vs Radeon AI PRO R9700 llama-bench results
- llama.cpp GitHub issue tracker, RX 9070 XT Vulkan and HIP results
- Hardware Corner, GPU ranking for local LLMs
- ComputingForGeeks, RTX 5060 Ti vs RTX 5070 Ti for local AI
- Vache Sarkissian, Vulkan vs ROCm on RDNA 4
- LMSYS Org, NVIDIA DGX Spark in-depth review
- AIMultiple, DGX Spark alternatives
- MacStories, M5 Ultra Mac Studio review
- Ars Technica, Mac Studio M5 Ultra coverage
- Apple Newsroom, Mac Studio with M5 Max and M5 Ultra announcement
- PopularAI, best CPU for running local LLMs and RAM speed analysis
- Puget Systems, Threadripper 9000 review
- Phoronix, Threadripper 9980X and 9970X Linux benchmarks
- ServeTheHome, DGX Spark launch review
Frequently Asked Questions
What is a good tokens per second for a local LLM?
For chat, about 15 to 20 tokens per second feels comfortable and 30 to 40 feels instant for most people. Below about 8 tokens per second, the text appears slower than you read.
Why does my GPU show thousands of tokens per second in one test and 50 in another?
The thousands figure is almost certainly prompt processing, which runs in parallel. The 50 figure is generation, which happens one token at a time and is limited by memory bandwidth.
Is the RTX 5090 worth it over the RTX 4090 for local AI?
If you can find both, the 5090 is about 1.5 to 1.8 times faster in single-user generation according to the sources above and adds 8GB of VRAM. Whether that justifies the price difference depends on whether you need 32GB.
Can I run a 70B model on a gaming PC?
Only with compromises. A 70B model at 4-bit needs roughly 40GB, so on a single 24GB or 32GB card you must drop to 2-bit or 3-bit, or offload layers to system RAM and accept much slower speeds. Two 24GB cards or a unified-memory machine handle it more gracefully.
Is a Mac faster than a PC for local LLMs?
For large models, a high-end Mac Studio often generates faster than unified-memory PCs because of its higher bandwidth. A discrete NVIDIA card is still faster for any model that fits in its VRAM, especially at prompt processing.
Does RAM speed matter for local AI?
Yes, when the CPU does any of the work. Public results show 17 to 18 percent gains going from DDR5-4800 to DDR5-6000 for CPU inference, and MoE expert offloading is even more sensitive to memory bandwidth.
Which software gives the best speed?
llama.cpp is the common baseline across all hardware. On Macs, MLX is often faster. On Radeon, compare ROCm and Vulkan builds. For serving many users, vLLM and SGLang scale better than llama.cpp.
How often do these numbers change?
Often. Software updates have raised results on new hardware by large margins within months, so re-check benchmarks from the last few months before buying.





