If you want to run large language models locally with Ollama, LM Studio, llama.cpp, or ExLlamaV2, forget FPS charts. For LLMs, VRAM capacity comes first, memory bandwidth second, and raw gaming performance third. Buy too little VRAM and your model either will not load at all or it spills into system RAM and drops to 1-2 tokens per second.

I have tested 7B to 70B quants on a mix of NVIDIA, AMD, and Intel cards. The pattern is consistent: a slower GPU with more VRAM beats a faster GPU with less VRAM for local AI. A 24GB RTX 3090 will run circles around a 12GB RTX 4070 for 33B and 70B models, even though the 4070 is newer and more efficient for gaming.

What actually matters for local LLMs

VRAM determines what size model you can fit. As a rough rule with 4-bit quantization (Q4_K_M), you need about 5GB for an 8B model, 10-12GB for a 14B, 20-22GB for a 32B, and 40GB+ for a 70B. If you want full context length, Chat + RAG, or run diffusion alongside the LLM, add 20-30 percent headroom.

Software support is the second filter. NVIDIA CUDA just works on Windows and Linux with every backend: llama.cpp, Ollama, LM Studio, Oobabooga, GPT4All, Stable Diffusion, Whisper, and RVC. AMD on Linux with ROCm is now usable for LLMs, but on Windows it is still a compromise. Intel Arc has made progress with IPEX and Vulkan in llama.cpp, but you will hit missing features and slower token speeds.

Memory bandwidth controls tokens per second once the model is loaded. This is why an RTX 3090 with 936 GB/s and a 384-bit bus often outruns an RTX 4070 Ti Super for inference, despite Ada Lovelace being more efficient. For training or fine-tuning with LoRA, you also want Tensor Cores, large VRAM, and ideally NVLink or at least headroom for optimizer states.

Best NVIDIA options for most people

For most buyers, NVIDIA is still the safe answer. If you are starting fresh and want 7B to 14B models at fast speeds with room for 32B quants, a 16GB card like the RTX 4070 Ti Super is the sweet spot. It does about 60-90 tok/s on 8B Q4 and 25-35 tok/s on 14B Q4 in llama.cpp with CUDA, depending on context size. The failure mode is simple: you cannot fit 70B without heavy offloading to RAM, which kills speed.

If you want to run 33B comfortably and experiment with 70B Q4 with partial CPU offload, get 24GB. New 24GB Ada cards are expensive, which is why so many local LLM builders buy a used RTX 3090 24GB on Amazon or second-hand. A 3090 is power-hungry at 350W and runs hot, you need a real 850W+ PSU and good case airflow, but 24GB for often half the price of a 4090 is hard to beat for 32B models.

The no-compromise single-GPU pick is the RTX 4090 24GB. It is not just about 24GB, it is 24GB plus 1,008 GB/s bandwidth and vastly faster CUDA cores. Expect roughly 110-140 tok/s on 8B Q4 and usable 12-18 tok/s on 70B Q3 with offload. It is overkill if you only chat with Mistral 7B and Llama 3 8B. It makes sense if you run large context, multi-model agents, or 70B daily. Do not buy it for future-proofing alone unless you will actually use the VRAM now.

On a tight budget, the old RTX 3060 12GB remains the cheapest usable entry. It runs 7B Q4 fully in VRAM at 35-50 tok/s and can handle 13B Q4 with small context. It struggles with 14B+ and large 8k context because 12GB fills fast and bandwidth is only 360 GB/s. Buy it only if your use is light chat, coding assist, and learning Ollama. If you can stretch to a 16GB RTX 4070 Ti Super graphics card, you will feel the difference immediately in context size and multitasking.

Where AMD and Intel fit

AMD makes sense when you want maximum VRAM per dollar for inference on Linux and do not need CUDA training tools. The RX 7900 XTX with 24GB and the RX 7900 XT with 20GB are excellent value for llama.cpp with ROCm or Vulkan on Linux, often matching or beating a 4070 Ti Super in tok/s per dollar for 7B to 32B models.

The catch is Windows. Ollama and LM Studio now support AMD on Windows much better than a year ago, but you will still see slower updates, occasional broken backends after driver updates, and poor performance in ExLlamaV2, bitsandbytes, and many training scripts. If you live in Windows and just want plug-and-play, NVIDIA will save you hours. If you run Linux and mostly do inference, a RX 7900 XTX 24GB card for local LLMs is a legitimate cheaper alternative to a 4090 for 32B workloads.

Intel Arc, specifically the Arc A770 16GB, is a budget experiment. With 16GB for very little money, it can load 14B Q4 and run 7B models at usable speeds in llama.cpp Vulkan builds. Token speed is well behind NVIDIA equivalents, often 30-50 percent slower, and many Python CUDA libraries simply have no Intel path. Buy it only if your budget is strict, you tinker on Linux, and you accept troubleshooting. For everyone else, a used NVIDIA 12GB+ card is a better gift of your time.

Quick comparison by model size

Match the card to the largest model you will run weekly, not the smallest demo you tried once. Partial GPU + CPU offload works, but once more than 20-30 percent of layers sit in system RAM, speed collapses.

GPU VRAM Best fit Trade-off to know
RTX 3060 12GB 12GB 7B Q4, small 13B Slow with 8k context, no 32B
RTX 4070 Ti Super 16GB 8B-14B fast, 32B with offload Pricey per GB, 70B not practical
Intel Arc A770 16GB Budget 7B-14B on Linux Weak software support, slower tok/s
RX 7900 XT / XTX 20GB / 24GB 14B-32B on Linux ROCm quirks, weak on Windows + training
RTX 3090 24GB 32B Q4, entry 70B offload 350W, heat, used-market risk
RTX 4090 24GB Fastest single-GPU 32B-70B Very expensive, needs big PSU/case

How to buy without wasting money

Do not buy two cheap 8GB cards hoping to combine VRAM. Consumer multi-GPU pooling for LLMs is messy and most apps will only use one card. One 24GB card beats two 12GB cards for simplicity and compatibility every time, unless you specifically build for llama.cpp with tensor split across two identical cards.

Check the rest of your system. For partial offload of 70B models, 32GB to 64GB of fast DDR4/DDR5 system RAM helps a lot, and an NVMe SSD speeds model loading but does not fix tok/s. Power matters too: a 3090 or 4090 on a cheap 650W PSU will crash under sustained LLM load. Budget for PSU and cooling, not just the card.

If you are unsure between 16GB and 24GB, buy for the context you use. If you run 4k context chat only, 16GB is fine and cheaper is fine. If you run 32k context, Retrieval-Augmented Generation with large documents, or run an LLM plus Stable Diffusion at the same time, get 24GB. VRAM pressure shows up as stutter, failed loads, and forced lower quants like Q3 instead of Q5, which noticeably hurts reasoning quality.

FAQ

How much VRAM do I need for Llama 3 70B locally?

For Q4 quantization you need roughly 40-42GB for full GPU inference. With a single 24GB card you can run it with partial CPU offload at 8-15 tok/s if you have 32GB+ system RAM, but it will not be fast. For smooth 70B, most people need two 24GB cards or an Apple Silicon Mac with 64GB+ unified memory.

Is AMD good for local LLMs or should I stick with NVIDIA?

Stick with NVIDIA if you use Windows, want one-click apps like LM Studio and Oobabooga, or do LoRA training. AMD is viable if you run Linux, mainly do chat inference in Ollama or llama.cpp, and want cheaper 20-24GB. Expect more setup work and slower support for new models.

Is a used RTX 3090 still worth it for AI in 2025-2026?

Yes, if the price is right and you verify thermals, memory temps, and that it was not abused for mining without repadding. It remains the cheapest way to get 24GB plus high bandwidth. Buy from a seller with returns, stress-test VRAM immediately, and make sure your case and PSU can handle 350W sustained.

Will an Intel Arc GPU run Ollama and LM Studio?

Basic inference can work via Vulkan / IPEX paths, especially with an Intel Arc A770 16GB for local AI builds, but expect slower tokens and occasional unsupported models. Treat it as a budget learning card, not a primary LLM workstation GPU.

Related guides

Browse all Graphics Cards guides →