If you are buying a GPU for machine learning, ignore gaming benchmarks. What matters is VRAM capacity, compute libraries, and whether your framework will actually use the card. For most people in PyTorch, TensorFlow, or JAX, that still means NVIDIA first, AMD or Intel only in specific cases.

What actually matters for ML

VRAM is the hard limit. If your model + batch + optimizer states do not fit, you do not train. System RAM does not save you. An 8GB card can fine-tune small vision models and run inference on 7B quantized LLMs, but it will not train a 7B model from scratch or fine-tune at full precision. 12GB is the practical minimum for learning, 16GB is far more comfortable, and 24GB is where full fine-tuning with LoRA/QLoRA on 7B-13B models becomes realistic.

Second is software support. NVIDIA CUDA plus cuDNN is what almost every tutorial, GitHub repo, and cloud image assumes. PyTorch with CUDA installs in one command and just works. AMD ROCm has improved a lot on Linux, but Windows support is still limited and many custom CUDA kernels in research code will not run. Intel XPU support in PyTorch is real for Arc cards, but expect more setup and slower community help when something breaks.

Tensor Cores and precision support also count. Modern NVIDIA cards accelerate FP16, BF16, and INT8/FP8, which can double or triple training throughput versus plain FP32. Do not buy based on gaming TFLOPS. Check VRAM bus width, power draw, and whether your case and PSU can handle a 300W+ card running at 100% for hours. ML loads expose weak cooling and weak power supplies that gaming never finds.

Quick comparison

Use case VRAM to target Sensible pick type Trade-off
Learning PyTorch, small CNNs, tabular 12GB – 16GB Mid-range NVIDIA 12GB+ card Slow for LLMs, but cheap and low power
Fine-tuning LLMs with LoRA/QLoRA (7B-13B) 16GB – 24GB High-VRAM consumer NVIDIA High upfront cost, needs good PSU/cooling
Stable Diffusion / local inference 12GB – 24GB NVIDIA with strong Tensor throughput 8GB works but limits batch size and resolution
Linux-only experimentation on budget 16GB – 24GB AMD high-VRAM card More VRAM per dollar, more software friction

Best overall for most people: 24GB NVIDIA

If you can afford one card to do everything, a 24GB NVIDIA card is still the safest buy for local ML. The reason is not raw speed, it is compatibility. Hugging Face Transformers, PEFT, bitsandbytes, llama.cpp with CUDA, TensorRT, vLLM, and most Docker images expect NVIDIA. You will spend time tuning models, not drivers.

The classic example is the RTX 3090 on the used market versus a new RTX 4090-class card. A used 3090 gives you 24GB for much less money, NVLink is not needed for single-GPU work, and it handles QLoRA on 7B models, SDXL at high resolution, and large batch vision work. The failure mode is power and heat: it is a 350W card with GDDR6X that runs hot in a closed case, and used cards may have worn fans from mining. Check VRAM temperatures under load, not just core temps.

If buying new, look at current RTX 4090 listings only if you actually need the speed for daily training. It is roughly 50-60% faster than a 3090 in FP16/BF16 and far more efficient, but you pay double for VRAM you already get on the older card. For learning and occasional fine-tunes, that premium is not worth it. For someone training every day, time saved is real.

Best value for learning: 12GB to 16GB NVIDIA

You do not need 24GB to learn ML. A 12GB card trains ResNets, YOLO variants, BERT-base fine-tunes, and small diffusion experiments without issue. Where it fails is LLM fine-tuning at full precision and large-image Stable Diffusion batches. You will hit CUDA out of memory errors if you try to load a 13B model in FP16, which needs about 26GB before optimizer overhead.

This is where a 12GB card like the RTX 3060 12GB still makes sense despite its age. It has more VRAM than the faster 8GB cards that replaced it, and for ML, 12GB slow beats 8GB fast. The 4060 Ti 16GB is the cleaner new buy: lower power, better BF16 efficiency, and that extra 4GB matters for LoRA on 7B quantized models. Browse RTX 4060 Ti 16GB options if you want new with a warranty and a 550-650W PSU is all you have.

Who should buy here: students, software engineers adding ML skills, and anyone who will also use cloud GPUs for big runs. Train small locally, rent A100/H100 by the hour for the final run. That hybrid approach is cheaper than buying 24GB you use twice a month.

AMD and Intel: when they make sense

AMD gives you more VRAM per dollar. A 24GB Radeon can be excellent for inference with llama.cpp, ONNX, or DirectML workflows on Windows, and for PyTorch on Linux with ROCm 6.x support. If your workload is mostly inference in FP16 or INT8 and you run Linux, it is a legitimate way to get 24GB cheap.

The failure mode is missing kernels. Bitsandbytes, xFormers, FlashAttention, and many research repos ship CUDA-only code. You can often find ROCm forks, but you will be debugging instead of training. I would only recommend AMD for ML if you are comfortable with Docker, Linux drivers, and reading GitHub issues. It is not a good first card if you just want tutorials to work.

Intel Arc, like the A770 16GB, is even more niche. Price per GB is good, and Intel has done real work on PyTorch XPU support. For learning basics and running small models, it works. For anything with custom CUDA extensions, assume it will not run. Buy Intel to experiment, not as your only ML GPU. Check current Intel Arc A770 16GB prices if you already have an NVIDIA machine and want a cheap second box for inference tests.

What to avoid

Do not buy an 8GB card for LLM work thinking you will upgrade later. You will spend more time fighting offloading to system RAM, extreme quantization, and batch size 1 than learning. 8GB is fine for learning CNNs and running Stable Diffusion 1.5 at 512×512, but be honest about that limit.

Do not buy a workstation or data-center card unless you know why you need it. Used Tesla cards without display outputs often lack cooling, need server airflow, and have poor FP16 support on older architectures. ECC VRAM and certified drivers help for long multi-day runs, but for home use a consumer card with 24GB is faster per dollar.

Also check your system: you need at least 32GB system RAM if you work with large datasets, an NVMe SSD for checkpointing, and a power supply with 150-200W headroom above the GPU rating. ML sustains full load for hours. Undervolt slightly if thermals are an issue, it rarely hurts training throughput more than 3-5% but can drop VRAM temps 8-10C.

FAQ

How much VRAM do I need for LLMs?

For inference only, 12GB handles 7B models in 4-bit quantization comfortably. For QLoRA fine-tuning a 7B model, 16GB is workable and 24GB is comfortable. Full fine-tuning without quantization needs far more and is usually better done in the cloud.

Is AMD usable for PyTorch now?

Yes on Linux for many standard models with ROCm, no for a lot of cutting-edge research code that depends on CUDA-only extensions. If you follow tutorials exactly, expect friction. If you run standard training loops, it is much better than two years ago.

Should I buy used for ML?

Yes if you test it. A used 24GB card is the cheapest way to get serious VRAM. Stress-test VRAM with a memory test and a 30-minute training loop, check hotspot temps, and buy from a seller with returns. Avoid cards with modified BIOS or missing screws.

Is cloud GPU better than buying?

If you train less than 20-30 hours per month on large models, renting is usually cheaper. If you experiment daily, a local 12GB-24GB card pays for itself in 4-6 months versus on-demand prices and gives you instant iteration without setup overhead.

Related guides

Browse all Graphics Cards guides →