If you are buying a GPU for deep learning in 2026, forget gaming benchmarks. What matters is VRAM capacity, CUDA support, and whether the card can run for hours at full load without throttling or crashing. NVIDIA still owns this space, not because AMD hardware is slow, but because PyTorch, TensorFlow, JAX, and most tutorials assume CUDA.

What actually matters for deep learning

VRAM is the hard limit. If your model plus batch plus optimizer states do not fit, you do not train. System RAM and NVMe swap with offloading can help for inference, but for training they are painfully slow. For fine-tuning 7B-13B LLMs with QLoRA, 16GB is the practical minimum, 24GB is comfortable, and 32GB+ is where full fine-tuning becomes realistic.

Second is software support. NVIDIA’s CUDA and cuDNN work out of the box on Windows and Linux. AMD’s ROCm has improved a lot, but it is still Linux-first, version-sensitive, and breaks with many custom kernels like FlashAttention, xFormers, and Triton on certain cards. Intel’s PyTorch XPU backend is usable for learning, but expect missing features.

Third is sustained power and cooling. A card that boosts to 2.8 GHz in games for 30 seconds may throttle after 20 minutes of training. Triple-fan 3-slot coolers survive this. Single-fan blower cards and cramped ITX cases do not, unless you undervolt and add airflow. Budget at least 750W for a 16GB midrange card and 1000W for a 4090 / 5090 class rig.

Quick comparison

These are the cards that make sense to buy new or used right now. Prices shift, but the VRAM hierarchy does not.

Card VRAM Best for Watch out for
RTX 5090 32GB 32GB GDDR7 LLM fine-tuning, large vision models, local inference Huge price, 575W power draw, needs new PSU cable
RTX 4090 24GB 24GB GDDR6X Serious training without enterprise pricing Still expensive used, large 3.5-slot cooler
RTX 3090 24GB 24GB GDDR6X Best VRAM per dollar if bought used Hot GDDR6X, high used-market risk, no DLSS 3 frame gen
RTX 4070 Super 12GB 12GB GDDR6X Learning PyTorch, small CNNs, LoRA experiments 12GB hits OOM fast on LLMs
RTX 4060 Ti 16GB 16GB GDDR6 Budget LLM inference and QLoRA on 7B models Narrow 128-bit bus, slower than 4070 for vision
RX 7900 XTX 24GB 24GB GDDR6 Large VRAM on Linux if you avoid CUDA-only tools ROCm gaps on Windows, many wheels need building from source

Best overall: 24GB to 32GB NVIDIA

If you train regularly and can afford it, get 24GB or more on NVIDIA. The RTX 4090 graphics card remains the sweet spot for most solo researchers because it offers 24GB, excellent FP16 / BF16 throughput, and mature support for FlashAttention 2, bitsandbytes 8-bit optimizers, and TensorRT-LLM. In practice it fine-tunes a 7B model with QLoRA at batch sizes that a 12GB card simply cannot load.

The RTX 5090 with 32GB is faster and lets you run 13B QLoRA or 30B+ GGUF inference with larger context, but it is not 50 percent better for 50 percent more money. It also draws up to 575W and spikes higher. Failure mode I see often: buyers put it on an old 850W PSU with a third-party 12VHPWR adapter, then get random black screens under load. If you buy this tier, buy a native ATX 3.1 PSU rated 1000W or more and check case clearance — most cards are over 330mm long.

Who should skip this tier? If you mostly follow courses, train ResNets, YOLO, BERT-size models, or do Kaggle tabular work, you will not use 24GB. The cheaper option is fine.

Best value for VRAM: used RTX 3090

The RTX 3090 is still the VRAM-per-dollar king because of its 24GB frame buffer and full NVLink support on some models for dual-GPU pooling. For Stable Diffusion XL training, DreamBooth, and 7B fine-tuning, it performs within 10-20 percent of a 4090 once you optimize batch size and gradient accumulation.

The trade-off is used-market risk. Many 3090s were mined on for two years. Ask for photos of the backplate and memory junction temps under load, not just core temps. GDDR6X on the back side runs hot, and degraded thermal pads cause throttling at 110C even if the core reads 70C. Repadding fixes it, but that means disassembling a $800 used card. Also, these cards are 350W and triple-slot — small cases and daisy-chained PCIe cables will cause crashes.

Buy used only if you are comfortable testing with OCCT VRAM test, FurMark, and a real training run in the return window. Otherwise a new 16GB card with warranty is safer.

Best budget starter: 16GB NVIDIA

For students and first projects, the RTX 4060 Ti 16GB is the most honest entry point. The GPU core is midrange and the 128-bit bus limits bandwidth for large vision transformers, but 16GB lets you actually load a quantized 7B LLM, run Stable Diffusion XL inference, and finish PyTorch tutorials without constant out-of-memory errors. That is more useful than a faster 12GB card that OOMs.

If your work is mostly computer vision and small-to-medium PyTorch jobs, the RTX 4070 Super graphics card trains faster thanks to more CUDA cores and higher memory bandwidth, but 12GB is a ceiling. You will be using gradient checkpointing, batch size 1-2, and 8-bit Adam just to fit. That is fine for learning, not for 13B experimentation.

Common failure here is expecting full fine-tuning. On 12-16GB, plan on LoRA, QLoRA, and frozen backbones. Full FP32 fine-tuning of even a 1B model can OOM if you keep optimizer states and large batches.

Where AMD and Intel fit

On paper, the RX 7900 XTX with 24GB looks like a cheap 3090. On Linux with ROCm 6.x, PyTorch works for many standard models, and inference with llama.cpp or vLLM ROCm builds is decent. The problem is the long tail: a GitHub repo with a custom CUDA extension, bitsandbytes, or a new attention kernel often has no ROCm wheel. You end up building from source or switching to CPU.

Choose AMD only if you run Linux, are comfortable with Docker ROCm images, and mainly do inference or standard training without exotic kernels. Intel Arc A770 16GB is similar: good VRAM for the price and improving PyTorch XPU support, but still a tinkering card, not a time-saving card. For coursework that grades on CUDA notebooks, neither is worth the hassle.

What to avoid

Do not buy an 8GB card for deep learning in 2026 unless it is strictly for learning syntax. You will spend more time shrinking batches than learning. Avoid dual old cards with no NVLink for LLM work — data-parallel helps for throughput, but VRAM does not pool, so two 8GB cards do not equal 16GB usable.

Also check your system: 32GB system RAM minimum, fast NVMe for datasets, and Ubuntu 22.04 / 24.04 if you want the fewest driver headaches. On Windows, use WSL2 for most PyTorch workflows. Native Windows CUDA works, but many research repos assume Linux paths.

FAQ

How much VRAM do I need for LLMs?

For QLoRA fine-tuning a 7B model, 16GB is workable and 24GB is comfortable. For 13B QLoRA, aim for 24GB minimum. Full fine-tuning needs roughly 3-4x more VRAM than inference, so most people should stick to LoRA methods under 24GB.

Is a used mining GPU okay for deep learning?

Sometimes, but test it hard. Run a memory stress test and a 1-hour training loop. Check hotspot and memory junction temperatures. If the seller will not allow returns or provide load temps, pass.

Is AMD ROCm good enough now?

Good enough for many PyTorch and inference tasks on Linux, but not equal to CUDA. If you rely on new research code with custom CUDA kernels, you will hit compatibility issues. Beginners should stick with NVIDIA to avoid debugging the stack instead of learning models.

Should I just use cloud GPUs instead?

If you train less than 10-15 hours per week or need H100s occasionally, cloud is cheaper. If you experiment daily, tune hyperparameters, or run local inference for privacy, a 16GB to 24GB local card pays for itself in a few months.

Related guides

Browse all Graphics Cards guides →