Will it run this LLM?
Pick a model, quantization, and context length to see exactly how much VRAM it needs — with a full weights/KV/activation breakdown, fit verdicts, speed estimates, multi-GPU and fine-tuning modes, and the best-value card if you're shopping.
Llama 3.1 8B @ Q4 needs ~6.6 GB
32.0 GB available · 21% usedNVIDIA RTX 5090 32GB (32.0 GB) vs ~6.6 GB needed.
249
tok/s (est.)
2.7 s
TTFT (est.)
249
throughput (est.)
| GPU | Memory | Verdict | Est. speed |
|---|---|---|---|
| NVIDIA RTX 5090 32GBreview | 32 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3090 Ti 24GB | 24 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 5080 16GBreview | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 7900 XTX 24GB | 24 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3090 24GB | 24 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3080 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3080 Ti 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 5070 Ti 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 7900 XT 20GB | 20 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3080 10GB | 10 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4080 Super | 16 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4080 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4070 Ti Super | 16 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 5070 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4090 24GB | 24 GB | Fits | Blazing (>50 tok/s)● |
| AMD RX 9070 XT 16GBreview | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 9070 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| Mac Studio M4 Max 128GB | 96 GB (unified) | Fits | Blazing (>50 tok/s) |
| AMD RX 7800 XT 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 2080 Ti 11GB | 11 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3070 Ti 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 6950 XT 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 7900 GRE 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| Intel Arc A770 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 6800 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 6800 XT 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 6900 XT 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| Intel Arc A580 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| Intel Arc A750 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| Intel Arc A770 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4070 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4070 Super 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 4070 Ti 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 2080 Super 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| Intel Arc B580 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 2060 Super 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 2070 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 2070 Super 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 2080 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3060 Ti 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3070 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 5060 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 5060 Ti 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 5060 Ti 16GB | 16 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 5700 XT 8GB | 8 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 9070 GRE 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 6750 XT 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 7700 XT 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| AMD RX 6700 XT 12GB | 12 GB | Fits | Blazing (>50 tok/s) |
| Intel Arc B570 10GB | 10 GB | Fits | Blazing (>50 tok/s) |
| NVIDIA RTX 3060 12GB | 12 GB | Fits | Fast (25–50 tok/s) |
| AMD RX 9060 XT 8GB | 8 GB | Fits | Fast (25–50 tok/s) |
| AMD RX 9060 XT 16GB | 16 GB | Fits | Fast (25–50 tok/s) |
| Mac Mini M4 Pro 64GB | 48 GB (unified) | Fits | Fast (25–50 tok/s) |
| Ryzen AI Max+ 395 (Strix Halo) 128GB | 96 GB (unified) | Fits | Fast (25–50 tok/s) |
| NVIDIA RTX 4060 Ti 16GB | 16 GB | Fits | Fast (25–50 tok/s) |
| NVIDIA RTX 4060 Ti 8GB | 8 GB | Fits | Fast (25–50 tok/s) |
| AMD RX 6650 XT 8GB | 8 GB | Fits | Fast (25–50 tok/s) |
| AMD RX 7600 8GB | 8 GB | Fits | Fast (25–50 tok/s) |
| AMD RX 7600 XT 16GB | 16 GB | Fits | Fast (25–50 tok/s) |
| NVIDIA RTX 4060 8GB | 8 GB | Fits | Fast (25–50 tok/s) |
| AMD RX 6600 8GB | 8 GB | Fits | Fast (25–50 tok/s) |
| NVIDIA RTX 2060 6GB | 6 GB | Fits w/ offload | Slow (3–10 tok/s) |
| Intel Arc A380 6GB | 6 GB | Fits w/ offload | Slow (3–10 tok/s) |
NVIDIA RTX 3060 12GB
$289 · Price as of Jul 14, 2026
VRAM figures are computed from each model's architecture (GQA-aware KV cache). Speed numbers (tok/s, TTFT, throughput) are bandwidth-derived estimates — ● marks measured decode rates. Unified-memory machines list usable GPU-allocatable memory, not total RAM.
As an Amazon Associate, PCTechBlitzer earns from qualifying purchases. Full disclosure →
Memory required = model weights at the chosen quantization plus a GQA-aware KV cache that scales with context length, batch, and concurrent users, plus activation and framework overhead. Fine-tuning mode adds gradient and optimizer state (Full ≈ 16 bytes/parameter; LoRA and QLoRA train a small adapter over a frozen base). A GPU fits when its memory covers the total; fits with offload when it's close and you accept CPU spill.
How much VRAM does a 70B model need?
At Q4 quantization, Llama 3.1 70B needs about 44 GB of memory (weights plus a GQA-aware KV cache at 8K context); at Q8 that grows to about 77 GB. No single consumer GPU fits it fully at Q4 — you need a unified-memory machine, multiple GPUs, or partial CPU offload.
Can an RTX 4090 run Llama 70B?
No — Llama 3.1 70B at Q4 needs about 44 GB, far beyond the NVIDIA RTX 4090 24GB's 24 GB, even allowing for CPU offload.
What quantization should I use?
Q4 is the sweet spot for local inference: roughly a quarter of the FP16 memory footprint with minimal quality loss for most tasks. Use Q8 when you have VRAM to spare and want quality headroom; FP16 is rarely worth it locally.
Does longer context need more VRAM?
Yes — the KV cache grows linearly with context length. Llama 3.1 70B at Q4 needs about 44 GB at 8K context, and more as you extend it. Quantizing the KV cache to Q8 or Q4, or fewer concurrent sequences, brings it back down.
Can I run LLMs on a Mac?
Yes. Apple Silicon's unified memory lets the GPU address most of system RAM — a Mac Studio M4 Max 128GB exposes ~96 GB to models, enough to run Llama 3.1 70B at Q4 fully in memory at usable speeds. Bandwidth, not capacity, is the constraint versus a discrete GPU.
What does 'fits with offload' mean?
The model is bigger than your GPU memory, but close enough that llama.cpp/Ollama can keep most layers on the GPU and run the rest on the CPU. It works, but expect the speed tier to drop to Slow regardless of how fast the card is.
How accurate are the numbers?
VRAM is computed from each model's real architecture — layer count, hidden size, and GQA KV-head count — so the weights and KV-cache figures are precise for the settings you pick. Tokens-per-second, time-to-first-token, and throughput are bandwidth-derived estimates bucketed into coarse tiers; measured decode rates from our reviews override them and are marked ●.