QUICK GUIDE ยท GPU BUYING
AI GPU or gaming GPU: the same purchase?
A fast gaming card can be limited in AI by VRAM or software. Conversely, a card useful for a local model may offer poor value for your display.
SHORT ANSWERFor local AI, verify software support first, then required VRAM. For gaming, start with resolution, refresh rate and games.
Choose an AI GPU from the workload backwards
Write down the model, quantization, context length, simultaneous users and software before shopping for VRAM. A card that loads the weights may still run out of memory during a long conversation. A larger card is not useful if the required backend cannot use it on your operating system.
| Task | Record before purchase | Acceptance test |
|---|
| Local chat | Model digest, quantization, context and concurrency | No unexpected CPU offload at the required context; record latency |
|---|
| Image generation | Checkpoint, resolution, batch, VAE and extensions | The complete workflow finishes repeatedly, including upscaling |
|---|
| GPU rental | Accepted GPU/backend, workload demand, wall power and fees | Observed paid hours and net costs, not a gaming FPS score |
|---|
The weights are only the first line of the memory budget
As a rough lower-bound calculation, 8 billion parameters at 4 bits occupy 4 billion bytes; 32 billion at 4 bits occupy 16 billion bytes. Quantization metadata, runtime buffers, KV cache and other allocations add to that. A 16 GB GPU therefore cannot be promised to fit every 32B model at Q4. Record the actual memory peak with the desired context before choosing a capacity tier.
A repeatable test is more useful than a blanket NVIDIA-versus-AMD verdict
For Ollama, record the application and driver versions, exact model digest and context setting. Run the same prompt after a warm-up, then a long-context prompt; inspect the GPU/CPU allocation with ollama ps. Report prompt processing separately from output generation and repeat the run. Consult the supported-hardware list for the exact OS and backend; compatibility changes over time. The observations from our dual-RTX 5090 host are documented separately and do not establish Radeon performance.
Read the nine revenue observations from our host โContinue with the complete check
This summary answers one specific question. Before ordering, also verify the exact model, dimensions, PSU, connectors, display outputs and return policy.