Compare NVIDIA, AMD, Intel, and Apple Silicon on memory bandwidth, VRAM, and bandwidth-per-euro value. Real specs, real prices — no marketing numbers.
Filter by model size to see which GPUs can run it, then sort by what matters: throughput, value, or VRAM.
| GPU | VRAM | Mem BW | TDP | Street | BW/€ | |
|---|---|---|---|---|---|---|
| Loading data… | ||||||
Three angles — fastest bandwidth under budget, best bandwidth per €, and most VRAM.
Bandwidth-derived speed estimates from hardware specs. Real street prices where available, MSRP fallback otherwise.
Token generation at batch size 1 is memory-bandwidth-bound: every forward pass streams the full model weights from VRAM. Higher GB/s means faster generation, regardless of compute units. Bandwidth is therefore the most honest single-number proxy for inference speed without model-specific assumptions.
VRAM, memory bandwidth, TDP, and clock data from official manufacturer specs and verified press releases.
Real market prices from Geizhals.de where available. MSRP used as fallback for workstation and Apple Silicon.
All figures are from manufacturer specs. Use them to compare GPUs against each other — not as absolute performance targets. Real-world inference speed depends on driver version, software stack, model quantisation, and context length.