The three questions that decide the GPU
Which NVIDIA GPU for Which Workload
The GPU is the single most consequential — and most expensive — choice in an AI or visualisation server. Getting it wrong is costly in two directions: an over-specified data-centre GPU sitting idle on a graphics workload, or an under-specified card that cannot hold your model in memory. This guide decodes the current NVIDIA data-centre and professional line-up by the three things that actually decide a purchase: how much memory (and what type), what form factor, and how much power. Every figure below is drawn from NVIDIA and OEM datasheets and cited.
The three questions that decide the GPU
1. How much memory, and is it HBM or GDDR? A model has to fit in GPU memory. High-Bandwidth Memory (HBM3/HBM3e) — used on the H100, H200 and B200 — delivers multiple terabytes per second and is what large-language-model training and memory-bound inference need. GDDR6 — used on the L40S and RTX 6000 Ada — is slower per second but far cheaper per gigabyte, and is ideal for graphics, rendering and mixed inference where raw memory bandwidth is not the bottleneck.
2. SXM or PCIe? SXM is NVIDIA's socketed module that mounts on an HGX baseboard with NVLink/NVSwitch, giving every GPU a fast, direct path to every other GPU — essential for training a model too big for one card. It runs hotter (up to 700–1,000 W) and typically needs a purpose-built or liquid-assisted chassis. PCIe cards drop into a standard server slot, are air-coolable, and are the pragmatic choice for inference, graphics and smaller-scale training. Some PCIe cards (H200 NVL, A100) add 2-way or 4-way NVLink bridges for a middle ground.
3. What is the power and cooling envelope? A single H100 SXM draws up to 700 W; a B200 up to 1,000 W. Eight of them in one chassis is a 10 kW-plus machine with matching cooling and PSU demands. A 350 W L40S or 300 W RTX 6000 Ada slots into a mainstream 2U/4U server on air. Match the GPU to the room, not just the workload.
The decode table
| GPU | Architecture | Memory | Form factor | Approx. board power | Primary workload |
|---|---|---|---|---|---|
| B200 | Blackwell | 180 GB HBM3e | SXM (HGX 8-GPU) | up to 1,000 W | Frontier-scale LLM training & inference |
| H200 SXM | Hopper | 141 GB HBM3e | SXM5 (HGX/DGX) | up to 700 W | Large-model training; memory-bound inference |
| H200 NVL | Hopper | 141 GB HBM3e | PCIe dual-slot | up to 600 W | LLM inference in air-cooled racks |
| H100 SXM | Hopper | 80 GB HBM3 | SXM5 (HGX/DGX) | up to 700 W | Large-scale training & inference |
| H100 PCIe | Hopper | 80 GB HBM2e | PCIe dual-slot | 300–350 W | Mainstream training/inference |
| A100 80GB SXM | Ampere | 80 GB HBM2e | SXM4 (HGX) | 400 W | Training & HPC (prior generation) |
| A100 80GB PCIe | Ampere | 80 GB HBM2e | PCIe dual-slot | 300 W | Training / inference / HPC |
| L40S | Ada Lovelace | 48 GB GDDR6 (ECC) | PCIe dual-slot (Gen4) | 350 W | Mixed AI + graphics; inference; fine-tuning; rendering |
| RTX 6000 Ada | Ada Lovelace | 48 GB GDDR6 (ECC) | PCIe dual-slot | 300 W | Pro-viz workstation; rendering; mixed AI |
| Jetson AGX Orin 64GB | Ampere (module) | 64 GB LPDDR5 | System-on-module | 15–60 W (configurable) | Edge AI, robotics, autonomous machines |
Matching the GPU to the job
Training large models — H100 / H200 SXM, or B200
When the model is large enough to span multiple GPUs, you want SXM on an HGX baseboard so NVLink can pool the cards. The H100 SXM (80 GB HBM3, 3.35 TB/s, 700 W) is the established workhorse. The H200 SXM keeps the same 700 W envelope but nearly doubles memory to 141 GB HBM3e at 4.8 TB/s — a straight upgrade for larger models and longer context. The B200 (Blackwell, 180 GB HBM3e, up to 1,000 W, 5th-gen NVLink at 1.8 TB/s GPU-to-GPU) is the current frontier part where budget and power allow.
Inference and mid-scale training — H100 PCIe / H200 NVL / A100
For serving models and smaller training runs in standard, air-cooled servers, PCIe is the sensible path. H100 PCIe (80 GB HBM2e, 300–350 W) drops into a mainstream 2U/4U GPU server. H200 NVL (141 GB HBM3e, up to 600 W, air-cooled, with 2-/4-way NVLink bridges) is purpose-built for LLM inference in enterprise racks that cannot take liquid-cooled SXM. The A100 (prior-generation Ampere, 40/80 GB HBM2/HBM2e) remains a strong value option — common in the refurbished channel — for training, inference and HPC.
Mixed AI + graphics, rendering and pro-viz — L40S / RTX 6000 Ada
When the workload blends inference with 3D graphics, rendering, video or simulation, the Ada Lovelace pair wins on flexibility and cost. The L40S (48 GB GDDR6 ECC, 864 GB/s, 350 W, PCIe Gen4) is the data-centre universal card — generative-AI inference, fine-tuning, Omniverse, rendering and video in one SKU. The RTX 6000 Ada (48 GB GDDR6 ECC, 300 W) is its workstation sibling for engineers, data scientists and creatives who need large memory on the desk or in a pedestal/rack workstation.
Edge and embedded AI — Jetson AGX Orin
For AI outside the data centre — robots, cameras, autonomous machines, industrial gateways — the Jetson AGX Orin 64GB module delivers up to 275 TOPS of INT8 AI performance from a 12-core Arm Cortex-A78AE CPU and an Ampere-architecture GPU (2,048 CUDA cores, 64 Tensor cores) with 64 GB LPDDR5, in a system-on-module whose power is configurable from 15 W to 60 W. It is a different class of product entirely: not a rack card but an embedded compute brick for deployment at the point of data.
Quick rules of thumb
- Model won't fit / multi-GPU training → HBM + SXM (H100, H200, B200 on HGX).
- Serving models in a normal rack → H100 PCIe or H200 NVL (air-cooled, NVLink-bridged).
- AI and graphics on one card → L40S (data centre) or RTX 6000 Ada (workstation).
- AI at the edge → Jetson AGX Orin.
- Budget training/HPC → A100 (prior gen, strong refurb availability).
Precision-specific FLOPS/TOPS figures vary by datasheet revision and by sparsity assumptions; always confirm against the cited NVIDIA datasheet for the exact SKU before quoting a performance number to a customer.
The Server Hub briefing
South African IT hardware news, once a week.
What’s new, what it costs in rand, and what it means for the kit you run — servers and storage, networking, backup power, surveillance and print. Every claim checked against a named source.