Contact
AI compute

The three questions that decide the GPU

AI compute6 min read

By Humphrey Theodore K. Ng’ambi

Updated 13 September 2026

Nvidia logo

Which NVIDIA GPU for Which Workload

The GPU is the single most consequential — and most expensive — choice in an AI or visualisation server. Getting it wrong is costly in two directions: an over-specified data-centre GPU sitting idle on a graphics workload, or an under-specified card that cannot hold your model in memory. This guide decodes the current NVIDIA data-centre and professional line-up by the three things that actually decide a purchase: how much memory (and what type), what form factor, and how much power. Every figure below is drawn from NVIDIA and OEM datasheets and cited.

The three questions that decide the GPU

1. How much memory, and is it HBM or GDDR? A model has to fit in GPU memory. High-Bandwidth Memory (HBM3/HBM3e) — used on the H100, H200 and B200 — delivers multiple terabytes per second and is what large-language-model training and memory-bound inference need. GDDR6 — used on the L40S and RTX 6000 Ada — is slower per second but far cheaper per gigabyte, and is ideal for graphics, rendering and mixed inference where raw memory bandwidth is not the bottleneck.

2. SXM or PCIe? SXM is NVIDIA's socketed module that mounts on an HGX baseboard with NVLink/NVSwitch, giving every GPU a fast, direct path to every other GPU — essential for training a model too big for one card. It runs hotter (up to 700–1,000 W) and typically needs a purpose-built or liquid-assisted chassis. PCIe cards drop into a standard server slot, are air-coolable, and are the pragmatic choice for inference, graphics and smaller-scale training. Some PCIe cards (H200 NVL, A100) add 2-way or 4-way NVLink bridges for a middle ground.

3. What is the power and cooling envelope? A single H100 SXM draws up to 700 W; a B200 up to 1,000 W. Eight of them in one chassis is a 10 kW-plus machine with matching cooling and PSU demands. A 350 W L40S or 300 W RTX 6000 Ada slots into a mainstream 2U/4U server on air. Match the GPU to the room, not just the workload.

The decode table

GPUArchitectureMemoryForm factorApprox. board powerPrimary workload
B200Blackwell180 GB HBM3eSXM (HGX 8-GPU)up to 1,000 WFrontier-scale LLM training & inference
H200 SXMHopper141 GB HBM3eSXM5 (HGX/DGX)up to 700 WLarge-model training; memory-bound inference
H200 NVLHopper141 GB HBM3ePCIe dual-slotup to 600 WLLM inference in air-cooled racks
H100 SXMHopper80 GB HBM3SXM5 (HGX/DGX)up to 700 WLarge-scale training & inference
H100 PCIeHopper80 GB HBM2ePCIe dual-slot300–350 WMainstream training/inference
A100 80GB SXMAmpere80 GB HBM2eSXM4 (HGX)400 WTraining & HPC (prior generation)
A100 80GB PCIeAmpere80 GB HBM2ePCIe dual-slot300 WTraining / inference / HPC
L40SAda Lovelace48 GB GDDR6 (ECC)PCIe dual-slot (Gen4)350 WMixed AI + graphics; inference; fine-tuning; rendering
RTX 6000 AdaAda Lovelace48 GB GDDR6 (ECC)PCIe dual-slot300 WPro-viz workstation; rendering; mixed AI
Jetson AGX Orin 64GBAmpere (module)64 GB LPDDR5System-on-module15–60 W (configurable)Edge AI, robotics, autonomous machines

Matching the GPU to the job

Training large models — H100 / H200 SXM, or B200

When the model is large enough to span multiple GPUs, you want SXM on an HGX baseboard so NVLink can pool the cards. The H100 SXM (80 GB HBM3, 3.35 TB/s, 700 W) is the established workhorse. The H200 SXM keeps the same 700 W envelope but nearly doubles memory to 141 GB HBM3e at 4.8 TB/s — a straight upgrade for larger models and longer context. The B200 (Blackwell, 180 GB HBM3e, up to 1,000 W, 5th-gen NVLink at 1.8 TB/s GPU-to-GPU) is the current frontier part where budget and power allow.

Inference and mid-scale training — H100 PCIe / H200 NVL / A100

For serving models and smaller training runs in standard, air-cooled servers, PCIe is the sensible path. H100 PCIe (80 GB HBM2e, 300–350 W) drops into a mainstream 2U/4U GPU server. H200 NVL (141 GB HBM3e, up to 600 W, air-cooled, with 2-/4-way NVLink bridges) is purpose-built for LLM inference in enterprise racks that cannot take liquid-cooled SXM. The A100 (prior-generation Ampere, 40/80 GB HBM2/HBM2e) remains a strong value option — common in the refurbished channel — for training, inference and HPC.

Mixed AI + graphics, rendering and pro-viz — L40S / RTX 6000 Ada

When the workload blends inference with 3D graphics, rendering, video or simulation, the Ada Lovelace pair wins on flexibility and cost. The L40S (48 GB GDDR6 ECC, 864 GB/s, 350 W, PCIe Gen4) is the data-centre universal card — generative-AI inference, fine-tuning, Omniverse, rendering and video in one SKU. The RTX 6000 Ada (48 GB GDDR6 ECC, 300 W) is its workstation sibling for engineers, data scientists and creatives who need large memory on the desk or in a pedestal/rack workstation.

Edge and embedded AI — Jetson AGX Orin

For AI outside the data centre — robots, cameras, autonomous machines, industrial gateways — the Jetson AGX Orin 64GB module delivers up to 275 TOPS of INT8 AI performance from a 12-core Arm Cortex-A78AE CPU and an Ampere-architecture GPU (2,048 CUDA cores, 64 Tensor cores) with 64 GB LPDDR5, in a system-on-module whose power is configurable from 15 W to 60 W. It is a different class of product entirely: not a rack card but an embedded compute brick for deployment at the point of data.

Quick rules of thumb

  • Model won't fit / multi-GPU training → HBM + SXM (H100, H200, B200 on HGX).
  • Serving models in a normal rack → H100 PCIe or H200 NVL (air-cooled, NVLink-bridged).
  • AI and graphics on one card → L40S (data centre) or RTX 6000 Ada (workstation).
  • AI at the edge → Jetson AGX Orin.
  • Budget training/HPC → A100 (prior gen, strong refurb availability).

Precision-specific FLOPS/TOPS figures vary by datasheet revision and by sparsity assumptions; always confirm against the cited NVIDIA datasheet for the exact SKU before quoting a performance number to a customer.

The Server Hub briefing

South African IT hardware news, once a week.

What’s new, what it costs in rand, and what it means for the kit you run — servers and storage, networking, backup power, surveillance and print. Every claim checked against a named source.

One email a week. No third-party sharing, and unsubscribe from any issue.