NVIDIA-Certified Servers Explained
NVIDIA-Certified Servers Explained
When a customer asks for an "AI server" or a "GPU server", they are almost always asking for one of a small set of OEM platforms designed to hold NVIDIA data-centre GPUs at density, with the power and cooling to keep them fed. This guide explains what these systems are, what the specifications actually mean, and how the mainstream buyer-requested models — Dell R760xa, HPE DL380a Gen11, Lenovo SR675 V3 and the Supermicro GPU SuperServer families — differ.
What "NVIDIA-Certified" actually means
NVIDIA-Certified Systems is a validation programme: NVIDIA and the server OEM jointly test a specific server-plus-GPU-plus-networking configuration and certify that it delivers the expected performance, manageability and security for AI and data-analytics workloads. It is a badge on a tested configuration, not a blanket claim about a chassis. The practical value to a buyer is assurance: a certified configuration is a known-good pairing of server, GPU and firmware, so you are not the one discovering a thermal or PCIe-lane limitation in production.
Separately, systems built around the NVIDIA HGX baseboard (the 4-GPU or 8-GPU SXM board with NVLink/NVSwitch) are a distinct, higher tier — these are the machines for training models that must span multiple GPUs.
Reading a GPU-server spec sheet
Six fields decide whether a server fits the job:
- Form factor (rack U). Density and cooling headroom. A 2U accelerator server (R760xa, DL380a) holds up to ~4 double-width GPUs; a 4U (Supermicro 421GE) up to ~10 PCIe GPUs; an 8U (Supermicro 821GE) carries the full 8-GPU SXM baseboard.
- GPU capacity — double-width vs single-width. Data-centre GPUs are double-width (dual-slot) cards. "4 double-wide or 8 single-wide" means four full-power accelerators, or eight thinner cards. This is the number that matters most.
- PCIe vs SXM. PCIe = cards in slots, air-coolable, flexible. SXM = socketed modules on an HGX baseboard with NVLink — mandatory for large multi-GPU training, but needs a purpose-built (often liquid-assisted) chassis.
- CPU platform. Current generation is Intel Xeon Scalable 4th/5th Gen or AMD EPYC 9004/9005, all with DDR5 and PCIe Gen5. PCIe Gen5 doubles per-lane bandwidth to the GPU over Gen4 — it matters for GPU-dense boxes.
- Memory and drives. DDR5 capacity for the CPU side; NVMe Gen5 bays for feeding data to the GPUs.
- Power and cooling. GPU servers are high-wattage. An 8-GPU SXM node is a 10 kW-class machine; even a 4-GPU PCIe 2U wants multi-kilowatt redundant Titanium PSUs and serious airflow.
The mainstream models compared
| Model | Form | Sockets / CPU | Max double-width GPUs | GPU topology | Notable |
|---|---|---|---|---|---|
| Dell PowerEdge R760xa | 2U | 2× Intel Xeon 4th/5th Gen | 4 (or 12 single-width) | PCIe Gen5 | GPU bays moved to chassis front for cold-air-first cooling; up to 2,800 W Titanium PSU |
| HPE ProLiant DL380a Gen11 | 2U | 2× Intel Xeon 4th/5th Gen | 4 (or 8 single-wide) | PCIe Gen5 | 'a' = accelerator-optimised; supports H100, H200 NVL, L40S; up to 3 TB DDR5 |
| Lenovo ThinkSystem SR675 V3 | 3U | 2× AMD EPYC 9004/9005 | 8 | PCIe or HGX 4-GPU SXM5 | One model does either 8× PCIe (H200/L40S) or the 4-GPU NVLink SXM baseboard |
| Supermicro SYS-421GE-TNRT | 4U | 2× Intel Xeon 4th/5th Gen | 10 | PCIe Gen5 (dual-root) | High-count PCIe workhorse; H100/A100/L40S; 4× 2,700 W Titanium |
| Supermicro SYS-821GE-TNHR | 8U | 2× Intel Xeon 4th/5th Gen | 8 (SXM) | HGX 8-GPU SXM5 + NVSwitch | Flagship NVLink node: HGX H100 80GB or HGX H200 141GB; 6× 3,000 W PSUs |
Choosing between them
- A mainstream on-prem AI/pro-viz build, standard rack, air cooling → Dell R760xa or HPE DL380a Gen11. Two-socket, 2U, up to four double-width GPUs (H100 PCIe, H200 NVL, L40S). The pragmatic majority of enterprise AI purchases land here.
- More PCIe GPUs in one box, still air-cooled → Lenovo SR675 V3 (8× PCIe) or Supermicro SYS-421GE-TNRT (up to 10× PCIe). Good for inference farms and mixed AI+graphics at scale.
- Multi-GPU training that must pool GPU memory over NVLink → the HGX SXM machines: Lenovo SR675 V3 in its 4-GPU SXM form, or the Supermicro SYS-821GE-TNHR 8-GPU node. These are the training-class systems and carry training-class power, cooling and price.
A note for the refurbished/new buyer
All five models above are current-generation (2023 platforms: Intel Xeon 4th/5th Gen or AMD EPYC 9004/9005, DDR5, PCIe Gen5) and are the safe modern floor for an AI build. Their prior-generation siblings — Dell R750xa, HPE DL380 Gen10 Plus, Lenovo SR670 V2, Supermicro X12 GPU systems — are the cheaper refurbished tier but sit on PCIe Gen4 and DDR4, which throttles GPU bandwidth and caps the newest accelerators. The 8-GPU SXM class (Supermicro 821GE and equivalents) is documented here for completeness, but a full HGX node is rarely a mainstream refurbished-channel item — treat it as a specify-and-source, build-to-order machine rather than an off-the-shelf refurb.
Always confirm the exact GPU support, PSU rating and cooling class against the model's current OEM QuickSpecs / spec sheet before quoting — GPU compatibility lists are revised as new accelerators (H200, B200) are qualified.
The Server Hub briefing
South African IT hardware news, once a week.
What’s new, what it costs in rand, and what it means for the kit you run — servers and storage, networking, backup power, surveillance and print. Every claim checked against a named source.