NTS Elite APEX HGX GPU server for AI training clusters
GPU sizing guide

Size GPUs for training, inference, and HPC

Practical sizing for LLM training, fine-tuning, inference, and HPC—memory, node count, fabric, and facility power from Fremont architects.

Sizing guide

Start from workload parameters—not brochure TFLOPS

Document model family and parameter scale, training vs inference, framework (PyTorch, JAX, vLLM, TensorRT-LLM), checkpoint size, rack kW, cooling posture, and growth horizon before locking silicon.

Large-model training typically needs 80GB+ HBM per GPU and multi-node high-bandwidth fabric. Fine-tuning often fits 4–8 GPUs per node with a fast NVMe checkpoint tier. Inference may reduce per-GPU memory with FP8/FP4 and tensor parallelism.

Facility limits—air vs DLC, depth clearance, and power—gate form factor as much as GPU class. NTS flags those gates in architecture review.

Sizing paths

Workload class, platforms, and next steps

Move from RFQ inputs into HGX, PCIe, cooling, and contract-ready quoting.

What to document in your RFQ RFQ inputs What to document in your RFQ Model scale, training vs inference, framework, checkpoints, rack kW, cooling, and growth horizon.
  • Parameter scale
  • Latency / batch SLA
  • Facility limits
Explore What to document in your RFQ →
8U HGX training nodes HGX · 8-GPU 8U HGX training nodes Maximum scale-out training with NVLink/NVSwitch—facility readiness review required before PO.
  • 80GB+ HBM class
  • Multi-node fabric
  • B200 / B300 paths
Explore 8U HGX training nodes →
PCIe multi-GPU hosts 2U / 4U PCIe PCIe multi-GPU hosts 2U 4-GPU for labs and edge inference; 4U 8-GPU for dense training when rack kW allows.
  • Departmental training
  • Inference density
  • Watch rack kW
Explore PCIe multi-GPU hosts →
Air vs liquid cooling Cooling Air vs liquid cooling DLC when air cannot hold clocks on dense HGX or high-kW rows—CDU worksheets included in review.
  • Facility plant first
  • Clock stability
  • CDU sizing
Explore Air vs liquid cooling →
HGX B300 buyer’s guide B300 HGX B300 buyer’s guide Air vs DLC decision framework for Blackwell Ultra / B300 platforms.
  • Cooling choice
  • TCO drivers
  • CLIN quotes
Explore HGX B300 buyer’s guide →
HGX B200 rack design B200 DLC HGX B200 rack design CDU-to-acceptance reference architecture for liquid-cooled B200 training rows.
  • Design inputs
  • Fabric attach
  • Acceptance
Explore HGX B200 rack design →
AI infrastructure guide Companion AI infrastructure guide Broader cluster, storage, and fabric context that pairs with this GPU sizing guide.
  • Cluster patterns
  • Storage attach
  • Facility planning
Explore AI infrastructure guide →
Request a sizing review Next step Request a sizing review Share framework, GPU count targets, rack kW, and vehicle—receive node count, fabric sketch, and quote path.
  • Validated BOM
  • Fabric sketch
  • Contract packaging
Explore Request a sizing review →

Frequently asked questions

Large-model training typically needs 80GB+ HBM per GPU with multi-node fabric. Fine-tuning often fits 4–8 GPUs per node. Inference may reduce per-GPU memory with FP8/FP4 and tensor parallelism—NTS maps this to your framework and SLA.

  • Training vs inference first
  • Checkpoint tier matters
  • HPC sized separately from LLM stacks

2U 4-GPU suits departmental training and edge inference. 4U 8-GPU raises density but watch rack kW and cooling. 8U HGX maximizes scale-out training and requires a facility readiness review.

  • Lab / pilot → 2U
  • Dense training → 4U
  • Scale-out NVLink → HGX

Model family and parameter scale, training vs inference with batch/latency SLA, framework, checkpoint size and ingest rate, rack kW and cooling posture, target node count, and 6–18 month growth horizon.

  • Workload parameters
  • Facility gates
  • Growth horizon

Yes. Validated node counts and fabric sketches quote on SEWP V, ITES-4H, GSA MAS, and SLED cooperatives with CLIN-aligned attachments.

  • Architecture review first
  • Multi-vehicle packaging
  • Integration CLINs available