GPU sizing guide
Size GPUs for training, inference, and HPC
Practical sizing for LLM training, fine-tuning, inference, and HPC—memory, node count, fabric, and facility power from Fremont architects.
Sizing guide
Document model family and parameter scale, training vs inference, framework (PyTorch, JAX, vLLM, TensorRT-LLM), checkpoint size, rack kW, cooling posture, and growth horizon before locking silicon.
Large-model training typically needs 80GB+ HBM per GPU and multi-node high-bandwidth fabric. Fine-tuning often fits 4–8 GPUs per node with a fast NVMe checkpoint tier. Inference may reduce per-GPU memory with FP8/FP4 and tensor parallelism.
Facility limits—air vs DLC, depth clearance, and power—gate form factor as much as GPU class. NTS flags those gates in architecture review.
Sizing paths
Move from RFQ inputs into HGX, PCIe, cooling, and contract-ready quoting.
Large-model training typically needs 80GB+ HBM per GPU with multi-node fabric. Fine-tuning often fits 4–8 GPUs per node. Inference may reduce per-GPU memory with FP8/FP4 and tensor parallelism—NTS maps this to your framework and SLA.
2U 4-GPU suits departmental training and edge inference. 4U 8-GPU raises density but watch rack kW and cooling. 8U HGX maximizes scale-out training and requires a facility readiness review.
Model family and parameter scale, training vs inference with batch/latency SLA, framework, checkpoint size and ingest rate, rack kW and cooling posture, target node count, and 6–18 month growth horizon.
Yes. Validated node counts and fabric sketches quote on SEWP V, ITES-4H, GSA MAS, and SLED cooperatives with CLIN-aligned attachments.