Skip to main content

Introduction

This document outlines the specifications required to be a Runpod secure cloud partner. These requirements establish the baseline, however for new partners, Runpod will perform a due diligence process prior to selection encompassing business health, prior performance, and corporate alignment. Meeting these technical and operational requirements does not guarantee selection. New partners
  • All specifications will apply to new partners on November 1, 2024.
Existing partners
  • Hardware specifications (Sections 1, 2, 3, 4) will apply to new servers deployed by existing partners on December 15, 2024.
  • Compliance specification (Section 5) will apply to existing partners on April 1, 2025.
A new revision will be released in October 2025 on an annual basis. Minor mid-year revisions may be made as needed to account for changes in market, roadmap, or customer needs.

Minimum deployment size

100kW of GPU server capacity is the minimum deployment size.

1. Hardware Requirements

1.1 GPU Compute Server Requirements

GPU Requirements

NVIDIA GPUs no older than Ampere generation.

CPU

Bus Bandwidth

Exceptions list:
  1. PCIe 4.0 x16 - A100 80GB PCI-E

Memory

Main system memory must have ECC.

Storage

There are two types of required storage, boot and working arrays. These are two separate arrays of hard drives which provide isolation between host operating system activity (boot array) and customer workloads (working array).

Boot array

Working array

1.2 Storage Cluster Requirements

Each datacenter must have a storage cluster which provides shared storage between all GPU servers. The hardware is provided by the partner, storage cluster licensing is provided by Runpod. All storage servers must be accessible by all GPU compute machines.

Baseline Cluster Specifications

Server Specifications

Storage Cluster Server Boot Array

Storage Cluster Server Working Array

Servers should have spare disk slots for future expansion without deployment of new servers. Even distribution among machines (e.g., 7 TB x 8 disks x 4 servers = 224 TB total space).

Dedicated Metadata Server for Large-Scale Clusters

Once a storage cluster exceeds 90% single core CPU on the leader node during peak hours, a dedicated metadata server is required. Metadata tracking is a single process operation, and single threaded performance is the most important metric.

1.3 CPU Server Requirements

Each datacenter should have a CPU server that to accommodate CPU-only Pods and Serverless workers. Runpod will also use this server to host additional features for which a GPU is not required (e.g., the S3-compatible API).

Baseline Cluster Specifications

Server Specifications

Storage

Boot Drive

2. Software Requirements

Operating System

Ubuntu Server 22.04 LTS Linux kernel 6.5.0-15 or later production version (Ubuntu HWE Kernel) SSH remote connection capability

BIOS Configuration

IOMMU disabled for non-VM systems Update server BIOS/firmware to latest stable version

Drivers and Software

HGX SXM System Addendum

  • NVIDIA Fabric Manager installed, activated, running, and tested
  • Fabric Manager version must match NVIDIA drivers and Kernel drivers headers
  • CUDA Toolkit, NVIDIA NSCQ, and NVIDIA DCGM installed
  • Verify NVLINK switch topology using nvidia-smi and dcgmi
  • Ensure SXM performance using dcgmi diagnostic tool

3. Data Center Power Requirements

4. Network Requirements

5. Compliance Requirements

To qualify as a Runpod secure cloud partner, the parent organization must adhere to at least one of the following compliance standards:
  • SOC 2 Type I (System and Organization Controls)
  • ISO/IEC 27001:2013 (Information Security Management Systems)
  • PCI DSS (Payment Card Industry Data Security Standard)
Additionally, partners must comply with the following operational standards: Runpod will review evidence of:
  • Physical access logs
  • Redundancy checks
  • Refueling agreements
  • Power system test results and maintenance logs
  • Power monitoring and capacity planning reports
  • Network infrastructure diagrams and configurations
  • Network performance and capacity reports
  • Security audit results and incident response plans
For detailed information on maintenance scheduling, power system management, and network operations, please refer to our documentation.

Release log

  • 2025-11-01: Initial release.