How to run a GPU Cloud RFP
A GPU cloud evaluation comes down to two questions: can the provider deliver the GPUs you need, where and when you need them, and will that capacity perform and stay healthy under your workloads. The example schedule below runs about 10 weeks from kickoff to contract award; if you need a large reserved cluster with a future delivery date, start capacity conversations earlier and allow for acceptance testing after signature.
- Characterize your workloads and size the demand (weeks 1-2). Separate training, fine-tuning and inference, and for each one record the model sizes, the largest single job in GPUs, expected GPU-hours per month, data volumes and latency targets. Write down the numbers vendors will price against: [GPU models required], GPU counts, regions, start date, term, storage capacity and expected data transfer out.
- Assemble the buying team. Include ML research and engineering leads, the platform or infrastructure team that will run the clusters, networking, security, privacy and data protection, finance, procurement and legal. Bring finance in early: decide whether you can commit to reserved capacity for a term, and how GPU costs will be charged back to teams.
- Choose the capacity model, and decide whether you need an RFI (weeks 2-3). Decide how much demand should be reserved and how much can run on demand or on interruptible capacity, and whether you need bare-metal nodes, managed Kubernetes, managed Slurm or several of these. If you haven't compared renting with buying your own cluster, do that first. If you don't know which providers can deliver your GPU models in your regions by your start date, issue a short RFI asking for available capacity, lead times, access models and pricing structure, and use the answers to build a shortlist.
- Tailor and issue the RFP (weeks 3-6). Delete requirements that don't apply, adjust priorities, and add your workload details to Section 1. Describe a reference training job and an inference workload you will use in the proof of concept, so vendors can plan for it. Give vendors three weeks to respond, with a written question-and-answer window in the first week.
- Score the responses (weeks 6-7). Have each evaluator score independently against the rubric before comparing notes, and set aside any vendor that misses a Must-have. Check every 'yes' for whether the capacity is installed and running today or still to be delivered, and whether each feature is generally available or still on the roadmap.
- Run a proof of concept with two or three vendors (weeks 7-9). Run your reference training job across several nodes on the same GPU models and network fabric you would be buying, not a demo environment. Measure how throughput scales as nodes are added, data loading and checkpoint times, and inference latency under realistic load. Ask the vendor to show what happens when a node fails mid-job, and log how quickly support answers your questions.
- Negotiate and award (weeks 9-10). Use the pricing table to compare like for like. Settle capacity delivery dates and remedies, acceptance testing before billing starts, node replacement times, commitment flexibility, renewal caps and exit terms before you sign, not after.
GPU Cloud RFP requirements
Each line is written to paste straight into your RFP. Must-have means a vendor that cannot meet it is out; Nice-to-have earns extra points. Change the priorities to fit your situation, and delete what doesn't apply.
Functional
| ID |
Requirement |
Priority |
| F1 |
Guaranteed capacity of [number] [GPU models required] in [regions] for the committed term, available from a committed delivery date. State whether reserved capacity is dedicated to us or allocated from a shared pool. |
Must-have |
| F2 |
Burst capacity above our reservation, on demand or reserved at short notice. State how on-demand capacity is allocated when demand is high, and the usual lead time for each GPU model. |
Must-have |
| F3 |
Spot or preemptible capacity for interruptible jobs, with a documented interruption notice period. |
Nice-to-have |
| F4 |
Access to GPU nodes as bare metal or virtual machines. State which options are available for each GPU model, and any documented virtualization overhead. |
Must-have |
| F5 |
A managed scheduler for our workloads: Kubernetes, Slurm or both [delete as appropriate]. State which parts the provider operates and upgrades, and which parts we operate. |
Must-have |
| F6 |
Gang scheduling for distributed jobs, so a multi-node job starts only when all of its GPUs are available, with job queueing, priority-based preemption and topology-aware placement of the nodes in a job. |
Must-have |
| F7 |
Multi-tenancy for our internal teams: separate projects or namespaces, per-team GPU quotas, priorities and fair-share policies, with GPU usage and cost reported by team and project. |
Must-have |
| F8 |
GPU partitioning or fractional GPU allocation for development, notebooks and small inference workloads. |
Nice-to-have |
| F9 |
Validated machine and container images with GPU drivers, communication libraries and [frameworks we use], kept current by the provider. |
Must-have |
| F10 |
Automatic health checks that detect failing GPUs, memory errors, overheating and network interconnect faults, take the affected node out of scheduling, and alert us. |
Must-have |
| F11 |
Automatic restart of interrupted training jobs from the last checkpoint once a failed node has been replaced. |
Nice-to-have |
| F12 |
GPU observability: utilization, memory use, power draw, temperature and error counts per GPU, node, job and team, in the console and through an API, with alerts for idle GPUs, failed jobs and hardware errors. |
Must-have |
| F13 |
Managed inference endpoints for our own models and open-weight models, with autoscaling on request rate, latency or queue depth, including scale to zero. |
Nice-to-have |
| F14 |
Inference serving features such as dynamic batching, quantized models, several models sharing one GPU, and canary or blue-green rollout of new model versions with rollback. |
Nice-to-have |
Technical and architecture
| ID |
Requirement |
Priority |
| T1 |
Published specifications for each proposed node type: GPU model and count, GPU memory, CPU, system memory, local NVMe capacity, and the GPU-to-GPU interconnect topology within the node. |
Must-have |
| T2 |
A dedicated RDMA network fabric (InfiniBand or RoCE) for multi-node training. For a cluster of [number] GPUs, state the per-GPU fabric bandwidth, the topology, and any oversubscription between switch tiers. |
Must-have |
| T3 |
A non-blocking fabric across our full cluster size. If the fabric is oversubscribed at any tier, state where and by how much. |
Nice-to-have |
| T4 |
All nodes in a reserved cluster located in one data center on one fabric. State the largest cluster of [GPU models required] you can provide on a single fabric, and whether we can add nodes to the same fabric later. |
Must-have |
| T5 |
Shared file storage for training data and checkpoints (a parallel or distributed file system) sized for our cluster. State sustained read throughput per node and in aggregate, and checkpoint write throughput, to be verified in the proof of concept. |
Must-have |
| T6 |
Object storage for datasets, artifacts and model weights with a widely used object storage API, and stated throughput between object storage and GPU nodes. |
Must-have |
| T7 |
GPU capacity, storage and backups in [regions]. State which GPU models are available in each region today, and which are planned, with dates. |
Must-have |
| T8 |
Private connectivity to our data centers and cloud environments (dedicated links or IPsec VPN), and an option for moving [data volume] of training data in bulk at onboarding. |
Must-have |
| T9 |
Control over GPU driver, firmware and system software versions, including version pinning, advance notice of upgrades, and maintenance scheduled around our long training runs. |
Must-have |
Integration
| ID |
Requirement |
Priority |
| I1 |
Single sign-on to the console and API via SAML 2.0 or OIDC, with role-based access control. State whether user and group provisioning via SCIM is supported. |
Must-have |
| I2 |
A documented REST API and CLI for all provisioning and management tasks, plus infrastructure-as-code support (for example, a Terraform provider). |
Must-have |
| I3 |
Export of GPU, node and job metrics to our monitoring tools through a Prometheus-compatible endpoint or OpenTelemetry, and of audit and security logs to our SIEM. |
Must-have |
| I4 |
Use of our own container images and private container registry, and integration with our existing MLOps, experiment tracking and model registry tools [list tools], without requiring the provider's own tooling. |
Must-have |
| I5 |
Export of billing and usage data, by API or scheduled file, to our cost management tools. |
Nice-to-have |
Security and compliance
| ID |
Requirement |
Priority |
| S1 |
A current SOC 2 Type II report and ISO/IEC 27001 certification covering the proposed services and data center locations, available under NDA. |
Must-have |
| S2 |
Attestations required for our data, such as [HIPAA, PCI DSS or FedRAMP at the required level]. Make this Must-have if these apply to your training or inference data. |
Nice-to-have |
| S3 |
Isolation between tenants at the compute, network fabric and storage layers, with documentation of how it is enforced wherever nodes, fabric or storage are shared. |
Must-have |
| S4 |
Dedicated single-tenant nodes or clusters for sensitive workloads. Make this Must-have if your training data or models require it. |
Nice-to-have |
| S5 |
Secure wiping of local storage, clearing of GPU memory and firmware integrity checks before any node we used is reassigned to another customer. |
Must-have |
| S6 |
Encryption in transit and at rest for all storage, including datasets, checkpoints and model weights. State whether customer-managed keys are supported for each storage type. |
Must-have |
| S7 |
Confidential computing options that protect data in use on the GPU, with remote attestation. |
Nice-to-have |
| S8 |
A data processing agreement and a contractual commitment that our data, model weights, logs and backups stay in [regions or jurisdictions], with a current list of data center locations and sub-processors. |
Must-have |
| S9 |
An exportable audit log of all console and API actions, including any access by provider staff, which must require our approval and come only from [permitted locations]. |
Must-have |
| S10 |
Documented patch timelines for GPU drivers, firmware and platform components, and notification of security incidents affecting our data within a defined time. State the time. |
Must-have |
Implementation and support
| ID |
Requirement |
Priority |
| M1 |
An onboarding plan covering account and identity setup, networking, storage, data transfer and migration of our first workloads, with named roles and the effort expected from both sides. |
Must-have |
| M2 |
Burn-in and acceptance testing of reserved clusters before handover (GPU, fabric and storage tests), with results shared with us and billing starting only after acceptance. |
Must-have |
| M3 |
Repair or replacement of failed nodes in reserved capacity within a stated time, with service credits or billing paused while a node is unavailable. State the replacement time. |
Must-have |
| M4 |
A service-level agreement for the availability of the console, API, managed scheduler and storage, with service credits. State how availability is measured. |
Must-have |
| M5 |
24×7 support for critical issues, with response-time targets by severity and access to engineers who can troubleshoot distributed training, fabric and storage problems. State your targets. |
Must-have |
| M6 |
Advance notice of planned maintenance, a public status page, and root-cause reports for incidents that affect us. |
Must-have |
| M7 |
A named technical account manager and periodic capacity planning reviews. |
Nice-to-have |
| M8 |
Engineering help tuning our training and inference workloads for the proposed hardware. |
Nice-to-have |
Commercial and pricing
| ID |
Requirement |
Priority |
| C1 |
Pricing per GPU-hour for each GPU model and capacity type (reserved by term, on-demand and spot), showing list price, discount and net price. |
Must-have |
| C2 |
Every other charge itemized: file and object storage, networking, IP addresses, data egress, managed services and support. State which charges are metered and which are fixed. |
Must-have |
| C3 |
Billing increments for on-demand and spot capacity (per second, minute or hour), and any minimum charges. |
Must-have |
| C4 |
Commitment terms: length, prepayment options, and what happens to committed capacity or spend we don't use. |
Must-have |
| C5 |
Flexibility to convert a reservation to a newer GPU generation, move it to another region, or change its size during the term. |
Nice-to-have |
| C6 |
Remedies if committed capacity is not delivered by the agreed date or falls short during the term, including service credits and a right to terminate. |
Must-have |
| C7 |
A cap on price increases at renewal. State the cap. |
Must-have |
| C8 |
A proof of concept on the proposed GPU models and fabric at no charge or a reduced charge. |
Nice-to-have |
| C9 |
Exit terms covering export of our data, checkpoints and model weights, any data transfer charges at exit, transition assistance and certified deletion at the end of the contract. |
Must-have |
Questions to ask GPU Cloud vendors
These questions are designed to separate vendors, not to collect brochure answers. Ask for evidence: a demo, a document or a reference.
- For [number] [GPU models required] in [regions], what capacity can you commit, and from what date? Is that capacity installed and running today, or still to be delivered? What remedies do you offer if delivery slips?
- Is our reserved capacity dedicated to us or allocated from a shared pool? Can it be reclaimed or reduced for any reason during the term?
- Describe the network fabric a [number]-GPU training job would run on: the technology, per-GPU bandwidth, topology and switch tiers. Is it non-blocking at that size? If not, where is it oversubscribed?
- Will you run our reference training job across several nodes of the exact hardware and fabric you are proposing during the proof of concept, and share the results and the configuration used?
- How do you detect a failing GPU, a memory error or a fabric fault? Walk us through what happens to a running multi-node job, from detection to a replacement node back in service.
- What burn-in and acceptance tests do you run before handing over a cluster? Will you share the results, and can we run our own tests before billing starts?
- Who controls GPU driver, firmware and library versions? Can we pin versions, and how much notice do we get before maintenance that would interrupt a long training run?
- What read throughput can your shared storage sustain to every node in our cluster at once, and how long does it take to write a checkpoint of [checkpoint size]? Show both in the proof of concept.
- How are tenants isolated on shared nodes, the network fabric and shared storage? What happens to local disks and GPU memory before a node we used goes to another customer?
- Where are our data, checkpoints, model weights, logs and backups stored? Where can your support staff access them from, and how is that access approved and recorded?
- Show a sample monthly bill for a workload like ours, with every line item: GPU-hours, storage, networking, IP addresses, egress, managed services and support. Which charges are metered and which are fixed?
- Can we convert a reservation to a newer GPU generation, move it to another region or change its size during the term? On what terms?
- Which GPU models, regions and features in your proposal are generally available today, and which are in preview or on the roadmap? Give dates for the roadmap items.
- At the end of the contract, how would we move [data volume] of data, checkpoints and model weights out, how long would it take, and what would it cost?
- Provide a reference customer running multi-node training at a similar scale on the GPU models you are proposing.
Scoring rubric
Score each criterion from 1 to 5, multiply by its weight, and add the results. The weights below are a starting point; agree on your own before any proposals arrive so the scoring can't be bent around a favorite.
| Criterion |
Weight |
What a strong response shows |
| Capacity commitments and availability |
20% |
Committed capacity of the GPU models you need, in your regions, by your dates, with installed hardware or a credible delivery plan and remedies if dates slip. |
| Training performance: network fabric and storage |
20% |
A fabric design stated plainly (bandwidth, topology, oversubscription) and storage throughput that held up when tested with your own jobs in the proof of concept. |
| Reliability, failure handling and support |
15% |
Automatic detection of failing hardware, a stated node replacement time backed by credits, acceptance testing before billing, and support engineers who understand distributed training. |
| Platform: scheduling, inference and observability |
10% |
A managed scheduler that fits how your teams work, per-team quotas, GPU-level metrics and alerts you can export, and inference serving if you need it. |
| Security, isolation and data residency |
15% |
Documented tenant isolation, node sanitization between customers, contractual data residency and support access that you approve and can audit. |
| Implementation and onboarding |
5% |
A realistic onboarding plan with named roles, a workable approach to moving your training data, and clear effort expected from your team. |
| Commercials and total cost |
15% |
Every charge itemized, clear commitment terms with some flexibility, a renewal cap and fair exit terms over a three-year view. |
| Total |
100% |
|
Scoring scale
- 5 Exceeds: meets every Must-have and most Nice-to-haves, shown in a demo or proof of concept, with a matching reference customer.
- 4 Strong: meets every Must-have and some Nice-to-haves, with clear evidence.
- 3 Adequate: meets most Must-haves; gaps have a credible workaround or a dated roadmap commitment.
- 2 Weak: misses one or more Must-haves, or the answer is vague.
- 1 Poor: does not meet the requirement, or no answer.
Frequently asked questions
What should a GPU cloud RFP include?
The GPU models, counts, regions and start dates you need, and how much capacity must be reserved; requirements for cluster networking, storage, scheduling, isolation, data residency, observability and node failure handling, each marked Must-have or Nice-to-have; inference serving if you'll use it; support and SLA terms; an itemized pricing table covering GPU-hours, storage and egress; exit terms; and the rubric you'll score with. This template includes all of them.
Should we reserve GPU capacity or use on-demand?
Reserved capacity guarantees specific GPUs for a term in exchange for a commitment, which suits steady training and inference demand. On-demand capacity needs no commitment, but there's no guarantee the GPU model you want will be available when you need it. If your demand is steady, consider reserving your baseline and covering peaks with on-demand or short-term capacity, and ask each vendor to quote both so you can compare them at your expected utilization.
What is a non-blocking network fabric for GPU clusters?
In a non-blocking fabric, the links between switch tiers carry as much bandwidth as the links into them, so every GPU can exchange data with any other at full speed at the same time. In an oversubscribed fabric, traffic that crosses switch tiers shares less bandwidth, which can slow multi-node training jobs. InfiniBand and RoCE (RDMA over Converged Ethernet) fabrics can be built either way, so ask each vendor whether the fabric is non-blocking at your cluster size and, if not, where it is oversubscribed.
How is a GPU cloud RFP different from an IaaS RFP?
A general IaaS RFP covers compute, storage and networking for a broad range of applications. A GPU cloud RFP focuses on what AI workloads need: guaranteed access to specific GPU models, an RDMA network fabric for multi-node training, storage that can keep GPUs fed and absorb checkpoints, GPU-aware scheduling and fast replacement of failed nodes. If you're buying both, run them separately or add these requirements as a section of your IaaS RFP.
How long does a GPU cloud RFP take?
The example schedule in this template runs about 10 weeks from kickoff to contract award, including three weeks for vendor responses and about three weeks for a proof of concept. A large reserved cluster may also have a delivery lead time after signature, plus time for acceptance testing, so ask each vendor for both in the RFP.
Further reading
Related RFP templates