GPU Cloud RFP Template: AI Infrastructure Requirements

Preview Download Ms Word Template
13 pages
0 downloads
Updated October 5, 2026

Use this GPU cloud RFP template when you're renting GPU capacity for AI training, fine-tuning or inference, whether as a reserved cluster, on-demand instances or a mix of both.

It's written for the ML engineering, infrastructure, security and procurement team running the evaluation, and it covers the whole process: when to issue an RFI first, the requirements to send, the questions that separate providers, and how to score the answers.

What’s inside this template

  • A step-by-step plan for running the evaluation, including when to issue an RFI first
  • Requirements for GPU capacity, cluster networking, storage, scheduling, isolation, observability and inference serving, each marked Must-have or Nice-to-have
  • Questions that expose capacity that isn't installed yet, oversubscribed networks, slow node replacement and charges left off the quote
  • A weighted scoring rubric with 1-to-5 definitions
  • An editable Word RFP with response codes, a pricing table and a timeline

Download the editable Microsoft Word version below. Every requirement, vendor question and scoring weight on this page is in the document, ready to tailor and send.

More Templates

LLM Gateway RFP Template (Free): AI Platform Requirements

Paste-ready LLM gateway requirements (routing and fallback, per-team keys, budgets, caching, logging, guardrails, data residency), vendor questions that separate platforms, and a scoring rubric.
View Template

AI Governance RFP Template (Free): Platform Requirements

Paste-ready AI governance platform requirements (AI inventory, risk assessments, framework mapping, model documentation, testing evidence, vendor AI risk), vendor questions and a scoring rubric.
View Template
Synthetic Data Generation Solution RFP Template

Synthetic Data Generation Solution RFP Template

Identifies and selects a comprehensive synthetic data generation platform that can create artificial datasets mimicking real-world data patterns while maintaining privacy and statistical accuracy.
View Template

How to run a GPU Cloud RFP

A GPU cloud evaluation comes down to two questions: can the provider deliver the GPUs you need, where and when you need them, and will that capacity perform and stay healthy under your workloads. The example schedule below runs about 10 weeks from kickoff to contract award; if you need a large reserved cluster with a future delivery date, start capacity conversations earlier and allow for acceptance testing after signature.

  1. Characterize your workloads and size the demand (weeks 1-2). Separate training, fine-tuning and inference, and for each one record the model sizes, the largest single job in GPUs, expected GPU-hours per month, data volumes and latency targets. Write down the numbers vendors will price against: [GPU models required], GPU counts, regions, start date, term, storage capacity and expected data transfer out.
  2. Assemble the buying team. Include ML research and engineering leads, the platform or infrastructure team that will run the clusters, networking, security, privacy and data protection, finance, procurement and legal. Bring finance in early: decide whether you can commit to reserved capacity for a term, and how GPU costs will be charged back to teams.
  3. Choose the capacity model, and decide whether you need an RFI (weeks 2-3). Decide how much demand should be reserved and how much can run on demand or on interruptible capacity, and whether you need bare-metal nodes, managed Kubernetes, managed Slurm or several of these. If you haven't compared renting with buying your own cluster, do that first. If you don't know which providers can deliver your GPU models in your regions by your start date, issue a short RFI asking for available capacity, lead times, access models and pricing structure, and use the answers to build a shortlist.
  4. Tailor and issue the RFP (weeks 3-6). Delete requirements that don't apply, adjust priorities, and add your workload details to Section 1. Describe a reference training job and an inference workload you will use in the proof of concept, so vendors can plan for it. Give vendors three weeks to respond, with a written question-and-answer window in the first week.
  5. Score the responses (weeks 6-7). Have each evaluator score independently against the rubric before comparing notes, and set aside any vendor that misses a Must-have. Check every 'yes' for whether the capacity is installed and running today or still to be delivered, and whether each feature is generally available or still on the roadmap.
  6. Run a proof of concept with two or three vendors (weeks 7-9). Run your reference training job across several nodes on the same GPU models and network fabric you would be buying, not a demo environment. Measure how throughput scales as nodes are added, data loading and checkpoint times, and inference latency under realistic load. Ask the vendor to show what happens when a node fails mid-job, and log how quickly support answers your questions.
  7. Negotiate and award (weeks 9-10). Use the pricing table to compare like for like. Settle capacity delivery dates and remedies, acceptance testing before billing starts, node replacement times, commitment flexibility, renewal caps and exit terms before you sign, not after.

GPU Cloud RFP requirements

Each line is written to paste straight into your RFP. Must-have means a vendor that cannot meet it is out; Nice-to-have earns extra points. Change the priorities to fit your situation, and delete what doesn't apply.

Functional

ID Requirement Priority
F1 Guaranteed capacity of [number] [GPU models required] in [regions] for the committed term, available from a committed delivery date. State whether reserved capacity is dedicated to us or allocated from a shared pool. Must-have
F2 Burst capacity above our reservation, on demand or reserved at short notice. State how on-demand capacity is allocated when demand is high, and the usual lead time for each GPU model. Must-have
F3 Spot or preemptible capacity for interruptible jobs, with a documented interruption notice period. Nice-to-have
F4 Access to GPU nodes as bare metal or virtual machines. State which options are available for each GPU model, and any documented virtualization overhead. Must-have
F5 A managed scheduler for our workloads: Kubernetes, Slurm or both [delete as appropriate]. State which parts the provider operates and upgrades, and which parts we operate. Must-have
F6 Gang scheduling for distributed jobs, so a multi-node job starts only when all of its GPUs are available, with job queueing, priority-based preemption and topology-aware placement of the nodes in a job. Must-have
F7 Multi-tenancy for our internal teams: separate projects or namespaces, per-team GPU quotas, priorities and fair-share policies, with GPU usage and cost reported by team and project. Must-have
F8 GPU partitioning or fractional GPU allocation for development, notebooks and small inference workloads. Nice-to-have
F9 Validated machine and container images with GPU drivers, communication libraries and [frameworks we use], kept current by the provider. Must-have
F10 Automatic health checks that detect failing GPUs, memory errors, overheating and network interconnect faults, take the affected node out of scheduling, and alert us. Must-have
F11 Automatic restart of interrupted training jobs from the last checkpoint once a failed node has been replaced. Nice-to-have
F12 GPU observability: utilization, memory use, power draw, temperature and error counts per GPU, node, job and team, in the console and through an API, with alerts for idle GPUs, failed jobs and hardware errors. Must-have
F13 Managed inference endpoints for our own models and open-weight models, with autoscaling on request rate, latency or queue depth, including scale to zero. Nice-to-have
F14 Inference serving features such as dynamic batching, quantized models, several models sharing one GPU, and canary or blue-green rollout of new model versions with rollback. Nice-to-have

Technical and architecture

ID Requirement Priority
T1 Published specifications for each proposed node type: GPU model and count, GPU memory, CPU, system memory, local NVMe capacity, and the GPU-to-GPU interconnect topology within the node. Must-have
T2 A dedicated RDMA network fabric (InfiniBand or RoCE) for multi-node training. For a cluster of [number] GPUs, state the per-GPU fabric bandwidth, the topology, and any oversubscription between switch tiers. Must-have
T3 A non-blocking fabric across our full cluster size. If the fabric is oversubscribed at any tier, state where and by how much. Nice-to-have
T4 All nodes in a reserved cluster located in one data center on one fabric. State the largest cluster of [GPU models required] you can provide on a single fabric, and whether we can add nodes to the same fabric later. Must-have
T5 Shared file storage for training data and checkpoints (a parallel or distributed file system) sized for our cluster. State sustained read throughput per node and in aggregate, and checkpoint write throughput, to be verified in the proof of concept. Must-have
T6 Object storage for datasets, artifacts and model weights with a widely used object storage API, and stated throughput between object storage and GPU nodes. Must-have
T7 GPU capacity, storage and backups in [regions]. State which GPU models are available in each region today, and which are planned, with dates. Must-have
T8 Private connectivity to our data centers and cloud environments (dedicated links or IPsec VPN), and an option for moving [data volume] of training data in bulk at onboarding. Must-have
T9 Control over GPU driver, firmware and system software versions, including version pinning, advance notice of upgrades, and maintenance scheduled around our long training runs. Must-have

Integration

ID Requirement Priority
I1 Single sign-on to the console and API via SAML 2.0 or OIDC, with role-based access control. State whether user and group provisioning via SCIM is supported. Must-have
I2 A documented REST API and CLI for all provisioning and management tasks, plus infrastructure-as-code support (for example, a Terraform provider). Must-have
I3 Export of GPU, node and job metrics to our monitoring tools through a Prometheus-compatible endpoint or OpenTelemetry, and of audit and security logs to our SIEM. Must-have
I4 Use of our own container images and private container registry, and integration with our existing MLOps, experiment tracking and model registry tools [list tools], without requiring the provider's own tooling. Must-have
I5 Export of billing and usage data, by API or scheduled file, to our cost management tools. Nice-to-have

Security and compliance

ID Requirement Priority
S1 A current SOC 2 Type II report and ISO/IEC 27001 certification covering the proposed services and data center locations, available under NDA. Must-have
S2 Attestations required for our data, such as [HIPAA, PCI DSS or FedRAMP at the required level]. Make this Must-have if these apply to your training or inference data. Nice-to-have
S3 Isolation between tenants at the compute, network fabric and storage layers, with documentation of how it is enforced wherever nodes, fabric or storage are shared. Must-have
S4 Dedicated single-tenant nodes or clusters for sensitive workloads. Make this Must-have if your training data or models require it. Nice-to-have
S5 Secure wiping of local storage, clearing of GPU memory and firmware integrity checks before any node we used is reassigned to another customer. Must-have
S6 Encryption in transit and at rest for all storage, including datasets, checkpoints and model weights. State whether customer-managed keys are supported for each storage type. Must-have
S7 Confidential computing options that protect data in use on the GPU, with remote attestation. Nice-to-have
S8 A data processing agreement and a contractual commitment that our data, model weights, logs and backups stay in [regions or jurisdictions], with a current list of data center locations and sub-processors. Must-have
S9 An exportable audit log of all console and API actions, including any access by provider staff, which must require our approval and come only from [permitted locations]. Must-have
S10 Documented patch timelines for GPU drivers, firmware and platform components, and notification of security incidents affecting our data within a defined time. State the time. Must-have

Implementation and support

ID Requirement Priority
M1 An onboarding plan covering account and identity setup, networking, storage, data transfer and migration of our first workloads, with named roles and the effort expected from both sides. Must-have
M2 Burn-in and acceptance testing of reserved clusters before handover (GPU, fabric and storage tests), with results shared with us and billing starting only after acceptance. Must-have
M3 Repair or replacement of failed nodes in reserved capacity within a stated time, with service credits or billing paused while a node is unavailable. State the replacement time. Must-have
M4 A service-level agreement for the availability of the console, API, managed scheduler and storage, with service credits. State how availability is measured. Must-have
M5 24×7 support for critical issues, with response-time targets by severity and access to engineers who can troubleshoot distributed training, fabric and storage problems. State your targets. Must-have
M6 Advance notice of planned maintenance, a public status page, and root-cause reports for incidents that affect us. Must-have
M7 A named technical account manager and periodic capacity planning reviews. Nice-to-have
M8 Engineering help tuning our training and inference workloads for the proposed hardware. Nice-to-have

Commercial and pricing

ID Requirement Priority
C1 Pricing per GPU-hour for each GPU model and capacity type (reserved by term, on-demand and spot), showing list price, discount and net price. Must-have
C2 Every other charge itemized: file and object storage, networking, IP addresses, data egress, managed services and support. State which charges are metered and which are fixed. Must-have
C3 Billing increments for on-demand and spot capacity (per second, minute or hour), and any minimum charges. Must-have
C4 Commitment terms: length, prepayment options, and what happens to committed capacity or spend we don't use. Must-have
C5 Flexibility to convert a reservation to a newer GPU generation, move it to another region, or change its size during the term. Nice-to-have
C6 Remedies if committed capacity is not delivered by the agreed date or falls short during the term, including service credits and a right to terminate. Must-have
C7 A cap on price increases at renewal. State the cap. Must-have
C8 A proof of concept on the proposed GPU models and fabric at no charge or a reduced charge. Nice-to-have
C9 Exit terms covering export of our data, checkpoints and model weights, any data transfer charges at exit, transition assistance and certified deletion at the end of the contract. Must-have

Questions to ask GPU Cloud vendors

These questions are designed to separate vendors, not to collect brochure answers. Ask for evidence: a demo, a document or a reference.

  1. For [number] [GPU models required] in [regions], what capacity can you commit, and from what date? Is that capacity installed and running today, or still to be delivered? What remedies do you offer if delivery slips?
  2. Is our reserved capacity dedicated to us or allocated from a shared pool? Can it be reclaimed or reduced for any reason during the term?
  3. Describe the network fabric a [number]-GPU training job would run on: the technology, per-GPU bandwidth, topology and switch tiers. Is it non-blocking at that size? If not, where is it oversubscribed?
  4. Will you run our reference training job across several nodes of the exact hardware and fabric you are proposing during the proof of concept, and share the results and the configuration used?
  5. How do you detect a failing GPU, a memory error or a fabric fault? Walk us through what happens to a running multi-node job, from detection to a replacement node back in service.
  6. What burn-in and acceptance tests do you run before handing over a cluster? Will you share the results, and can we run our own tests before billing starts?
  7. Who controls GPU driver, firmware and library versions? Can we pin versions, and how much notice do we get before maintenance that would interrupt a long training run?
  8. What read throughput can your shared storage sustain to every node in our cluster at once, and how long does it take to write a checkpoint of [checkpoint size]? Show both in the proof of concept.
  9. How are tenants isolated on shared nodes, the network fabric and shared storage? What happens to local disks and GPU memory before a node we used goes to another customer?
  10. Where are our data, checkpoints, model weights, logs and backups stored? Where can your support staff access them from, and how is that access approved and recorded?
  11. Show a sample monthly bill for a workload like ours, with every line item: GPU-hours, storage, networking, IP addresses, egress, managed services and support. Which charges are metered and which are fixed?
  12. Can we convert a reservation to a newer GPU generation, move it to another region or change its size during the term? On what terms?
  13. Which GPU models, regions and features in your proposal are generally available today, and which are in preview or on the roadmap? Give dates for the roadmap items.
  14. At the end of the contract, how would we move [data volume] of data, checkpoints and model weights out, how long would it take, and what would it cost?
  15. Provide a reference customer running multi-node training at a similar scale on the GPU models you are proposing.

Scoring rubric

Score each criterion from 1 to 5, multiply by its weight, and add the results. The weights below are a starting point; agree on your own before any proposals arrive so the scoring can't be bent around a favorite.

Criterion Weight What a strong response shows
Capacity commitments and availability 20% Committed capacity of the GPU models you need, in your regions, by your dates, with installed hardware or a credible delivery plan and remedies if dates slip.
Training performance: network fabric and storage 20% A fabric design stated plainly (bandwidth, topology, oversubscription) and storage throughput that held up when tested with your own jobs in the proof of concept.
Reliability, failure handling and support 15% Automatic detection of failing hardware, a stated node replacement time backed by credits, acceptance testing before billing, and support engineers who understand distributed training.
Platform: scheduling, inference and observability 10% A managed scheduler that fits how your teams work, per-team quotas, GPU-level metrics and alerts you can export, and inference serving if you need it.
Security, isolation and data residency 15% Documented tenant isolation, node sanitization between customers, contractual data residency and support access that you approve and can audit.
Implementation and onboarding 5% A realistic onboarding plan with named roles, a workable approach to moving your training data, and clear effort expected from your team.
Commercials and total cost 15% Every charge itemized, clear commitment terms with some flexibility, a renewal cap and fair exit terms over a three-year view.
Total 100%

Scoring scale

  • 5 Exceeds: meets every Must-have and most Nice-to-haves, shown in a demo or proof of concept, with a matching reference customer.
  • 4 Strong: meets every Must-have and some Nice-to-haves, with clear evidence.
  • 3 Adequate: meets most Must-haves; gaps have a credible workaround or a dated roadmap commitment.
  • 2 Weak: misses one or more Must-haves, or the answer is vague.
  • 1 Poor: does not meet the requirement, or no answer.

Frequently asked questions

What should a GPU cloud RFP include?

The GPU models, counts, regions and start dates you need, and how much capacity must be reserved; requirements for cluster networking, storage, scheduling, isolation, data residency, observability and node failure handling, each marked Must-have or Nice-to-have; inference serving if you'll use it; support and SLA terms; an itemized pricing table covering GPU-hours, storage and egress; exit terms; and the rubric you'll score with. This template includes all of them.

Should we reserve GPU capacity or use on-demand?

Reserved capacity guarantees specific GPUs for a term in exchange for a commitment, which suits steady training and inference demand. On-demand capacity needs no commitment, but there's no guarantee the GPU model you want will be available when you need it. If your demand is steady, consider reserving your baseline and covering peaks with on-demand or short-term capacity, and ask each vendor to quote both so you can compare them at your expected utilization.

What is a non-blocking network fabric for GPU clusters?

In a non-blocking fabric, the links between switch tiers carry as much bandwidth as the links into them, so every GPU can exchange data with any other at full speed at the same time. In an oversubscribed fabric, traffic that crosses switch tiers shares less bandwidth, which can slow multi-node training jobs. InfiniBand and RoCE (RDMA over Converged Ethernet) fabrics can be built either way, so ask each vendor whether the fabric is non-blocking at your cluster size and, if not, where it is oversubscribed.

How is a GPU cloud RFP different from an IaaS RFP?

A general IaaS RFP covers compute, storage and networking for a broad range of applications. A GPU cloud RFP focuses on what AI workloads need: guaranteed access to specific GPU models, an RDMA network fabric for multi-node training, storage that can keep GPUs fed and absorb checkpoints, GPU-aware scheduling and fast replacement of failed nodes. If you're buying both, run them separately or add these requirements as a section of your IaaS RFP.

How long does a GPU cloud RFP take?

The example schedule in this template runs about 10 weeks from kickoff to contract award, including three weeks for vendor responses and about three weeks for a proof of concept. A large reserved cluster may also have a delivery lead time after signature, plus time for acceptance testing, so ask each vendor for both in the RFP.

Further reading

Related RFP templates

Download Ms Word Template