LLM Gateway RFP Template (Free): AI Platform Requirements

Preview Download Ms Word Template
14 pages
0 downloads
Updated October 5, 2026

Use this LLM gateway RFP template when your applications call large language models from several providers, or from your own self-hosted models, and you need one governed layer in front of them for access, routing, cost control, logging and guardrails.

It's written for the platform engineering, security and procurement team running the evaluation, and it covers the whole process: when to issue an RFI first, the requirements to send, the questions that separate vendors, and how to score the answers.

What’s inside this template

  • A step-by-step plan for running the evaluation, including when an RFI should come first
  • Requirements for routing and fallback, a single API, per-team keys, budgets, cost attribution, caching, logging, guardrails and model catalog governance, each marked Must-have or Nice-to-have
  • Questions that expose added latency, provider features lost in translation, open-source versus paid-edition gaps and roadmap-only features
  • A weighted scoring rubric with 1-to-5 definitions
  • An editable Word RFP with response codes, a pricing table that separates gateway fees from model usage, and a timeline

Download the editable Microsoft Word version below. Every requirement, vendor question and scoring weight on this page is in the document, ready to tailor and send.

More Templates

GPU Cloud RFP Template: AI Infrastructure Requirements

Paste-ready GPU cloud requirements (capacity commitments, cluster networking, storage, scheduling, isolation, inference, pricing), vendor questions that separate providers, and a scoring rubric.
View Template

AI Governance RFP Template (Free): Platform Requirements

Paste-ready AI governance platform requirements (AI inventory, risk assessments, framework mapping, model documentation, testing evidence, vendor AI risk), vendor questions and a scoring rubric.
View Template
Synthetic Data Generation Solution RFP Template

Synthetic Data Generation Solution RFP Template

Identifies and selects a comprehensive synthetic data generation platform that can create artificial datasets mimicking real-world data patterns while maintaining privacy and statistical accuracy.
View Template

How to run an LLM Gateway RFP

An LLM gateway sits in the path of every model call your applications make, so the evaluation involves platform engineering, application teams, security, privacy and finance at once, and most of the work is agreeing on scope before vendors get involved. The example schedule below runs about 12 weeks from kickoff to contract award; stretch it if you need a self-hosted deployment in several regions or a formal privacy review of prompt logging.

  1. Inventory current LLM use and agree on scope (weeks 1-2). List every application that calls a model today, which providers and models it uses, where its API keys are stored and who pays the bill. Include self-hosted models. Record the numbers vendors will price against: applications, teams, monthly requests and tokens, current spend and the regions you operate in.
  2. Assemble the buying team. Include the platform or AI engineering team that will run the gateway, two or three application teams that will use it, security, privacy and legal, AI governance, network, finance or FinOps, and procurement. Bring privacy in early: the gateway can log every prompt and response, so decide up front what may be logged, what must be redacted, how long logs are kept and where they are stored.
  3. Settle the deployment model, and decide whether you need an RFI (weeks 2-4). Decide whether prompts may pass through a vendor-hosted service, or whether the gateway must run in your own cloud account or data center. Decide too whether you need a gateway only, or a broader AI platform with prompt management, evaluation and agent tooling. If either question is still open, if you are weighing an open-source gateway against a commercial one, or if your long list has more than five vendors, issue a short RFI first and use the answers to build a shortlist.
  4. Tailor and issue the RFP (weeks 4-7). Delete requirements that don't apply, adjust priorities, and add your environment details to Section 1. Prepare a test set now for the proof of concept: representative prompts, prompts containing masked personal data, and prompt-injection attempts. Give vendors three weeks to respond, with a written question-and-answer window in the first week.
  5. Score the responses (weeks 7-8). Have each evaluator score independently against the rubric before comparing notes, and set aside any vendor that misses a Must-have. Check every 'yes' for whether the feature is generally available today, in preview or on the roadmap, and whether it needs a paid edition or an add-on.
  6. Run a proof of concept with two or three vendors (weeks 8-11). Route real traffic from one or two representative applications through each gateway. Force a provider outage, hit a hard budget limit, measure the latency the gateway adds with guardrails on and off, run your test set through the guardrails, and check that logs are redacted as configured. Reconcile the gateway's cost report with a provider invoice, and call references of a similar size.
  7. Negotiate and award (weeks 11-12). Use the pricing table to compare like for like, keeping gateway fees separate from model usage. Settle whether model usage is billed by the vendor or directly by providers, renewal caps, overage terms and exit assistance before you sign, not after.

LLM Gateway RFP requirements

Each line is written to paste straight into your RFP. Must-have means a vendor that cannot meet it is out; Nice-to-have earns extra points. Change the priorities to fit your situation, and delete what doesn't apply.

Functional

ID Requirement Priority
F1 One API endpoint through which our applications reach hosted model providers and our self-hosted models, compatible with a widely used chat-completions-style API and with the API formats our applications already use, so an application can change models through configuration rather than code. Must-have
F2 Support for chat and embedding requests, streamed responses, tool (function) calling and structured output. State which provider-specific features are passed through unchanged, which are translated into a common format, and which are not supported. Must-have
F3 Support for image, audio and document inputs, and for providers' batch processing interfaces, through the same API. Nice-to-have
F4 Routing rules based on application, team, model alias, request attributes or data classification, which administrators can change without redeploying applications. Must-have
F5 Automatic retries and fallback to alternate models, providers, regions or self-hosted deployments on errors, timeouts or provider rate limits, with a fallback order we configure. Must-have
F6 Load balancing across several deployments, accounts or regions of the same model, and weighted traffic splitting so a new model or model version can be tested on a share of traffic before full rollout. Must-have
F7 Gateway-issued API keys for each team, application or project, scoped to approved models and budgets, with expiry and revocation. Provider credentials are held centrally and never exposed to applications. Must-have
F8 Token and spend budgets per team, application, project or key, with alerts at thresholds we set (by email, chat or webhook) and hard limits that block requests when a budget is used up. Must-have
F9 Rate limits on requests and tokens per minute, and concurrency limits, per key, team or application, returning error responses that applications can handle. Must-have
F10 Cost attribution for every request by team, application, project, user and model, using provider price lists (including differently priced token types such as cached input), our negotiated rates and our own cost rates for self-hosted models, with showback and chargeback reports we can export. Must-have
F11 Response caching for identical requests, with a time-to-live set per application, cache scoping per application or tenant, and a way for an application to bypass the cache on a request. Must-have
F12 Semantic caching that can return a stored response for a similar prompt, with a configurable similarity threshold, scoping by tenant and data permission, and cache invalidation. Nice-to-have
F13 A log record for every request: application, user, model, provider, tokens, latency, cost, status and policy decisions. Capture of full prompts and responses must be configurable per application, including an option to store no content at all. Must-have
F14 Redaction or masking of sensitive data in prompts and responses before they are written to logs, and retention periods we configure per application. Must-have
F15 Input and output guardrails that detect personal data (PII) and custom patterns we define, such as account numbers or credentials, and redact, mask or block them, with policies set per application. Must-have
F16 Content filters for harmful content categories, with thresholds and actions (block, warn or log) set per application, provided natively or through integration with guardrail services we choose. Must-have
F17 Detection of prompt-injection and jailbreak attempts, in user input and in content returned by retrieval or tools, with block, warn or log actions. Nice-to-have
F18 Hooks or plug-ins that call our own or third-party guardrail and policy services before and after each model call. Must-have
F19 A model catalog that administrators govern: which models and versions are approved, for which teams, regions and data classifications, with model versions pinned per application. State how you tell us when a provider schedules a model we use for retirement. Must-have
F20 Dashboards for request volume, latency (including time to first token), error rates, token usage and cost by application, team, model and provider. Must-have
F21 Prompt templates stored in the platform, with versioning, approval before production use and rollback, which applications call by reference. Nice-to-have
F22 Evaluation hooks: sampling of production requests into evaluation datasets, scoring of live traffic (including model-graded checks), an API to record user feedback, and export to our evaluation tools. Nice-to-have
F23 Governance of agent tool use through the gateway, such as proxying Model Context Protocol (MCP) servers with authentication, tool allowlists and logging. Nice-to-have

Technical and architecture

ID Requirement Priority
T1 State which deployment models you offer: vendor-hosted SaaS, a data plane in our cloud account with a vendor-managed control plane, or fully self-hosted (state whether self-hosted deployment works in disconnected environments, with offline updates). For each, identify which components process or store prompt and response content, and what data or telemetry is sent to you. Must-have
T2 A gateway tier that scales horizontally with no single point of failure, deployable across availability zones and regions. State what happens to traffic if the control plane can't be reached. Must-have
T3 Added latency of no more than [X] milliseconds at the 95th percentile at [requests per second], with streamed tokens passed to the application as they arrive. State your measured figures with logging and guardrails on and off, and how you measured them. Must-have
T4 Data residency controls: requests from a given application or data classification are routed only to model endpoints in regions we approve, and logs are stored in a region we choose. Must-have
T5 Private connectivity between our applications, the gateway and cloud model services (private endpoints, peering or an equivalent), so model traffic does not have to cross the public internet. State the options for each cloud and deployment model. Must-have
T6 Support for self-hosted open-weight models behind inference servers with compatible APIs, with health checks and the same routing, limits, logging and guardrails as hosted models. Must-have
T7 Management of routes, keys, budgets and guardrail policies as code, through an API and an infrastructure-as-code or GitOps workflow, with promotion between development, test and production environments. Nice-to-have

Integration

ID Requirement Priority
I1 Single sign-on to the administration console with our identity provider via SAML 2.0 or OIDC, with user and group provisioning via SCIM. Must-have
I2 Authentication of applications and end users with tokens from our identity provider (for example, OAuth 2.0 or OIDC tokens) as an alternative to static keys, so requests can be attributed to an individual user. Nice-to-have
I3 Storage of provider credentials in our secrets manager or an encrypted vendor store, with rotation that does not interrupt traffic. Must-have
I4 Export of traces and metrics through OpenTelemetry to our observability platform, with spans for each model call, retry, fallback and guardrail check. State whether you follow the OpenTelemetry semantic conventions for generative AI. Must-have
I5 Streaming of request logs and administrator audit logs to our SIEM and data platform in a documented format, plus a REST API for logs, usage and cost data. Must-have

Security and compliance

ID Requirement Priority
S1 A current SOC 2 Type II report and ISO/IEC 27001 certification covering the proposed hosted services, available under NDA. Must-have
S2 ISO/IEC 42001 certification for AI management, and FedRAMP authorization or equivalent regional certifications for public-sector use. Make the latter Must-have if you are a public-sector buyer. Nice-to-have
S3 A contractual commitment that our prompts, responses, logs and metadata are not used to train or improve any model, and a statement of the data-retention terms that apply when requests use model provider accounts held by the vendor. Must-have
S4 Encryption in transit and at rest for all stored prompts, responses and logs. State whether customer-managed encryption keys are supported. Must-have
S5 Administrator access through SSO with MFA, role-based permissions (for example administrator, developer and auditor), separation so one team cannot read another team's logs, and an audit log of every configuration and policy change. Must-have
S6 A data processing agreement, a current list of sub-processors (including any model providers the vendor uses on our behalf), and support for our privacy obligations (for example, GDPR, or HIPAA with a business associate agreement where it applies). Must-have
S7 A contractual commitment to notify us of security incidents affecting our data within a defined time. State the time. Must-have

Implementation and support

ID Requirement Priority
M1 A phased migration plan: inventory of current provider integrations and keys, onboarding of applications in agreed waves, and then enforcement of gateway-only access to model providers, with named roles and the effort expected from both sides. Must-have
M2 Identification of who delivers the implementation (vendor, partner or both) and the relevant experience of the named team. Must-have
M3 A service-level agreement for the availability of every vendor-hosted component, with service credits. State how availability is measured. Must-have
M4 24×7 support for critical issues, with response-time targets by severity, for both hosted and self-hosted deployments. State your targets. Must-have
M5 Advance notice of changes that affect routing, guardrail behavior or API compatibility, a public changelog and status page, and root-cause reports for incidents that affect us. Must-have
M6 Training for platform administrators, and onboarding material and code examples for application developers. Nice-to-have

Commercial and pricing

ID Requirement Priority
C1 Pricing broken down by component, metric (per request, per token volume, per gateway instance, per user or platform fee) and term, showing list price, discount and net price. Must-have
C2 A statement of whether model usage is billed through you (with any markup or fee) or directly by providers on our own accounts. Price both if you offer both. Must-have
C3 Identification of every proposed feature that requires a paid edition, a higher tier or an add-on module, including features missing from any open-source or community edition. Must-have
C4 A cap on price increases at renewal. State the cap. Must-have
C5 Terms that apply if we exceed licensed request, token or instance volumes, with no interruption of service before we have been notified. Must-have
C6 A proof of concept at no charge, with the vendor's engineering support. Nice-to-have
C7 Exit terms covering export of configuration (routes, policies, budgets and model catalog), logs and evaluation data in usable formats, and transition assistance at the end of the contract. Must-have

Questions to ask LLM Gateway vendors

These questions are designed to separate vendors, not to collect brochure answers. Ask for evidence: a demo, a document or a reference.

  1. During the proof of concept, force an outage of one provider. Show the fallback, what the calling application receives, and what happens to a streamed response that fails partway through.
  2. How much latency does the gateway add at the median and the 95th percentile at our expected request volume, with logging and guardrails on and off? How do you measure it, and where should the gateway run to keep it low?
  3. For each provider we use, which features (tool calling, structured output, multimodal input, batch processing, provider-side prompt caching) pass through unchanged, which are translated into a common format, and which are lost?
  4. How quickly do you support a new model or a change to a provider's API? Give recent examples, with the provider's release date and the date you supported it.
  5. Are rate limits and budgets enforced exactly across all gateway instances and regions, or approximately? What happens to a request, including a streamed one, when a team reaches a hard budget limit?
  6. Show a cost report for one team and reconcile it with a provider invoice. How do you handle cached-input pricing, batch pricing, our negotiated rates and the cost of our self-hosted models?
  7. In each deployment option you offer, which components process or store prompt and response content, and what data or telemetry is sent back to you?
  8. Show PII redaction on one request: in the prompt before it reaches the model, in the response, and in the logs. Can redacted values be restored in the response for the user who supplied them, and how much latency does redaction add?
  9. How do you detect prompt injection, including instructions hidden in retrieved documents or tool outputs? What are the known limits, and will you run your detection against our own test prompts during the proof of concept?
  10. If you propose semantic caching, how do you stop a cached response from being served to a user who isn't allowed to see it, or after the underlying data has changed?
  11. Can the gateway govern agent traffic, such as calls to Model Context Protocol (MCP) servers and tools, with authentication, allowlists and logging? Which parts are generally available and which are in preview?
  12. How do we get production traffic into our evaluation process: sampling requests into datasets, scoring live traffic and capturing user feedback? Which evaluation tools do you integrate with?
  13. Which features in your proposal are in the open-source or community edition, which require a paid edition or add-on, and which are on the roadmap? Give dates for the roadmap items.
  14. If the gateway itself fails, what do our applications experience? Describe your most significant service outage in the last two years: the cause, the customer impact and what changed afterwards.
  15. Provide a reference customer of similar size that moved several applications from direct provider integrations onto your gateway.

Scoring rubric

Score each criterion from 1 to 5, multiply by its weight, and add the results. The weights below are a starting point; agree on your own before any proposals arrive so the scoring can't be bent around a favorite.

Criterion Weight What a strong response shows
Routing, reliability and API compatibility 20% Fallback demonstrated under a forced outage, measured added latency, and provider features passed through without loss.
Access control, quotas and cost attribution 15% Scoped per-team keys, budgets and rate limits that hold under load, and cost reports that reconcile with provider invoices.
Guardrails and data protection 15% PII redaction and content filters shown on your own test prompts, set per application, with honest limits stated for prompt-injection detection.
Logging, observability and evaluation 10% Complete request records with redaction, OpenTelemetry export, and a practical path from production traffic to evaluation.
Deployment, data residency and private connectivity 15% A deployment model that keeps prompts where your policies require, region controls for model calls and logs, and private network paths.
Implementation and support 10% A credible phased migration plan, an experienced named team, and clear SLAs and support targets.
Commercials and total cost 15% Gateway fees kept separate from model usage, no hidden markup, a renewal cap, and fair exit terms over a three-year view.
Total 100%

Scoring scale

  • 5 Exceeds: meets every Must-have and most Nice-to-haves, shown in a demo or proof of concept, with a matching reference customer.
  • 4 Strong: meets every Must-have and some Nice-to-haves, with clear evidence.
  • 3 Adequate: meets most Must-haves; gaps have a credible workaround or a dated roadmap commitment.
  • 2 Weak: misses one or more Must-haves, or the answer is vague.
  • 1 Poor: does not meet the requirement, or no answer.

Frequently asked questions

What should an LLM gateway RFP include?

Your current LLM usage and the numbers vendors will price against; requirements for routing and fallback, API compatibility, per-team keys, budgets, cost attribution, caching, logging, guardrails, observability and model catalog governance, each marked Must-have or Nice-to-have; deployment, data residency and private connectivity requirements; security and support terms; a pricing table that separates gateway fees from model usage; and the rubric you'll score with. This template includes all of them.

What is an LLM gateway?

An LLM gateway, also called an AI gateway, is a proxy between your applications and the models they call. Applications send requests to one API, and the gateway handles authentication, routing, fallback, rate limits, budgets, logging and guardrails before passing each request to a hosted provider or a self-hosted model. Some products add prompt management, evaluation or agent tooling and are sold as AI platforms, so decide which of those you need before you issue the RFP.

Should we use an open-source LLM gateway or a commercial one?

An open-source gateway that you run yourself keeps prompts inside your network and has no license fee, but your team owns upgrades, scaling, high availability and on-call support. A commercial or managed gateway moves that work to the vendor and typically adds enterprise features such as SSO, audit logs and support SLAs. If a vendor offers both, ask exactly which of your requirements need the paid edition.

Does an LLM gateway add latency?

Yes, some. Every request makes an extra network hop and passes through authentication, logging and policy checks, and guardrails that call a classifier or another model add more. How much depends on the product, where it runs relative to your applications and model endpoints, and which policies are switched on, so ask vendors for measured figures and repeat the measurement with your own traffic in the proof of concept.

How long does an LLM gateway RFP take?

The example schedule in this template runs about 12 weeks from kickoff to contract award, including three weeks for vendor responses and three to four weeks for a proof of concept. Allow more time if you need a self-hosted deployment in several regions, or if legal and privacy review of prompt logging is new to your organization.

Further reading

Related RFP templates

Download Ms Word Template