October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk11 min

The LLM Portability Drill: Redeploy an Open Model on a Second GPU Cloud

A reproducible drill for moving an open-model inference deployment from one GPU cloud to a second provider: what to record, what to check, and how to validate the endpoint.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can move an open-model inference deployment from one GPU cloud to another, but only if you treat the move as a test rather than a copy. Pin the model reference, serving image, launch arguments, environment variables, secrets, model cache, and health-check timing; redeploy on the second provider; then confirm that the endpoint answers real requests. The official vLLM Kubernetes guide describes the ingredients of a GPU deployment, but it does not promise that the same configuration behaves identically on every cloud. What follows is a reproducible drill that shows what transferred and what needed provider-specific changes.

Why treat portability as something to demonstrate

A deployment that works on one provider usually works for reasons that are never written down: a cached model on a volume, a token exported in a shell profile, a GPU label that the cluster happens to use, or a health probe whose timing was tuned by trial and error. Moving the workload exposes every one of those assumptions. The drill makes them explicit.

The vLLM project documents a Kubernetes route that covers GPU-backed deployment, a persistent model cache, an optional secret for gated models, and startup checks. The project also lists other Kubernetes deployment routes, and a managed-container service can be used instead. Choose one route and record it. Two runs on the same route are a fair comparison; a Kubernetes run against a Docker pod is a different experiment, and you should label it that way. Keep in mind that the vLLM Kubernetes page exists in both a stable and a latest version, and the two may differ. Record which one you followed and its date.

Record the deployment before touching the second cloud

The first deliverable is a written record complete enough that someone else could rebuild the deployment without asking you questions. Fill in every row of the table below for the source deployment. Leave a field as “not used” when it genuinely does not apply, rather than omitting it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs
Field What to record Why it matters on a second cloud
Model reference Repository name and revision or commit, if the hub provides one A floating reference can silently change the weights you test
Access conditions License terms and whether the model is gated Gated models need a token on the new cloud before the first download
Serving image Image name, tag, and digest A tag such as “latest” is not a reproducible pin
Launch command and arguments The exact entrypoint, model argument, and every flag Flags are the most common silent difference between runs
Environment variables Name and purpose of each variable; values only for non-secret settings Cache paths and cluster-specific settings often live here
Secrets Secret name, key name, and the consumer that reads it Secrets must be recreated in the destination’s own secret store
Model cache Mount path, volume type, size, and whether it is shared or per-pod Determines whether the second cloud downloads weights on every start
GPU request GPU count, GPU type or class, and the resource key used GPU type and memory differ across providers
Context and batching settings Maximum sequence length and any batching or memory-utilization arguments These determine whether the model fits in the memory you are given
Port and API shape Container port, service port, and the API routes you call Your client code depends on these
Probes Startup, readiness, and liveness paths, delays, periods, and failure thresholds Probe timing must cover model loading on the new hardware
Endpoint exposure Internal service, port-forward, load balancer, or public ingress Each provider exposes networking differently

The drill, step by step

Step 1: Establish a baseline you can repeat

Choose one open model you are permitted to access. The vLLM guide uses Mistral-7B-Instruct-v0.3 as its example, but nothing in the drill requires that model. Run the source deployment at least twice from the recorded inputs, and note the time from pod creation to a successful first request. If the two runs differ, your record is incomplete, and you should fix it before moving on.

Step 2: Split portable settings from provider settings

Keep the deployment in version control as a manifest or template, and separate its contents into two groups. This split is a recommended working method rather than a rule any source prescribes, but it makes the differences between clouds visible.

  • Portable settings: the image digest, model reference, serving arguments, environment variables, port numbers, and probe paths.
  • Provider settings: the storage class, GPU resource label or node selector, network and ingress configuration, and any region or zone.
  • Secrets: stored in the destination’s secret mechanism and referenced by name, never written into the image or the manifest.

When the second deployment needs a change, the change should land in the provider group. If you find yourself editing a portable setting to make the second cloud work, record that as a finding about portability.

Step 3: Choose the runtime route for the second cloud

The second cloud does not have to use the same runtime, but the route determines what you manage yourself. The table below compares the routes that appear in the sources.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system
Route Where it is documented What you manage Verify before deploying
Kubernetes with GPU resources vLLM’s Kubernetes guide Deployment, service, persistent volume, secret, and probes GPU resources are schedulable on the nodes you pay for
Managed Kubernetes Lambda’s Managed Kubernetes documentation Workloads on a managed cluster; the page describes GPU, InfiniBand, shared storage, and preinstalled NVIDIA GPU and Network Operators Which GPU types and regions the cluster offers; whether InfiniBand is needed for your model
Docker pod Runpod’s guide to deploying vLLM with Docker on Runpod Container image, environment, port exposure, and pod configuration How the platform maps ports and volumes, and whether the cache persists across pod restarts
GPU rental marketplace Vast.ai’s rental and model endpoint pages Selecting a host, image, and disk; host characteristics vary Listing-level GPU, VRAM, disk, network, and price terms, which the page describes as real-time and volatile
Serverless containers with GPUs Google’s codelab on running vLLM on Cloud Run GPUs Container deployment, service configuration, and the model download path Currently available GPU options, startup limits, and whether the model cache persists between instances; the codelab does not settle these

The table is a comparison of routes, not a ranking. The sources do not establish comparable pricing, availability, or service-level terms across these providers, so the verification column is where you must do the work.

Step 4: Confirm GPU capacity before you deploy

Capacity is the first thing to check, because a missing GPU is the fastest way to waste an afternoon. On a Kubernetes route, confirm that a node advertises the GPU resource your manifest requests:

  1. List nodes with GPU capacity: kubectl get nodes -o wide.
  2. Inspect a candidate node and look for the GPU resource under Allocatable: kubectl describe node <node-name>. For NVIDIA GPUs the resource key is typically nvidia.com/gpu, but confirm the key the provider actually exposes.
  3. Check that the node’s GPU memory is enough for the weights plus the cache required by your context length and batching settings. The sources do not give a universal minimum VRAM for any model or workload, so test with the settings you recorded rather than trusting a rule of thumb.

On a non-Kubernetes route, check the equivalent in the provider’s console or CLI: the GPU type and memory of the instance you are about to start.

Step 5: Plan the model download and cache

The vLLM guide uses a persistent volume for the model cache and describes that storage as optional. It also notes that the model may take a long time to download. In the drill, the cache is a measured part of the test. Run the deployment twice on the second cloud: once with an empty cache and once with a populated one. Record both times. If the cache does not persist across pod restarts on the second provider, every restart pays the full download cost, and that belongs in your report.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Step 6: Recreate secrets in the destination

Only gated models need an access token. When one is required, create the secret inside the destination environment and reference it from the deployment. A Kubernetes example, using placeholder names you should adapt:

kubectl create secret generic model-access --from-literal=HF_TOKEN=<your-token>

Do not place the token in the container image, a ConfigMap, or a manifest committed to version control. Confirm the variable reaches the process without printing it in logs.

Step 7: Set probes to match real load times

The vLLM Kubernetes documentation cautions that a startup or readiness threshold that is too low can cause the scheduler to kill a server that is still loading. Use the cold-cache timing from Step 5 to set the probes:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting
  1. Set the startup probe’s total budget to exceed the slowest cold start you measured on the second cloud, with margin.
  2. Keep readiness from passing until the model has loaded, so traffic never reaches a server that cannot answer.
  3. Set liveness conservatively, or delay it until after startup, so it does not restart a server that is still loading.
  4. Re-measure after any change of GPU type, storage, or region, because load time changes with them.

Confirm the probe paths against the version of the server image you pinned, since the endpoints the probes call are part of the serving setup you recorded.

Validate the deployment on the second cloud

Validation has to show that the server works, not merely that the pod is running. Run these checks in order and write down the result of each.

  1. Check readiness. Confirm the pod reports Ready: kubectl get pods -l app=vllm-server. Use the label your manifest sets.
  2. Read the startup log. Check that the model finished loading and that the server reports it is listening: kubectl logs deploy/vllm-server. Note any warnings about memory, unsupported flags, or fallbacks.
  3. Expose the endpoint for testing. Forward the service to your workstation: kubectl port-forward svc/vllm-server 8000:8000. Use a service name and port that match your recorded setup.
  4. List the served model. Send curl http://localhost:8000/v1/models and confirm the model name matches the one in your record.
  5. Send an inference request. Post a short prompt to the chat or completions route the server exposes, for example curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"<served-model-name>","messages":[{"role":"user","content":"Reply with one word: ready"}],"max_tokens":8}'. A well-formed response with generated text is the pass condition.
  6. Record the outcome. Note time from deployment to first successful request, the cold-cache and warm-cache load times, every flag or manifest field you had to change, and any error you saw.

Then test the exposure you will actually use. A port-forward confirms the server; it does not prove that your public or internal endpoint, its authentication, or its timeouts work. Repeat the request through the endpoint you plan to serve from, and record any difference.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What transferred and what needed provider-specific changes

Use this table as the checklist for your final report. Fill in the second column from your own run rather than from this guide’s expectations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
  • 0dB technology lets you enjoy light gaming in relative silence
Item Expected to carry over Commonly changes on a second provider
Image digest and serving arguments Yes, if the image is pulled from a registry the second provider can reach Architecture or driver requirements may force a different image tag
Model reference and revision Yes Access to a gated model may need a new token or approval path
Environment variables Most, excluding cluster-specific paths Cache paths, network interface names, and GPU visibility settings
Model cache The arrangement, if the volume type exists Storage class, persistence behavior, and download speed
GPU request The count and model class, if the provider offers them Resource key, node selector, GPU memory, and availability
Probes Paths and structure Timing, because load time depends on hardware and storage
Endpoint exposure The API shape your client calls Ingress, load balancer, authentication, and timeouts

Troubleshooting the common failures

  • The pod stays Pending. The scheduler cannot find a node with the requested GPU. Re-run the node check in Step 4, and confirm the resource key and any node selector match what the provider exposes.
  • The server exits during model load. Check the log for memory errors. Reduce the context length or memory settings only in the recorded settings, and note the change, because it alters what you are testing.
  • The pod restarts while the log shows loading. The probes are too aggressive. Increase the startup budget using the measured cold-start time from Step 5.
  • Model download returns an authorization error. The gated-model token is missing, invalid, or was not granted access to the model. Recreate the secret in the destination and check that the variable reaches the container.
  • The endpoint is unreachable. Confirm the service, port, and exposure type. A port-forward working while the public endpoint fails points to ingress, load balancer, or firewall configuration rather than the model server.
  • The model answers, but slowly. Compare GPU type, storage performance, and whether the cache is warm. Record these as provider differences rather than assuming the software changed.

Provider notes, with their limits

  • Lambda Managed Kubernetes. Lambda’s documentation describes managed Kubernetes with GPU and InfiniBand support, shared persistent storage across nodes, and preinstalled NVIDIA GPU and Network Operators. That makes it a managed-cluster route. It does not establish that every cluster or region offers every GPU type, so confirm availability in the console before planning the drill.
  • Vast.ai. The Vast.ai site describes choosing GPUs by model, VRAM, price, and availability, and deploying model endpoints. Prices shown there are real-time and can change. Host characteristics vary from listing to listing, so record the specific host you used.
  • Runpod. Runpod’s guide covers vLLM in a Docker pod and iterating on deployment configuration. It is a useful example of a pod-based route. It is not evidence that the same operational guarantees or costs apply elsewhere.
  • Google Cloud Run GPUs. Google’s codelab shows vLLM serving an open model on Cloud Run with GPUs. Available GPU options and deployment features change over time, so verify them in Google’s current documentation before you plan around a specific configuration.

The drill does not produce a provider ranking. Compare providers on the axes in the table above, using prices checked on the day you run the drill for the exact region and configuration you chose.

What a complete report contains

A useful report is the filled-in record from the first section, the two timing measurements from Step 5, the validation results from the checklist, the list of provider-specific changes, and the dated version of each source you followed. Report failures as carefully as passes. A deployment that required three undocumented changes is still a successful drill, because the changes are now part of the record.

Portability, in the end, is not a claim about a software stack. It is a result you can show: the same recorded inputs, run on a second provider, producing a working endpoint, with every difference named.

Quick Recap

Bestseller No. 1
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.99
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$814.99
SaleBestseller No. 5
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$829.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.