7 providers · spot · scale to zero

Run inference on the cheapest GPU anywhere

gpuscale shops 7 clouds and marketplaces, serves your model on whichever is cheapest right now, and releases the hardware the moment demand stops.

helm install gpuscale deploy/helm/gpuscale --set providers.vastai.enabled=true
Offers for RTX 4090 · 24 GB · spot, right now
live · cached 7s
  • vast.ai · fi-helprovisioned0.284$0.284 per hour. Clears your ceiling.
  • verda · nl-amsprovisioned0.331$0.331 per hour. Clears your ceiling.
  • runpod · us-orprovisioned0.390$0.390 per hour. Approaching.
  • tensordock · denext in line0.472$0.472 per hour. At the ceiling.
  • gcp · us-central1-aabove ceiling0.593$0.593 per hour. Refused — never bought.
  • aws · us-east-1above ceiling0.718$0.718 per hour. Refused — never bought.

Cheapest first. Nothing right of the dashed line is ever bought, however deep the queue — that is what the ceiling means. A failed offer is held out for 60 seconds; 5 attempts, 3 offers each.

7Providers shopped per request
$0.284Cheapest 4090 in this snapshot
$0.00Cost of an idle pool
10sBatch window before dispatch
30sSpot reclaim detection
Built for a bill you control

Ceilings that refuse, not ceilings that warn

Every pool carries a per-instance and a per-pool maximum. An offer above either is passed over, however deep the queue gets.

Ceilings that refuse

$0.50 per instance and $3.00 per pool by default. Both are enforced before any provider API is called, so an offer above either is never bought.

Scale to zero

Set minNodes: 0 and an idle pool costs nothing. A 2m cooldown stops a gap between requests from provisioning again.

Spot preemption handled

Provider APIs are polled every 30s. A reclaimed instance is replaced while demand remains, and queued requests wait rather than fail.

No Kubernetes on the GPU

The rented machine runs one agent and one engine. It opens an outbound tunnel, holds no credential of yours, and needs no inbound port.

Bin-packing

When the same models fit on fewer GPUs the controller consolidates them and drains what is left over.

Metrics included

Each worker's engine is scraped and re-exposed, so one endpoint carries control-plane and per-GPU series together.

Engines: vLLM, SGLang, Ollama. Providers: AWS, Azure, GCP, RunPod, Vast.ai, TensorDock, Verda.

The product

Everything on one screen

Capacity, models, cost and chat. The running total never leaves the header, because it is the only number that can hurt you.

gpuscale · gpu-workloadsaccruing $1.005/hr of $3.00
Rented GPUs, their hourly price and lifecycle phase
ClaimProviderGPU$/hrStatusAgeModel
clm-8f31a2vast.ai · fi-helRTX 4090 · 24 GB0.284Ready41mglm-4-7
clm-1c07deverda · nl-amsA100 · 80 GB0.331Ready12mqwen3-coder-30b
clm-b45119runpod · us-orRTX 4090 · 24 GB0.390Bootstrapping80sglm-4-7
clm-6ae802vast.ai · fi-helRTX 3090 · 24 GB0.211Draining2h 06m
clm-33f9c1vast.ai · pl-wawRTX 4090 · 24 GB0.298ReclaimedReplaced

Active requests by model

gpu_api_active_requests_by_model — the series that triggers dispatch and, on the way down, release.

Chat

YouWhich model is cheapest per token today?

glm-4-7 · vast.ai · fi-helglm-4-7 on the 4090 at $0.284/hr, and it is serving you now. qwen3-coder-30b costs 17% more per hour and only wins above roughly 30B.

0

No GPUs running, nothing billing. The pool released its last node six minutes ago, after the 2m cooldown expired. This is the pool working, not the pool broken.