Back to deployments
New: guided deploy

Deploy a model endpoint on dedicated GPUs

Follow the four steps below to get a private dedicated endpoint

1
Model
Choose a curated model, or search for a Hugging Face model.
2
Compute
Set the minimum and maximum number of replicas.
3
Scaling
Choose the warm replica baseline and peak capacity.
4
Review
Name the endpoint and deploy!
How auto-scaling works

Your minimum replicas stay warm and billable. Additional replicas start when traffic increases, up to your selected maximum.

requests
running replicas
Always-on baseline
Spike → more replicas
Quiet → back to baseline
1

Choose a model

Choose a recommended model, or search Hugging Face by model name or id.

Fast

up to 32 GB

Compact models that load quickly and are cheapest to run.

Heavy

65 – 192 GB

Large models that require more GPU memory.

Massive

193 GB or more

Very large models that require multiple high-memory GPUs.

2

Choose compute

Llama 3.1 8B Instruct needs at least 24 GB of GPU memory. We've preselected the smallest GPU that fits - you can pick a larger one for more throughput.

Loading GPU options…
3

Choose scaling

Minimum replicas stay running and billable. Maximum replicas define how far the endpoint can scale under load.

Selected

Queue load

A shorter queue scales sooner. A longer queue uses running replicas more fully before adding capacity.

Selected
4

Review and deploy

Name the endpoint, add any deployment-only settings, then confirm the cost before creating it.

You need to have accepted the model license on Hugging Face first.

Advanced runtime options

Optionally override the selected model with your own image, or add runtime variables.

No variables

Leave empty to deploy the selected model with the managed vLLM runtime.

No environment variables configured.

Runtime
Llama 3.1 8B Instruct
Compute
Select a GPU
Scaling
1 warm, scales to 3
Queue load
4 queued requests per replica
Runtime variables
None

What you'll pay

Quiet hours
$0.00/hour
1 replica always ready.
At peak
$0.00/hour
All 3 replicas running at peak load.
Required minimum balance
$0.00
We reserve enough balance for one peak hour so scaling can start immediately.

How it's calculated: $0.0000/hr × 1 GPU × running replicas. You're billed per second of actual run time.

Security token unavailable. Refresh the page before deploying.