Back to deployments
New: guided deploy
Deploy a model endpoint on dedicated GPUs
Follow the four steps below to get a private dedicated endpoint
1
Model
Choose a curated model, or search for a Hugging Face model.
2
Compute
Set the minimum and maximum number of replicas.
3
Scaling
Choose the warm replica baseline and peak capacity.
4
Review
Name the endpoint and deploy!
How auto-scaling works
Your minimum replicas stay warm and billable. Additional replicas start when traffic increases, up to your selected maximum.
requests
running replicas
Always-on baseline
Spike → more replicas
Quiet → back to baseline
