The deployment module lets your team serve open models from the built-in model library and fine-tune them with LoRA training jobs — entirely on your infrastructure with no data leaving your environment.

Deployer modes

Configure the deployer in gateway.yaml:
ModeConfigDescription
kubernetesdeployer.mode: kubernetesCreates vLLM Deployments, Services, and HPAs in your cluster
dstackdeployer.mode: dstackDrives an external dstack server for GPU scheduling across AWS, GCP, Azure, or on-prem
deployer:
  mode: kubernetes
  namespace: manylayers
  reconcile_interval: 5s
  cold_start_timeout: 2m

Model library

Browse the built-in model library to see what’s available for deployment:
curl http://localhost:8180/admin/library \
  -H "Authorization: Bearer $ADMIN_KEY"

Deployment lifecycle

1

Create a deployment

Select a model from the library and create a deployment. ManyLayers provisions the necessary infrastructure resources in your cluster or dstack environment.
2

Automatic registration

ManyLayers polls deployment status every reconcile_interval. Once the model is healthy, it is automatically registered in the gateway catalog and becomes available for routing.
3

Cold start handling

If a model has scaled to zero, the first incoming request triggers a wake. ManyLayers waits up to cold_start_timeout before returning 504.
4

Scale

Adjust replicas at any time via POST /admin/deployments/{id}/scale.

Deployment management API

MethodPathAuthDescription
GET/admin/deploymentsadminList all deployments
POST/admin/deploymentsdeployer.deployments.manageCreate a deployment
GET/admin/deployments/{id}adminGet deployment status
DELETE/admin/deployments/{id}deployer.deployments.manageDelete a deployment
POST/admin/deployments/{id}/scaleadminScale replicas