Deployer modes
Configure the deployer ingateway.yaml:
| Mode | Config | Description |
|---|---|---|
kubernetes | deployer.mode: kubernetes | Creates vLLM Deployments, Services, and HPAs in your cluster |
dstack | deployer.mode: dstack | Drives an external dstack server for GPU scheduling across AWS, GCP, Azure, or on-prem |
Model library
Browse the built-in model library to see what’s available for deployment:Deployment lifecycle
Create a deployment
Select a model from the library and create a deployment. ManyLayers provisions the necessary infrastructure resources in your cluster or dstack environment.
Automatic registration
ManyLayers polls deployment status every
reconcile_interval. Once the model is healthy, it is automatically registered in the gateway catalog and becomes available for routing.Cold start handling
If a model has scaled to zero, the first incoming request triggers a wake. ManyLayers waits up to
cold_start_timeout before returning 504.Deployment management API
| Method | Path | Auth | Description |
|---|---|---|---|
GET | /admin/deployments | admin | List all deployments |
POST | /admin/deployments | deployer.deployments.manage | Create a deployment |
GET | /admin/deployments/{id} | admin | Get deployment status |
DELETE | /admin/deployments/{id} | deployer.deployments.manage | Delete a deployment |
POST | /admin/deployments/{id}/scale | admin | Scale replicas |