deployer.mode: dstack, ManyLayers drives an external dstack server over REST. dstack handles GPU scheduling, pooling, autoscaling, and scale-to-zero across multiple infrastructure backends — AWS, GCP, Azure, Lambda Labs, or on-premises hardware.
Prerequisites
- dstack server 0.18.x
- A dstack project with at least one configured backend (cloud or on-prem)
Configuration
Helm values
How it works
- ManyLayers creates a dstack service run for each model deployment.
- dstack provisions GPU infrastructure and starts the inference container.
- ManyLayers polls the dstack service status and updates the deployment record.
- Once running, the model is registered in the gateway catalog and available for routing — just like any other model.
Docker Compose (for evaluation)
infra/docker/docker-compose.dev.yml ships a local dstack server and its own database in the optional profile. Name the service to start just those two:
http://localhost:3100, then configure ManyLayers to point at it.
dstack service runs are long-lived model serving endpoints. ManyLayers uses these exclusively — not one-shot tasks. Training jobs generate dstack task manifests separately; see Training.