When deployer.mode: dstack, ManyLayers drives an external dstack server over REST. dstack handles GPU scheduling, pooling, autoscaling, and scale-to-zero across multiple infrastructure backends — AWS, GCP, Azure, Lambda Labs, or on-premises hardware.

Prerequisites

  • dstack server 0.18.x
  • A dstack project with at least one configured backend (cloud or on-prem)

Configuration

deployer:
  mode: dstack
  dstack:
    server_url: http://dstack-server:3100
    token: ${DSTACK_TOKEN}
    project: main

Helm values

deployer:
  enabled: true
  mode: dstack
  dstack:
    serverURL: http://dstack.dstack.svc:3000
    project: main
    tokenSecret: dstack-token      # name of existing Kubernetes Secret
    tokenSecretKey: token           # key within the Secret

How it works

  1. ManyLayers creates a dstack service run for each model deployment.
  2. dstack provisions GPU infrastructure and starts the inference container.
  3. ManyLayers polls the dstack service status and updates the deployment record.
  4. Once running, the model is registered in the gateway catalog and available for routing — just like any other model.

Docker Compose (for evaluation)

infra/docker/docker-compose.dev.yml ships a local dstack server and its own database in the optional profile. Name the service to start just those two:
docker compose --env-file infra/docker/.env.dev -f infra/docker/docker-compose.dev.yml up -d dstack
Create a project and token in the dstack UI at http://localhost:3100, then configure ManyLayers to point at it.
dstack service runs are long-lived model serving endpoints. ManyLayers uses these exclusively — not one-shot tasks. Training jobs generate dstack task manifests separately; see Training.