When deployer.mode: kubernetes, ManyLayers creates vLLM-based model serving infrastructure in your cluster. Models are deployed as standard Kubernetes workloads with automatic scaling and health monitoring.

Prerequisites

  • Kubernetes 1.24+ with GPU nodes (NVIDIA GPU Operator recommended)
  • ManyLayers gateway deployed with the Helm chart (deployer.enabled: true)
The Helm chart automatically creates the ServiceAccount and RBAC permissions the deployer needs:
# values.yaml
deployer:
  enabled: true
  mode: kubernetes

Configuration

deployer:
  mode: kubernetes
  namespace: manylayers         # namespace where model workloads run
  reconcile_interval: 5s        # how often to check deployment status
  cold_start_timeout: 2m        # how long to wait for a model to wake from zero

What ManyLayers creates

When you create a deployment, ManyLayers provisions these Kubernetes resources:
  1. Deployment — runs the model inference container (vLLM) with the appropriate GPU resource requests
  2. Service — ClusterIP service for the model endpoint
  3. HPA (optional) — Horizontal Pod Autoscaler based on GPU utilization or request rate

Scaling

curl -X POST http://localhost:8180/admin/deployments/$ID/scale \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"replicas": 3}'
Scale to zero is supported. When a request arrives for a scaled-to-zero model, ManyLayers automatically triggers scale-up and waits up to cold_start_timeout before returning 504.

Model artifacts

Model artifacts can be staged in MinIO or any S3-compatible storage. The deployment includes a checksum-verified init container that pulls the model archive before starting inference. Start a local MinIO instance from infra/docker/docker-compose.dev.yml — it is in the optional profile, so name the service to start it on its own:
docker compose --env-file infra/docker/.env.dev -f infra/docker/docker-compose.dev.yml up -d minio
The API is on localhost:9100 and the console on localhost:9101.