deployer.mode: kubernetes, ManyLayers creates vLLM-based model serving infrastructure in your cluster. Models are deployed as standard Kubernetes workloads with automatic scaling and health monitoring.
Prerequisites
- Kubernetes 1.24+ with GPU nodes (NVIDIA GPU Operator recommended)
- ManyLayers gateway deployed with the Helm chart (
deployer.enabled: true)
Configuration
What ManyLayers creates
When you create a deployment, ManyLayers provisions these Kubernetes resources:- Deployment — runs the model inference container (vLLM) with the appropriate GPU resource requests
- Service — ClusterIP service for the model endpoint
- HPA (optional) — Horizontal Pod Autoscaler based on GPU utilization or request rate
Scaling
cold_start_timeout before returning 504.
Model artifacts
Model artifacts can be staged in MinIO or any S3-compatible storage. The deployment includes a checksum-verified init container that pulls the model archive before starting inference. Start a local MinIO instance frominfra/docker/docker-compose.dev.yml — it is in the optional profile, so name the service to start it on its own:
localhost:9100 and the console on localhost:9101.