Prerequisites
- Deployer configured (
deployer.mode: kubernetesordeployer.mode: dstack) - For Kubernetes: GPU nodes with NVIDIA GPU Operator installed
- For dstack: a running dstack server with configured GPU backends
Browse the model library
Create a deployment
Monitor deployment status
reconcile_interval (default 5s). Once the model is healthy, it is automatically registered in the model catalog.Use the deployed model
Once registered, the model appears in
/v1/models and works like any other model in the gateway:Cold start behavior
If a model has scaled to zero, the first incoming request triggers a wake. ManyLayers waits up tocold_start_timeout (default 2 minutes) for the model to become healthy. If it does not start in time, the request returns 504.