This guide shows you how to deploy a model from the built-in model library and make it available for inference through the gateway — running entirely on your infrastructure.

Prerequisites

  • Deployer configured (deployer.mode: kubernetes or deployer.mode: dstack)
  • For Kubernetes: GPU nodes with NVIDIA GPU Operator installed
  • For dstack: a running dstack server with configured GPU backends
1

Browse the model library

curl http://localhost:8180/admin/library \
  -H "Authorization: Bearer $ADMIN_KEY"
The library lists available models along with their GPU resource requirements.
2

Create a deployment

curl -X POST http://localhost:8180/admin/deployments \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3-8b",
    "replicas": 1
  }'
ManyLayers creates the necessary infrastructure resources in your cluster or dstack environment.
3

Monitor deployment status

curl http://localhost:8180/admin/deployments/$DEPLOYMENT_ID \
  -H "Authorization: Bearer $ADMIN_KEY"
ManyLayers polls status every reconcile_interval (default 5s). Once the model is healthy, it is automatically registered in the model catalog.
4

Use the deployed model

Once registered, the model appears in /v1/models and works like any other model in the gateway:
curl http://localhost:8180/v1/chat/completions \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3-8b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
5

Scale the deployment

curl -X POST http://localhost:8180/admin/deployments/$DEPLOYMENT_ID/scale \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"replicas": 3}'

Cold start behavior

If a model has scaled to zero, the first incoming request triggers a wake. ManyLayers waits up to cold_start_timeout (default 2 minutes) for the model to become healthy. If it does not start in time, the request returns 504.

Delete a deployment

curl -X DELETE http://localhost:8180/admin/deployments/$DEPLOYMENT_ID \
  -H "Authorization: Bearer $ADMIN_KEY"
This removes the model from the catalog and deletes the infrastructure resources in your cluster.