Prerequisites
- ManyLayers running with Deployment enabled
- GPU infrastructure (cloud or on-prem) configured with dstack or Kubernetes
- The model weights downloaded or accessible via Hugging Face
Steps
Configure the deployment target
Tell ManyLayers where to run the model. In your configuration, set up the compute backend:Using dstack:Using Kubernetes:
Create the deployment
Define the model deployment — which model to run, how many GPUs to use, and the serving configuration.From the workspace UI: go to Deployment → New deployment, select the model, and configure resources.Via API:ManyLayers provisions the inference server in the background. Check status:Wait until
status is running before proceeding.Register the deployment as a gateway provider
Once the deployment is running, register it as a provider in the gateway. This is what makes it callable through the standard API.
Create a logical model pointing to your deployment
Create a logical model so your team can call the model by name through the gateway.Add the model to your team’s allow-list:
Why this matters
Your self-hosted model is now a first-class provider in the gateway. You can:- Route traffic between your self-hosted model and commercial providers using canary routing
- Apply guardrails to your self-hosted model outputs
- Set budgets and rate limits per team, just like with external providers
- See all requests in audit logs and metrics
Next steps
- Use canary routing to gradually shift traffic from a commercial provider to your self-hosted model
- Configure failover so requests fall back to a commercial provider if your deployment is unavailable
- Set up training to fine-tune a model on your own data