Outcome: Run an open-weight model (Llama 3, Mistral, Qwen, or similar) on your own infrastructure with no external API calls. Your team accesses it through the same OpenAI-compatible API as any other model. This removes dependency on external providers for sensitive workloads, eliminates per-token API costs for high-volume inference, and lets you run models that aren’t available through commercial providers.

Prerequisites

  • ManyLayers running with Deployment enabled
  • GPU infrastructure (cloud or on-prem) configured with dstack or Kubernetes
  • The model weights downloaded or accessible via Hugging Face

Steps

1

Configure the deployment target

Tell ManyLayers where to run the model. In your configuration, set up the compute backend:Using dstack:
deployment:
  backend: dstack
  dstack:
    server_url: https://dstack.example.com
    project: ml-inference
    token: ${DSTACK_TOKEN}
Using Kubernetes:
deployment:
  backend: kubernetes
  kubernetes:
    namespace: inference
    image_pull_secret: registry-secret
2

Create the deployment

Define the model deployment — which model to run, how many GPUs to use, and the serving configuration.From the workspace UI: go to Deployment → New deployment, select the model, and configure resources.Via API:
curl -X POST http://localhost:8180/admin/deployments \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "llama-3-8b",
    "model_id": "meta-llama/Meta-Llama-3-8B-Instruct",
    "engine": "vllm",
    "gpu_count": 1,
    "gpu_type": "A100"
  }'
ManyLayers provisions the inference server in the background. Check status:
curl http://localhost:8180/admin/deployments/llama-3-8b \
  -H "Authorization: Bearer $ADMIN_KEY"
Wait until status is running before proceeding.
3

Register the deployment as a gateway provider

Once the deployment is running, register it as a provider in the gateway. This is what makes it callable through the standard API.
curl -X POST http://localhost:8180/admin/providers \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "self-hosted-llama",
    "type": "openai_compatible",
    "base_url": "http://llama-3-8b.inference.svc.cluster.local/v1"
  }'
4

Create a logical model pointing to your deployment

Create a logical model so your team can call the model by name through the gateway.
curl -X POST http://localhost:8180/admin/models \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "llama-3-8b",
    "provider_id": "self-hosted-llama",
    "upstream_model": "meta-llama/Meta-Llama-3-8B-Instruct"
  }'
Add the model to your team’s allow-list:
curl -X PATCH http://localhost:8180/admin/teams/$TEAM_ID \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"allowed_models": ["gpt-4o", "llama-3-8b"]}'
5

Make your first call

Your team can now call the self-hosted model using the exact same API as any other model:
curl http://localhost:8180/v1/chat/completions \
  -H "Authorization: Bearer $YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-3-8b",
    "messages": [{"role": "user", "content": "What is 2+2?"}]
  }'
All gateway controls apply — rate limits, budget checks, guardrails, audit logging — exactly as with any external provider.

Why this matters

Your self-hosted model is now a first-class provider in the gateway. You can:

Next steps

  • Use canary routing to gradually shift traffic from a commercial provider to your self-hosted model
  • Configure failover so requests fall back to a commercial provider if your deployment is unavailable
  • Set up training to fine-tune a model on your own data