API key returns 401
Symptom: Requests return401 Unauthorized even though you have an API key.
Likely causes and fixes:
- Key not assigned to a team — every key must belong to a team. Check that the key was created with a valid
team_id. In the admin UI, go to Admin → Keys and verify the team assignment. - Key has expired — if the key was created with an
expires_atdate, it stops working immediately at that time. Create a new key with a future or no expiry date. - Using the wrong Authorization format — the header must be
Authorization: Bearer ml-your-key. Check for extra spaces or missingBearer. - Admin key vs team key — admin endpoints (
/admin/...) require the admin key. Team keys only work for inference endpoints (/v1/...).
Requests are slow or timing out
Symptom: Requests complete but take much longer than expected, or return504 Gateway Timeout.
Likely causes and fixes:
- Upstream provider is slow — check the upstream provider’s status page. Latency from OpenAI, Anthropic, or other providers varies by model and time of day. Consider configuring failover routing so slow upstreams trigger automatic fallback.
-
request_timeout_msis too low — if you’ve set a custom timeout in your gateway config, increase it. Large prompts and slow models can take 30-60 seconds. - Guardrail webhook latency — if you’ve configured a webhook-based guardrail, the webhook is called synchronously and adds to request latency. Check your webhook’s response time. Consider hosting the webhook closer to ManyLayers or caching results.
- Cache miss on all requests — semantic caching requires the embedding model to be available. If the embedding provider is slow or rate-limited, cache lookups add latency. Check embedding model response times.
Guardrail is blocking valid requests
Symptom: Legitimate requests are being blocked with400 or a guardrail rejection message.
Likely causes and fixes:
-
Regex pattern is too broad — review the regex patterns in your guardrail policy. A pattern like
\d{4}will match any four-digit number, not just credit card fragments. Test patterns against real request samples before deploying. -
Policy is in block mode instead of monitor mode — if you’re still tuning a policy, set it to
monitormode first. You’ll see which requests trigger it in the audit log without blocking real traffic. -
Checking the audit log for details — the audit log records which guardrail triggered and what pattern matched. Query recent entries filtered by your team to see what’s being caught.
Knowledge base search returns poor results
Symptom: The workspace agent’s answers don’t reflect what’s in your documents, or searches return irrelevant chunks. Likely causes and fixes:-
Documents haven’t finished indexing — check indexing status. Uploading many documents or large PDFs takes time. Wait for all documents to reach
indexedstatus before evaluating quality. -
Embedding model mismatch — if you changed the embedding model after uploading documents, old embeddings use a different vector space. Trigger a re-embed job to regenerate embeddings with the current model.
- Chunk size is suboptimal — very small chunks lose context; very large chunks dilute relevance. For factual Q&A, try 512-768 token chunks. Adjust in the knowledge base settings and re-embed.
- Reranking is not enabled — enable a reranker model on the knowledge base to improve result ordering. Reranking significantly improves quality when initial retrieval returns mixed-relevance results.
Gateway not starting
Symptom: ManyLayers fails to start, exits immediately, or refuses connections. Likely causes and fixes:- Config YAML syntax error — run your YAML through a linter. A malformed value (an unquoted string that looks like a boolean or number, an unclosed brace) will prevent the config from loading.
-
Database not reachable — ManyLayers requires a PostgreSQL database. Check that the database is running, the connection string is correct, and network connectivity is working.
-
Migration not completed — on first boot, ManyLayers runs database migrations. If these fail (usually due to insufficient permissions), the service won’t start. Check logs for migration errors and ensure your database user has
CREATE TABLEandALTER TABLEpermissions. -
Port already in use — if port 8180 (or your configured port) is in use by another process, the gateway can’t bind to it. Check with
lsof -i :8180and stop the conflicting process. -
License not found (SaaS/enterprise mode) — if you’re running in a mode that requires a license, ensure the license file or environment variable is present and valid. Check logs for
licenseerror messages.