Vector store options
ManyLayers supports two vector store backends for RAG knowledge bases:| Backend | Config | Persistence | Multi-replica | Use case |
|---|---|---|---|---|
memory | rag.vector_store: memory | No (lost on restart) | No | Evaluation and single-node testing |
qdrant | rag.vector_store: qdrant | Yes | Yes | Production |
Configuring Qdrant
Kubernetes (in-cluster Qdrant)
The Helm chart can deploy a Qdrant instance alongside your gateway:qdrant.enabled: true, the chart automatically injects the Qdrant service URL into the gateway config.
Docker Compose
Qdrant is part of theoptional profile in infra/docker/docker-compose.dev.yml. Name the service to start it on its own, without the rest of the stack:
http://localhost:6433 (override with QDRANT_HOST_PORT).
Qdrant for semantic caching
You can also use Qdrant to share the semantic cache across all gateway replicas:ml_semantic_cache collection in Qdrant automatically. The embedding model must be available in the team’s allowed models.
Retrieval quality settings
These settings are managed through the admin UI (Admin → Settings → RAG) and stored in the database:| Setting | Default | Description |
|---|---|---|
rag.hybrid_search | false | Enable BM25 keyword scoring fused with vector search via RRF for better coverage |
rag.recency_half_life | 720h | Exponential decay for document timestamps; 0 = disabled |
rag.rerank_model | "" | Cross-encoder model for a rerank pass after initial retrieval |
rag.relevance_filter_enabled | false | LLM-based post-retrieval relevance classification to filter low-quality chunks |
rag.relevance_filter_model | "" | Model used for relevance classification (e.g. gpt-4o-mini) |