Vector store options

ManyLayers supports two vector store backends for RAG knowledge bases:
BackendConfigPersistenceMulti-replicaUse case
memoryrag.vector_store: memoryNo (lost on restart)NoEvaluation and single-node testing
qdrantrag.vector_store: qdrantYesYesProduction

Configuring Qdrant

rag:
  enabled: true
  vector_store: qdrant
  qdrant:
    url: http://qdrant:6433     # Qdrant REST API address
    api_key: ${QDRANT_API_KEY}  # omit or set to "" for unauthenticated OSS Qdrant

Kubernetes (in-cluster Qdrant)

The Helm chart can deploy a Qdrant instance alongside your gateway:
# values.yaml
qdrant:
  enabled: true
  image:
    repository: qdrant/qdrant
    tag: "v1.9.2"
  service:
    port: 6333
  persistence:
    enabled: true
    size: 20Gi
When qdrant.enabled: true, the chart automatically injects the Qdrant service URL into the gateway config.

Docker Compose

Qdrant is part of the optional profile in infra/docker/docker-compose.dev.yml. Name the service to start it on its own, without the rest of the stack:
docker compose --env-file infra/docker/.env.dev -f infra/docker/docker-compose.dev.yml up -d qdrant
It listens on http://localhost:6433 (override with QDRANT_HOST_PORT).

Qdrant for semantic caching

You can also use Qdrant to share the semantic cache across all gateway replicas:
cache:
  enabled: true
  ttl: 5m
  semantic:
    enabled: true
    embedding_model: text-embedding-3-small
    threshold: 0.92
    store: qdrant    # shares semantic cache across all replicas
This creates a ml_semantic_cache collection in Qdrant automatically. The embedding model must be available in the team’s allowed models.

Retrieval quality settings

These settings are managed through the admin UI (Admin → Settings → RAG) and stored in the database:
SettingDefaultDescription
rag.hybrid_searchfalseEnable BM25 keyword scoring fused with vector search via RRF for better coverage
rag.recency_half_life720hExponential decay for document timestamps; 0 = disabled
rag.rerank_model""Cross-encoder model for a rerank pass after initial retrieval
rag.relevance_filter_enabledfalseLLM-based post-retrieval relevance classification to filter low-quality chunks
rag.relevance_filter_model""Model used for relevance classification (e.g. gpt-4o-mini)