Files
club-3090/services/openwebui/docker-compose.yml
T
41b1c83e84 Wire SearXNG web search + Qdrant vector DB into Open WebUI (#472)
* Wire SearXNG web search + Qdrant vector DB into Open WebUI

Two integrations OWUI wasn't actually using despite both services running:

SearXNG (web search) — the compose used the pre-v0.9 env names
(ENABLE_RAG_WEB_SEARCH / RAG_WEB_SEARCH_*), which v0.9.6 ignores, so web
search was silently off. Renamed to the current names (ENABLE_WEB_SEARCH,
WEB_SEARCH_ENGINE, WEB_SEARCH_RESULT_COUNT, WEB_SEARCH_CONCURRENT_REQUESTS)
→ config path rag.web.search.*, seeded from env on a fresh volume.
SearXNG already serves JSON at :8088.

Qdrant (vector DB) — added VECTOR_DB=qdrant + QDRANT_URI (:6333) +
QDRANT_ON_DISK so document/web-search RAG vectors go to the qdrant service
instead of the embedded Chroma default. VECTOR_DB is a plain startup env
(applies every boot). Switching backends doesn't migrate existing Chroma
embeddings — re-index any knowledge.

Existing-volume note (this rig): web search also needed a one-time DB edit
(rag.web.search.enable/engine) because OWUI's PersistentConfig DB value
overrides the env once persisted; fresh installs get it from the env above.

Live-validated: ENABLE_WEB_SEARCH/engine=searxng at runtime; a web-search
round-trip through OWUI returned results AND created the qdrant collection
'open-webui_web-search' (3 points) — proving SearXNG + Qdrant work together
(search → embed → Qdrant, not Chroma).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* OWUI web search: bypass embed/retrieve + disable hybrid (reliable results)

Web search ran (SearXNG → embed → Qdrant) but the model never saw the
results (sources: 0): the embed→retrieve path needs a reranking model for
hybrid search, and with RAG_RERANKING_MODEL unset the rerank step returned
0 chunks, so nothing reached the prompt.

- BYPASS_WEB_SEARCH_EMBEDDING_AND_RETRIEVAL=true — inject SearXNG results
  straight into context (no vector round-trip).
- BYPASS_WEB_SEARCH_WEB_LOADER=true — use the search snippets, not full-page
  fetches → small, fast context.
- ENABLE_RAG_HYBRID_SEARCH=false — pure vector retrieval for document RAG too
  (hybrid needs a reranker; set RAG_RERANKING_MODEL + flip true for recall).

Live-validated: a web-search chat against gemma-4-12b-int8 now returns a
cited answer ("Anthropic recently announced ... [1]"), sources attached.
(Applied to the live rig via the PersistentConfig DB since env doesn't
override an existing volume; these env vars seed fresh installs.)

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* OWUI: enable hybrid RAG with bge-reranker-base (in-process, CPU)

Re-enable hybrid (BM25 + vector + rerank) document-RAG retrieval now that a
cross-encoder reranker is configured. bge-reranker-base (~280 MB) loads
IN-PROCESS in OWUI on CPU — no GPU contention, no separate service —
auto-downloaded from HF on first use. Document-RAG only (web search bypasses
retrieval). Swap to bge-reranker-v2-m3 (~2.3 GB) for higher recall.

Applied to the live rig via the PersistentConfig DB (rag.reranking_model +
rag.enable_hybrid_search); these env vars seed fresh installs.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-25 07:14:40 +05:00

92 lines
5.6 KiB
YAML

services:
open-webui:
# Pinned (was :main) — v0.9.6 live-validated 2026-06-09: image-gen + chat config
# survive the migration. Bump deliberately, re-validate the image-gen wiring.
image: ghcr.io/open-webui/open-webui:v0.9.6
container_name: open-webui
restart: unless-stopped
ports:
- "8080:8080"
volumes:
- open-webui-data:/app/backend/data
env_file:
# Image generation: Ideogram-4 → ComfyUI (:8188). PersistentConfig — applies on a
# FRESH volume; on an existing volume set it in Admin → Settings → Images instead.
- ./imagegen.env
environment:
# Chat models — bootstrap one connection PER BACKEND on a FRESH volume (PersistentConfig),
# pointing OWUI straight at each model's own port rather than at the LiteLLM gateway. OWUI
# hides models from an UNREACHABLE connection, so each model appears in the picker only while
# its gpu-mode scene is actually serving — scene-accurate, no listing dead routes. (Going via
# LiteLLM :4000 instead would always list every catalog model, since the gateway is always up;
# that masking is exactly what this avoids. LiteLLM stays running for other clients — it just
# leaves OWUI's picker.) Served names match the catalog, so model IDs are unchanged:
# :8090 = ai-studio qwen director (uncensored prompt crafter; studio lanes call it internally,
# and it doubles as a general chat model). Up only in the ai-studio scene.
# :8010 = qwen3.6-27b-autoround (gpu-mode qwen27b)
# :8051 = qwen3.6-35b-a3b-autoround (gpu-mode qwen35b-a3b)
# :8032 = gemma-4-31b-autoround (gpu-mode gemma-31b)
# :8038 = gemma-4-12b-int8 (gpu-mode gemma12b)
# :8199 = deckard-40b (gpu-mode deckard)
# Plural *_URLS / *_KEYS are ;-separated. On an EXISTING volume the stored config wins — re-run
# scripts/setup-ai-studio.sh (it registers these + drops a stale :4000) or set them in
# Admin → Settings → Connections. Backends are no-auth on localhost, so the key is a placeholder.
- OPENAI_API_BASE_URLS=http://host.docker.internal:8090/v1;http://host.docker.internal:8010/v1;http://host.docker.internal:8051/v1;http://host.docker.internal:8032/v1;http://host.docker.internal:8038/v1;http://host.docker.internal:8199/v1
- OPENAI_API_KEYS=sk-noauth;sk-noauth;sk-noauth;sk-noauth;sk-noauth;sk-noauth
# Ollama removed from the stack 2026-06-22 — disable the API so OWUI doesn't
# probe a dead :11434 on boot (chat routes through the OpenAI API above).
- ENABLE_OLLAMA_API=false
# Leave empty → Open WebUI generates a strong random key on first boot and persists it
# in the data volume (.webui_secret_key). NEVER ship a hardcoded shared secret (it would
# let anyone forge session tokens). Override via the host env only to share sessions
# across replicas: WEBUI_SECRET_KEY=<your-secret> docker compose up -d
- WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY:-}
# --- Web search (SearXNG) --- v0.9.x env names: the old ENABLE_RAG_WEB_SEARCH /
# RAG_WEB_SEARCH_* names were renamed upstream and are now IGNORED (→ web search silently
# off). SearXNG runs at :8088 with JSON output enabled; OWUI substitutes <query>.
# Config path rag.web.search.* — seeded from these on boot when unset in the DB.
- ENABLE_WEB_SEARCH=true
- WEB_SEARCH_ENGINE=searxng
- WEB_SEARCH_RESULT_COUNT=5
- WEB_SEARCH_CONCURRENT_REQUESTS=5
- SEARXNG_QUERY_URL=http://host.docker.internal:8088/search?q=<query>&format=json
# Inject SearXNG results straight into the prompt instead of embed→vector-retrieve. The
# retrieve path needs a reranker for hybrid search and was silently returning 0 chunks (web
# results never reached the model); bypassing it is the reliable path. web_loader off = use
# the search snippets, not full-page fetches → small, fast context. (rag.web.search.bypass_*)
- BYPASS_WEB_SEARCH_EMBEDDING_AND_RETRIEVAL=true
- BYPASS_WEB_SEARCH_WEB_LOADER=true
# --- Vector DB (Qdrant) --- document / knowledge RAG vectors live in the qdrant service
# (:6333) instead of the embedded Chroma default. VECTOR_DB is a plain STARTUP env (applies
# on every boot, not PersistentConfig). QDRANT_ON_DISK persists vectors to qdrant's volume.
# Switching backends does NOT migrate existing Chroma embeddings — re-index any knowledge.
- VECTOR_DB=qdrant
- QDRANT_URI=http://host.docker.internal:6333
- QDRANT_ON_DISK=true
# Fallback / additional providers (keys loaded if set in env)
- BRAVE_SEARCH_API_KEY=${BRAVE_SEARCH_API_KEY:-}
- TAVILY_API_KEY=${TAVILY_API_KEY:-}
- GOOGLE_PSE_API_KEY=${GOOGLE_PSE_API_KEY:-}
- GOOGLE_PSE_ENGINE_ID=${GOOGLE_PSE_ENGINE_ID:-}
- SERPER_API_KEY=${SERPER_API_KEY:-}
- JINA_API_KEY=${JINA_API_KEY:-}
# --- URL-paste / page loading ---
- ENABLE_RAG_LOCAL_WEB_FETCH=true
# --- Image/file upload, document RAG ---
# Hybrid (BM25 + vector + rerank) retrieval for document RAG. Needs a cross-encoder reranker:
# bge-reranker-base (~280 MB) runs IN-PROCESS in OWUI on CPU (no GPU contention, no separate
# service), auto-downloaded from HF on first use. Document-RAG only — web search bypasses
# retrieval (above). For higher recall swap to BAAI/bge-reranker-v2-m3 (~2.3 GB, still CPU).
- ENABLE_RAG_HYBRID_SEARCH=true
- RAG_RERANKING_MODEL=BAAI/bge-reranker-base
extra_hosts:
- "host.docker.internal:host-gateway"
volumes:
open-webui-data: