Files
club-3090/services/openwebui/docker-compose.yml
T
noonghunnaandClaude Opus 4.8 7e9f5cb930 serve: neutral primary model name (qwen3.6-27b / gemma-4-31b), keep -autoround alias
The shared served-model-name `qwen3.6-27b-autoround` (from #490) mislabels the
non-autoround 27b scenes (fp8 dual-max, lmcache): /v1/models advertises
"autoround" while `root` points at qwen3.6-27b-fp8. Same class on gemma-4-31b —
after the v0.24.0 consolidation the default is cyankiwi qat-AWQ-INT4 (bf16 KV),
yet the LiteLLM route still targeted `gemma-4-31b-autoround` (a latent #482 drift).

Fix WITHOUT breaking anything, via vLLM multi-served-name:
- Every 27b scene now serves `qwen3.6-27b <its-quant-name>`; every gemma-31b
  scene serves `gemma-4-31b <its-quant-name>`. The neutral name is PRIMARY
  (honest /v1/models id); the quant-specific name is retained as a live ALIAS.
- LiteLLM: add `qwen3.6-27b` / `gemma-4-31b` canonical public routes; keep the
  `-autoround` routes as back-compat aliases (same upstream). Repairs the gemma drift.
- Migrate our own MODEL= defaults + docs (bench/verify/quality/launch/setup, c3,
  tui-core, EXAMPLES, ...) to the neutral name. Weights slugs (`-autoround-int4`)
  untouched; CHANGELOG + results/ history left as-is.

Retiring the `-autoround` alias entirely is a deliberate later step once nothing
still asks for it.

Live-verified on-rig (single/minimal, stock v0.24.0): /v1/models lists BOTH names
(root=...-autoround-int4); chat to `qwen3.6-27b` AND `qwen3.6-27b-autoround` both
return 200; `qwen3.6-27b-fp8` correctly 404s. Full shell gate 59/59 (1 = known
worktree-fixture); c3 pytest 41 passed; served-name arg-order + YAML validated.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-02 11:28:31 +00:00

92 lines
5.6 KiB
YAML

services:
open-webui:
# Pinned (was :main) — v0.9.6 live-validated 2026-06-09: image-gen + chat config
# survive the migration. Bump deliberately, re-validate the image-gen wiring.
image: ghcr.io/open-webui/open-webui:v0.9.6
container_name: open-webui
restart: unless-stopped
ports:
- "8080:8080"
volumes:
- open-webui-data:/app/backend/data
env_file:
# Image generation: Ideogram-4 → ComfyUI (:8188). PersistentConfig — applies on a
# FRESH volume; on an existing volume set it in Admin → Settings → Images instead.
- ./imagegen.env
environment:
# Chat models — bootstrap one connection PER BACKEND on a FRESH volume (PersistentConfig),
# pointing OWUI straight at each model's own port rather than at the LiteLLM gateway. OWUI
# hides models from an UNREACHABLE connection, so each model appears in the picker only while
# its gpu-mode scene is actually serving — scene-accurate, no listing dead routes. (Going via
# LiteLLM :4000 instead would always list every catalog model, since the gateway is always up;
# that masking is exactly what this avoids. LiteLLM stays running for other clients — it just
# leaves OWUI's picker.) Served names match the catalog, so model IDs are unchanged:
# :8090 = ai-studio qwen director (uncensored prompt crafter; studio lanes call it internally,
# and it doubles as a general chat model). Up only in the ai-studio scene.
# :8010 = qwen3.6-27b (gpu-mode qwen27b)
# :8051 = qwen3.6-35b-a3b-autoround (gpu-mode qwen35b-a3b)
# :8032 = gemma-4-31b (gpu-mode gemma-31b)
# :8038 = gemma-4-12b-int8 (gpu-mode gemma12b)
# :8199 = deckard-40b (gpu-mode deckard)
# Plural *_URLS / *_KEYS are ;-separated. On an EXISTING volume the stored config wins — re-run
# scripts/setup-ai-studio.sh (it registers these + drops a stale :4000) or set them in
# Admin → Settings → Connections. Backends are no-auth on localhost, so the key is a placeholder.
- OPENAI_API_BASE_URLS=http://host.docker.internal:8090/v1;http://host.docker.internal:8010/v1;http://host.docker.internal:8051/v1;http://host.docker.internal:8032/v1;http://host.docker.internal:8038/v1;http://host.docker.internal:8199/v1
- OPENAI_API_KEYS=sk-noauth;sk-noauth;sk-noauth;sk-noauth;sk-noauth;sk-noauth
# Ollama removed from the stack 2026-06-22 — disable the API so OWUI doesn't
# probe a dead :11434 on boot (chat routes through the OpenAI API above).
- ENABLE_OLLAMA_API=false
# Leave empty → Open WebUI generates a strong random key on first boot and persists it
# in the data volume (.webui_secret_key). NEVER ship a hardcoded shared secret (it would
# let anyone forge session tokens). Override via the host env only to share sessions
# across replicas: WEBUI_SECRET_KEY=<your-secret> docker compose up -d
- WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY:-}
# --- Web search (SearXNG) --- v0.9.x env names: the old ENABLE_RAG_WEB_SEARCH /
# RAG_WEB_SEARCH_* names were renamed upstream and are now IGNORED (→ web search silently
# off). SearXNG runs at :8088 with JSON output enabled; OWUI substitutes <query>.
# Config path rag.web.search.* — seeded from these on boot when unset in the DB.
- ENABLE_WEB_SEARCH=true
- WEB_SEARCH_ENGINE=searxng
- WEB_SEARCH_RESULT_COUNT=5
- WEB_SEARCH_CONCURRENT_REQUESTS=5
- SEARXNG_QUERY_URL=http://host.docker.internal:8088/search?q=<query>&format=json
# Inject SearXNG results straight into the prompt instead of embed→vector-retrieve. The
# retrieve path needs a reranker for hybrid search and was silently returning 0 chunks (web
# results never reached the model); bypassing it is the reliable path. web_loader off = use
# the search snippets, not full-page fetches → small, fast context. (rag.web.search.bypass_*)
- BYPASS_WEB_SEARCH_EMBEDDING_AND_RETRIEVAL=true
- BYPASS_WEB_SEARCH_WEB_LOADER=true
# --- Vector DB (Qdrant) --- document / knowledge RAG vectors live in the qdrant service
# (:6333) instead of the embedded Chroma default. VECTOR_DB is a plain STARTUP env (applies
# on every boot, not PersistentConfig). QDRANT_ON_DISK persists vectors to qdrant's volume.
# Switching backends does NOT migrate existing Chroma embeddings — re-index any knowledge.
- VECTOR_DB=qdrant
- QDRANT_URI=http://host.docker.internal:6333
- QDRANT_ON_DISK=true
# Fallback / additional providers (keys loaded if set in env)
- BRAVE_SEARCH_API_KEY=${BRAVE_SEARCH_API_KEY:-}
- TAVILY_API_KEY=${TAVILY_API_KEY:-}
- GOOGLE_PSE_API_KEY=${GOOGLE_PSE_API_KEY:-}
- GOOGLE_PSE_ENGINE_ID=${GOOGLE_PSE_ENGINE_ID:-}
- SERPER_API_KEY=${SERPER_API_KEY:-}
- JINA_API_KEY=${JINA_API_KEY:-}
# --- URL-paste / page loading ---
- ENABLE_RAG_LOCAL_WEB_FETCH=true
# --- Image/file upload, document RAG ---
# Hybrid (BM25 + vector + rerank) retrieval for document RAG. Needs a cross-encoder reranker:
# bge-reranker-base (~280 MB) runs IN-PROCESS in OWUI on CPU (no GPU contention, no separate
# service), auto-downloaded from HF on first use. Document-RAG only — web search bypasses
# retrieval (above). For higher recall swap to BAAI/bge-reranker-v2-m3 (~2.3 GB, still CPU).
- ENABLE_RAG_HYBRID_SEARCH=true
- RAG_RERANKING_MODEL=BAAI/bge-reranker-base
extra_hosts:
- "host.docker.internal:host-gateway"
volumes:
open-webui-data: