Files
club-3090/docs
ce71d4376a beellama: bump pin v0.3.0 → v0.3.2-preview (commit-pinned) + validate (#389)
* beellama: bump pin v0.3.0 → v0.3.2-preview (commit-pinned) + validate

Maintainer chose the v0.3.2 preview over the v0.3.1 stable for the newer
build (adds experimental KVarN KV-compression). v0.3.2 is a rolling
pre-release — Anbeeld replaces its moving Docker tags with newer branch
builds — so we pin the COMMIT-suffixed tag for an immutable pin:
  install.spec → ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.3.2-317c65e27e1e

Validated on-rig (single 3090, q5ks-dflash): boots on the preview image,
verify-full all-pass (Paris / tool_calls / streaming / thinking), prose
coherent, DFlash spec-dec active (acceptance ~0.12-0.17 on short tasks,
not collapsed). beellama has no vendored patches, so nothing to rebase.

Composes STAY 🧪 experimental: preview ≠ stable. The first stable tag now
exists (v0.3.1, server-cuda-v0.3.1, non-prerelease — Qwen3 MTP post-norm +
CUDA KV-quant fixes); repoint install.spec there to un-park (#455) once it
passes the full gate. Multiarch fallback (compose-literal, sm_120 direct-
compose) left at v0.3.0 — rebuild at the chosen tag is a separate follow-up.

scripts/tests/*.sh green (the 2 reds are pre-existing untracked-experimental-
compose artifacts: qwopus-coder + nex-n2-mini, unrelated to this pin).

Co-Authored-By: Claude Opus 4.8 <[email protected]>

* beellama: repoint pin to KVarN build BY DIGEST (the -317c65 tag lacks KVarN)

The commit-suffixed preview tag server-cuda-preview-v0.3.2-317c65e27e1e
PREDATES the KVarN merge — its --cache-type-k rejects kvarn* (only
turbo/TCQ). KVarN is only in the latest rolling server-cuda-preview-v0.3.2
build (commit 98caf25), which has no immutable commit-suffixed tag, so we
pin its DIGEST (immutable + KVarN), same as the vLLM :gemma digest pin:
  install.spec → ghcr.io/anbeeld/beellama.cpp@sha256:858e7cfbfb0d5d…

Measured (single 3090, Qwen3.6-27B Q5_K_S, -np 1): KVarN lifts the single-
request ceiling from ~196K (q5_0/q4_1; 262K OOMs) to the full 262K —
kvarn4 (≈q5_0 quality) fits 262K tight (~1GB free), kvarn2 ~3GB free.
Recall-at-depth NIAH + Qwopus-coder re-validate on the digest build pending.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-13 02:56:45 +05:00
..

club-3090 docs index

Two tracks. Pick the one that matches what you're doing.

  • User track — "I have GPUs and a model; how do I serve it?"
  • Contributor / maintainer track — "I'm working on the v0.8.0 pull pipeline, patches, or the calibration loop."

Every link below resolves to a file in this repo.


User track

Start here if you want to run a model.

Doc What it is
GETTING_STARTED.md Start here — 5-minute clone-to-curl path. No decisions, no menus.
SINGLE_CARD.md 1× RTX 3090 — workload → curated config → quick start.
DUAL_CARD.md 2× RTX 3090 (PCIe / NVLink auto-detected) — workload → config → quick start.
MULTI_CARD.md 3+ GPUs — TP scaling math, derivation from dual.yml, valid TP values.
PULL.md Any HF safetensors repo — evaluate against the KV math, honest about confidence.
BRING_YOUR_OWN.md Serve + tune + validate your own model/compose (any engine, single or dual) without touching the catalog.
HARDWARE.md Card-class questions — 4090/5090, power caps, NVLink, laptop EC power.
GLOSSARY.md TPS / KV / MTP / TP and the rest of the vocabulary.
FAQ.md Common setup and operational questions.
COMPARISONS.md Self-host vs cloud APIs — cost crossover and when each wins.
EXAMPLES.md Worked end-to-end usage examples.
ai-studio/ Club 3090 AI Studio — chat-driven, all-modality creative studio (text · image · video · audio) on 2× 3090. Start here for the architecture + the full 8-lane matrix.
ai-studio/image.md Image lanes — HiDream-O1 (top-quality/photoreal) · Ideogram-4 (design/logo/text) · Chroma (uncensored).
ai-studio/video.md Video — LTX-2.3 (video+audio) + Sulphur (uncensored), text/image→video, 60 s+ chaining.
ai-studio/audio.md Audio — Step-Audio-EditX (voice clone+edit) · Kokoro (narration) · ACE-Step (music) · Stable Audio (SFX).

Contributor / maintainer track

The v0.8.0 pull pipeline (in pipeline order)

A model slug flows through these stages. Read them in order to understand the whole.

Stage Doc What it owns
[D] COMPOSE_GENERATOR.md The #141 compose generator — the substrate that owns the arch→patches matrix.
Gate PULL_GATE.md scripts/pull.sh — the locked 6-stratum abort taxonomy, [C0]/[C2a]/[B]/[C1] gates, §4.1 confidence×verdict table.
[E] PULL_EMIT_DERIVED.md Download → boot → smoke for a download-eligible derived model; writes the §6 capture artifacts.
[F] LOOP.md The calibration loop — reads the capture bundle, classifies, runs the inbound-trust pipeline, dedups failures into the tracker.

Patch & model contribution

Doc What it is
PATCH_POLICY.md When/how a patch ships, the local-overlay vs upstream rules.
PATCH_ATTRIBUTION.md The Phase-A patch-attribution matrix — arch → engine-pin → required patches.
ADDING_MODELS.md How a new model gets added to the curated catalog.
KV_MATH.md The KV-cache math the [B] fit verdict is computed from.

Stack reference & ops

Doc What it is
ARCHITECTURE.md Repo/stack architecture overview.
UPSTREAM.md Upstream PR / issue tracker for this stack.
NIGHTLY_BUMP_RUNBOOK.md Procedure for bumping the vLLM nightly pin.
CONTAINER_RUNTIMES.md Docker / container runtime notes.

Reference matrices & deep dives

These are cross-cutting references both tracks reach for.

Doc What it is
scripts/switch.sh --list (runtime command, not a doc) The authoritative compose × slug matrix. Registry-derived from scripts/lib/profiles/compose_registry.py, so it's always current — every launchable slug with its topology, model, engine, KV format, and max ctx. Run this rather than trusting any hand-maintained table; the static lists in the per-topology docs are illustrative, this is the source of truth.
engines/ Per-engine deep dives — vLLM, llama.cpp, SGLang.
INFERENCE_ENGINES.md Engine picker — which engine for which workload, and structural gaps.
CLIFFS.md The accumulated-context / prefill failure modes (Cliff 2, Cliff 2b) and how to detect them.
DTYPE_MATRIX.md Supported dtype × model × engine matrix.
KERNEL_MATRIX.md Quant-kernel availability and alignment constraints.
QUALITY_TEST.md The quality-test harness and what it measures.
RESULTS_CARD.md The standard 3-panel format (Serving · Quality · Takeaways) for sharing a config's measured results.
STRUCTURED_COT.md The bounded-thinking / structured-CoT compose path.
TQ3_MTP_GENESIS.md TQ3 KV × MTP × Genesis-patch results and config.