Investigation of a Discord cross-rig report (WSL, 3060Ti+4090: verify
fails on the pinned v0.3.2-preview digest, passes on the old noonghunna
snapshot; NOT reproducible on our 3090s — verify-full all-pass) surfaced
three doc-level truths worth recording:
- The 50-series gap is a TOOLCHAIN ceiling: Anbeeld's CI builds on the
Dockerfile-default CUDA 12.4, whose nvcc cannot target sm_120 (no
cubin, max PTX compute_90) — every official tag lacks Blackwell.
Upstream ask filed: Anbeeld#85 (CUDA_VERSION 12.8.1 + arch list incl
120); when it lands we retire the noonghunna snapshot entirely.
- Three registry status_notes claimed launchers inject 'server-cuda-
v0.3.0' — stale since the 2026-06-12 pin bump; they inject the
v0.3.2-preview digest from engines/beellama-local.yml install.spec.
Fixed all three (qwen dflash, gemma-12b, gemma dflash) + honest
labeling of the noonghunna snapshot as v0.3.0-feature-level and
unmaintained (predates KVarN + v0.3.1 fixes).
- The engine-notes self-build guidance was outdated: FA_ALL_QUANTS is
hardcoded in Anbeeld's cuda.Dockerfile since our PR Anbeeld#48, so a
self-build needs only CUDA_DOCKER_ARCH (+ CUDA_VERSION=12.8.1 for
sm_120). Recipe verified against his master Dockerfile 2026-07-04.
UPSTREAM.md beellama row updated (dated entry + next-triggers; the 'no
official image' claim struck through as historical). Gates: YAML +
registry import clean; status-drift / profiles-compat / switch-parity
green.
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm