docs: formalize the add-a-model workflow for the post-refactor architecture + agent discoverability

ADDING_MODELS.md was stale (pre-refactor) and AGENTS.md lacked the full catalog
flow; neither covered the profile-catalog compatibility class that hotfix #236
exposed. Refresh both for the <quant>/ layout + registry-as-single-source-of-truth,
and add a thin repo CLAUDE.md so users' AI agents discover the workflow.

- docs/ADDING_MODELS.md:
  - "Three paths" intro (serve safetensors via pull.sh · run a local GGUF · catalog)
    + a new "Run a local GGUF without the catalog" 3-step recipe (the pull.sh gap).
  - Fix the stale weights schema (a MAP keyed by quant-slug, not a list) + add
    kvcalc_key + default_port==PORT to the registry example.
  - New "Step 4b — Profile-catalog compatibility" (the #236 class: engine
    supported_model_families, hardware supported_kv_formats, canonical-scenario fit,
    patches.yml chat-template, catalog-size guard, +1 ../ mount depth, registry-
    derived launchers).
  - Rewrite Step 7 to run the FULL guard suite (table of what each gate guards) +
    the baseline-vs-last-tag rule. Update diagram + checklist. De-link the
    /opt/ai/CLAUDE.md reference (path leak + 404) -> point at AGENTS.md.
- AGENTS.md: new "Adding a model — full workflow" at-a-glance block linking
  docs/ADDING_MODELS.md, with the catalog steps + #236-compat + the catalog guard
  tests (the in-repo entry point for AI agents working in a clone).
- CLAUDE.md (new): thin pointer to AGENTS.md so Claude-Code agents pick up the same
  guidance without the two large files drifting.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
noonghunna
2026-05-26 22:16:11 +00:00
parent a74398d64f
commit 37871c6317
3 changed files with 118 additions and 28 deletions

View File

@@ -157,6 +157,15 @@ This rule applies to **shipped composes AND local-only test composes** — apply
When testing a new model, create the directory hierarchy from the start: `models/<new-model>/<engine>/compose/<topology>/<quant-slug>/<serving>.yml`. The quant slug must match the `weights_variant` key in `scripts/lib/profiles/models/<model>.yml` and `scripts/lib/profiles/weights.py`. When the model isn't Qwen3-Next, write `Genesis: N/A — Genesis is Qwen3-Next-specific` in the profile schema so readers don't expect Genesis-style perf folds where they don't apply.
#### Adding a model — full workflow → [`docs/ADDING_MODELS.md`](docs/ADDING_MODELS.md)
Read that doc before catalog work; the at-a-glance for agents:
- **Just serving, not cataloging?** Safetensors → `scripts/pull.sh <org/Model> --profile-like vllm/minimal`; a self-grabbed **GGUF** → copy an ik/llama compose and point `--model` at it (no registry/profile needed — see ADDING_MODELS "Run a local GGUF without the catalog"). The steps below are only for promoting a model into the **curated catalog**.
- **Catalog steps the compose alone doesn't cover:** (1) `scripts/lib/profiles/models/<id>.yml``weights:` is a **map keyed by quant-slug**, not a list; (2) a `compose_registry.py` entry (`weights_variant`=slug · `kvcalc_key` — vLLM `"<model>:<profile>"`, ik/llama `"SKIP"` · `default_port` == the compose's `${PORT:-NNNN}`); (3) launchers **auto-derive** from the registry — never edit `launch.sh`/`switch.sh`; promote a default via the `DEFAULTS` map.
- **Profile-catalog compatibility (easy to miss — hotfix #236):** the new `(model, engine, KV-format)` combo must validate or `test-profiles-compat` / `diagnose-profile` go red. Add the model's `family` to the engine's `supported_model_families` (`scripts/lib/profiles/engines/*.yml`), the KV format to the hardware profiles' `supported_kv_formats` (`scripts/lib/profiles/hardware/*.yml`); register any vendored chat-template in `scripts/lib/profiles/patches.yml` (with the symmetric-protocol `drift_guard`); bump the `test-compose-registry-disk` size-count.
- **Run the FULL catalog test suite**, not just the serving tests in [Tests](#tests): `for t in scripts/tests/*.sh; do bash "$t"; done`. Key gates: `test-compose-registry-disk`, `test-compose-mounts-resolve` (the `../` depth), `test-model-weights-registry`, `test-switch-registry-parity` + `test-launch-registry-parity`, `test-profiles-compat`, `test-patch-attribution`, plus `tools/kv-calc.py --calibration`. A narrow subset shipped a model with two real catalog gaps (#236) — and some failures are pre-existing/env, so **baseline against the last release tag** before treating one as a blocker.
#### Where do experimental / unvalidated composes live?
**Same directory as shipped composes, but kept untracked until validation passes.** Don't create a separate `experimental/` subdirectory — the relative paths to `../patches/...` and `../cache/...` are calibrated to the compose dir, and promoting an experiment from a sub-folder would require re-pathing every mount.