Commit Graph
8 Commits
Author SHA1 Message Date
noonghunna 2f8bade82c fix(docs): bump curl smoke-test max_tokens 30 → 200 (#14)
Qwen3.6 thinks before answering by default, so a "Capital of France?"
smoke with max_tokens=30 returns truncated mid-`<think>` content. apnar
hit this on a working stack (verify-full.sh all green) and wasted time
debugging a non-bug.

Bump all 7 user-facing curl examples to max_tokens=200 (covers a typical
think block + the one-sentence answer with headroom).

verify-full.sh / verify.sh / verify-stress.sh stay at max_tokens=30
because they already pass chat_template_kwargs.enable_thinking=false,
which skips the think block entirely.

EXAMPLES.md gets an inline note explaining the headroom + the alternative
(disable thinking via chat_template_kwargs) for users who want a tighter
smoke.
2026-04-30 21:59:17 +00:00
noonghunna 3d151b9edc feat(vllm): structured-CoT bounded-thinking compose (cross-rig port)
Port andthattoo/structured-cot to our stack — Qwen3.6-27B AutoRound INT4
dense / 1× RTX 3090 / vLLM nightly + MTP n=3 + TQ3 KV. Re-benched on
full HumanEval+ 164 + LiveCodeBench v6 50.

Headline (max_tokens=4096, greedy):
- HumanEval+ 164:  FSM 92.7% vs FREE 88.4% (+4.3pp), 30.7× compression
- LiveCodeBench v6 50: FSM 66.0% vs FREE 42.0% (+24pp), 26.2× compression

The +Δpp partly reflects FSM dodging the max_tokens=4096 truncation trap
rather than pure reasoning gain — see docs/STRUCTURED_COT.md "Honest
caveats" for the full picture.

Three port surprises worth keeping (all in docs):
1. vLLM dev205+ defaults StructuredOutputsConfig.enable_in_reasoning=False;
   grammar mask only fires post-</think> unless overridden.
2. Legacy extra_body={"guided_grammar": ...} is silently dropped on
   dev205+ (tip-off: identical FREE/FSM token counts). Use the new
   structured_outputs.grammar field.
3. Qwen3.6 chat template auto-prefixes <think>\n; drop the leading
   literal from upstream grammars when porting.

Files added:
- models/qwen3.6-27b/vllm/compose/docker-compose.bounded-thinking.yml
- docs/STRUCTURED_COT.md (public writeup)
- models/qwen3.6-27b/vllm/diagnostics/structured-cot-bench.md (internal)

Files updated:
- scripts/launch.sh wizard + scripts/switch.sh variant map
- models/qwen3.6-27b/README.md (recommended single-card list, patch surface)
- models/qwen3.6-27b/vllm/README.md (compose menu)
- docs/SINGLE_CARD.md (TL;DR table now four rows)
- CHANGELOG.md (new top entry)

Also reverts the long-text.yml experimental flag added during smoke
testing (the flag now lives only in bounded-thinking.yml) and adds the
Genesis pre-flight check we previously skipped on long-text.

Credit: andthattoo for the technique, the grammar files, and the eval
harness.
2026-04-30 21:54:13 +00:00
noonghunnaandClaude Opus 4.7 ebacba1efd fix: address open issues #1, #4, #7
Closes / addresses 3 reported issues + adds requested feature:

#7 vid (PORT not honored, MODEL_DIR vs MODELS_DIR confusion):
  - All 8 vLLM compose files now use "${PORT:-XXXX}:8000" so .env PORT
    flows through. Defaults preserved per-variant (8020 single, 8010-8013
    dual). llama.cpp composes already had this pattern.
  - scripts/switch.sh: load .env early; per-variant default-port table;
    new resolve_ready_url() picks PORT > variant default for the readiness
    probe.
  - scripts/launch.sh: same default-port table; final endpoint URL printed
    to user reflects actual mapped port.
  - .env.example: ⚠ box callout that variable names are CASE-SENSITIVE
    (MODEL_DIR singular, NOT MODELS_DIR plural — silently ignored).
    New PORT section documenting per-variant defaults.

#4 timxx (tools-text.yml fails "Free memory ... less than desired"):
  - docs/FAQ.md: new entry "Container fails to start: Free memory..."
    explaining the vLLM startup check, the two workarounds (free VRAM /
    lower mem-util), and which configs hit it most often (0.97+ mem-util).
  - Compose defaults unchanged (0.97 stays the right design target on
    headless rigs); the FAQ documents the workaround for users with X11.

#1 fabriciomalta (per-config VRAM column):
  - docs/SINGLE_CARD.md: TL;DR table now has VRAM column with mem-util.
  - docs/DUAL_CARD.md: TL;DR table same + footnote explaining per-card
    semantics and which dual configs would/wouldn't fit on 2× 20 GB cards
    (relevant to fabriciomalta's 2× 3080-20GB use case).

#2 tenitram (empty responses) — fixed in master via aab8ff4
(P68/P69 disabled). Closed with reply pointing at the fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-30 16:19:53 +00:00
noonghunnaandClaude Opus 4.7 48f93e550f docs: demote 48K/tools-text/minimal to fallback; lead with long-* + llama.cpp
User feedback: the small-ctx variants (48K default, tools-text 75K, minimal
32K) "offer very little context and not many people will find that as
viable options." Now that Cliff 1 is closed on the long-* variants via the
PN12 anchor sidecar, they're strictly more useful than the 48K/75K
alternatives for the workloads most users come for. The single residual
limitation is Cliff 2 on single-prompt >50K, addressed by llama.cpp.

SINGLE_CARD.md:
- TL;DR table reduced to 3 recommended options (long-vision · long-text ·
  llamacpp/default).
- Cliff 2 caveat promoted to a prominent ⚠️ callout right under the table —
  the one limitation users need to know.
- Old per-variant sections folded; small-ctx variants moved to an
  "Other variants in the repo" section as fallback / diagnostic.

scripts/launch.sh:
- Wizard leads with the 3 primary options (long-vision · long-text ·
  llamacpp/default). Diagnostic / niche options bundled at the end with a
  "[fallback]" prefix so they don't dominate the menu.

models/qwen3.6-27b/README.md:
- Single-card recommended-options bullet list now leads with the 3 primary
  variants and explicitly names Cliff 2 as the single shipped limitation.

docs/EXAMPLES.md:
- Cline section: stop pointing at tools-text; long-* now handle Cline's
  tool returns. Cliff 2 is the only remaining caveat to flag.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-30 13:21:14 +00:00
noonghunnaandClaude Opus 4.7 427d2f8aa9 docs+scripts+charts: propagate new ceilings (long-vision 198K, long-text 218K)
Sweep across all user-facing docs reflecting the post-PN12-anchor-fix
ceilings established in 287de1c → f3e5b52:

Docs touched:
- models/qwen3.6-27b/README.md — VRAM allocation paragraph + What's not
  working list + Genesis patches table (PN12/PN13/P101/P103 added)
- models/qwen3.6-27b/INTERNALS.md — forward-looking note pointing at
  CLIFFS.md for current state
- models/qwen3.6-27b/CHANGELOG.md — new 2026-04-30 PM entry
- docs/SINGLE_CARD.md — TL;DR table, VRAM budget bullet, frontier-context
  section, cliff status footer
- docs/FAQ.md — vLLM-vs-llama.cpp framing, ctx-drop question, Cliff 1/2
  explanations, troubleshooting list
- docs/engines/README.md — engine comparison table
- docs/engines/VLLM.md — feature bullets, TQ3 table, ctx tier description
- docs/engines/LLAMA_CPP.md — "why no cliffs" framing, single-card switch
  decision

Charts regenerated:
- tools/charts/gen-perf.py — labels updated (long-vision 198K, long-text 218K)
- tools/charts/gen-vram.py — added 218K text-only row, relabeled 198K row
  with mem-util note. SVG/PNG outputs regenerated via Docker matplotlib.

Scripts:
- scripts/launch.sh — wizard option labels
- scripts/switch.sh — header documentation

Cliff 1 status across all variants: closed.
Cliff 2 status: still applies single-prompt >50–60K on single-card.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-30 12:59:57 +00:00
noonghunnaandClaude Opus 4.7 37a4895f6d Remove fast-chat.yml; extend P68/P69 disable to default
fast-chat (20K, fp8, vision) and default docker-compose.yml (48K, TQ3,
vision) had effectively the same TPS post-PN8. fast-chat's only
remaining differentiator was "smaller context = ~3s faster boot," and
20K is actively bad for IDE-agent users (Copilot tool-schema preamble
alone hits 20K). Net negative — removed.

Default compose was missed in the previous P68/P69 fix — it had the
same env vars enabled and the same silent-stop bug above 8000 chars.
Both now disabled with the same explanatory comment.

Updated:
- scripts/switch.sh, scripts/launch.sh — drop the variant
- docs/SINGLE_CARD.md, FAQ.md, engines/VLLM.md, model + vllm + patches
  READMEs — references removed or pointed to default/tools-text
- All sibling compose YAML "see also" tables — fast-chat row removed,
  tools-text row repurposed for IDE-agent guidance
- CHANGELOG entry; old historical entries kept as-is (append-only)

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-29 19:29:10 +00:00
noonghunnaandClaude Opus 4.7 abc06c3e33 UX polish: pre-flight checks + cards-first wizard + PNG embeds
- scripts/preflight.sh (new) — sourceable library: docker, GPU >= N,
  disk free, GPU-idle warning, running-container note. Each error has
  an actionable Fix: hint instead of a cryptic mid-run crash.
- scripts/setup.sh + scripts/launch.sh wire pre-flight in early.
  launch.sh adds --no-preflight escape hatch.
- launch.sh wizard inverted: cards → workload → auto-pick engine.
  Newcomers can answer "how many GPUs" and "what do I want to do" but
  rarely "vLLM or llama.cpp" — engine falls out of the pick with a
  one-paragraph why. --engine override still works (filters the
  workload list to that engine).
- Embedded charts swapped SVG → PNG in README + SINGLE_CARD +
  DUAL_CARD + qwen3.6-27b/README. Clicking a PNG on GitHub opens a
  viewable image; SVGs open as raw XML. SVG remains the editable
  source — re-export PNG when SVG changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-29 14:29:19 +00:00
noonghunnaandClaude Opus 4.7 4b77ed5eb1 Add launch.sh wizard + switch.sh stateless variant switcher
scripts/switch.sh — stateless engine/variant switcher. Brings down
whatever's running (any vllm-qwen36-27b* or llama-cpp-qwen36-27b*
container — discovers compose file via docker labels), brings up
the new variant, waits for /v1/models to respond. Supports --list
(show all 13 variants) and --down (stop without booting). Variant
names are <engine>/<file-stem>: vllm/default, vllm/dual, vllm/
dual-turbo, vllm/dual-dflash, vllm/dual-dflash-noviz, vllm/long-
vision, vllm/long-text, vllm/fast-chat, vllm/tools-text, vllm/
no-genesis-mtp, vllm/minimal, llamacpp/default, llamacpp/concurrent.

scripts/launch.sh — interactive wizard for first-run users. Asks
engine → cards → workload, maps to variant, calls switch.sh, then
runs verify-full.sh to confirm clean serving. Also accepts flags
for non-interactive use:
  bash scripts/launch.sh --variant vllm/default
  bash scripts/launch.sh --engine vllm --cards 1   (asks the rest)

README.md — quick-start replaces "cd into compose dir + docker
compose up" with `bash scripts/launch.sh`. Click-throughs from the
launch tweet land in a guided flow instead of having to find the
right compose file by hand.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-28 22:09:12 +00:00