Files
club-3090/docs/EXAMPLES.md
T
noonghunnaandClaude Opus 4.7 5aa97a25d9 v0.20 migration + Genesis v7.65 dev tip + cold-start cache + env-var alignment
This branch migrates the entire vLLM stack from `dev205+g07351e088` + Genesis
v7.64 to `0.20.1rc1.dev16+g7a1eb8ac2` + Genesis v7.65 dev tip (commit
`d89a089`). v7.65 is on Sandermage's `dev` branch — explicitly the cross-rig
testing surface he requested in discussion #19; he'll merge dev→main once we
both confirm stable. Pin gates restated when that lands.

What changes
------------

Pin migration:
- vLLM image: nightly-07351e08... → nightly-7a1eb8ac2... (dev205 → v0.20.1rc1.dev16)
- Genesis: 64dd18b (v7.64) → d89a089 (v7.65 dev tip)

Sidecar churn:
- DROPPED: patch_pn12_ffn_pool_anchor.py (PN12 native on v0.20)
- DROPPED: patch_pn12_compile_safe_custom_op.py (Genesis P38B in-source hook)
- DROPPED: patch_fa_max_seqlen_clamp.py (Genesis PN17 + P15B)
- ADDED: patch_workspace_lock_disable.py (relaxes vllm#39226 strict assertion;
  P98 covers same surface but auto-skips on v0.20 due to drift-marker false
  positive — pending Sandermage marker fix)

Env-var alignment to Sandermage's PROD set (start_27b_int4_TQ_k8v4.sh@dev):
- FIXED naming bugs that silently no-op'd patches:
  - PN9_INDEPENDENT_DRAFTER_ATT → _ATTN (was silently OFF)
  - PN22 → PN22_LOCAL_ARGMAX_TP (was silently OFF)
  - PN26_BLOCK_KV → PN26_SPARSE_V_BLOCK_KV (fell back to default 4, not 8)
  - PN26_NUM_WARPS → PN26_SPARSE_V_NUM_WARPS
  - PN26_THRESHOLD → PN26_SPARSE_V_THRESHOLD (fell back to default 0.001, not 0.01)
- ADDED explicit-OFFs to match Sander's PROD verbatim:
  - P78_TOLIST_CAPTURE_GUARD=0 (we use our own patch_tolist_cudagraph.py)
  - P81_FP8_BLOCK_SCALED_M_LE_8=0 (FP8-specific, no-op on TQ3)
  - P82=0, P82_THRESHOLD_SINGLE=0.3
- Cap divergence (justified): PROFILE_RUN_CAP_M=4128 + PREALLOC_TOKEN_BUDGET=4128
  (Sander uses 4096 — vLLM `interface.py:639` forces our config's Mamba
  block_size to 4128 due to TQ3 + TP=1 page-size math; lower values
  AssertionError at boot)
- Carry-forward (intentional): P4 (hybrid TQ required), P65 (TQ spec-CG
  downgrade — pending v0.20 verification that #40880 closure makes it
  redundant)

Cold-start cache mounts (closes #22):
- All 10 composes now mount torch_compile_cache + Triton cache from
  `models/qwen3.6-27b/vllm/cache/`. First boot warms (~6 min); warm boot
  drops to ~3.2 min (47% faster). Per-stage savings on long-text:
  - Dynamo bytecode transform: 18s → 5s (-73%)
  - torch.compile: 57s → 9s (-85%)
  - Initial profiling/warmup: 51s → 7s (-87%)

Mamba block_size cap fix:
- v0.20 enforces `long_prefill_token_threshold >= block_size`; on hybrid
  Mamba+TQ3, vLLM forces block_size=4128. Bumped GENESIS_PROFILE_RUN_CAP_M
  and PREALLOC_TOKEN_BUDGET 4096→4128 across all 5 main composes.

Default 48K compose:
- Required workspace_lock_disable sidecar after initial v0.20 boot hit
  vllm#39226 strict assertion. Caught during validation, fixed.

Context restored vs dev205 backoffs (validated 33K + 50K stress on v0.20):
- long-text:        185K → 214K (+16%)
- long-vision:      140K → 198K (+41%)
- bounded-thinking: 185K → 214K (+16%)

Bench results (n=5, results/v0.20-migration/):
- long-text 214K        narr 49.74 / code 67.39 (CV 2.6/2.7%)
- long-vision 198K      narr 50.32 / code 66.12 (CV 2.3/4.1%)
- bounded-thinking 214K narr 49.77 / code 65.80 (CV 1.4/2.3%)
- tools-text 75K (fp8)  narr 53.32 / code 69.66 (CV 2.3/1.4%)
- dual-turbo 262K (TP=2) narr 58.33 / code 76.01 per-stream
                         269 TPS aggregate at n=4 streams (3.63x speedup)
- default 48K           narr 48.82 / code 65.98 (n=3)

Validation: verify-full 8/8 on every variant. verify-stress 33K AND 50K
tool-prefill PASS on every variant — the cliff that fired on EVERY dev205
config no longer reproduces.

Docs + charts:
- README + SINGLE_CARD + DUAL_CARD + CLIFFS + EXAMPLES + STRUCTURED_COT
  + FAQ + UPSTREAM + 3 engine docs + model README + INTERNALS + CHANGELOG
  all updated with new pin, ctx, TPS numbers, and "v0.20 unblock" section
- performance.{png,svg} + variants regenerated with measured TPS
- vram-budget.{png,svg} + variants regenerated with measured VRAM
- UPSTREAM tracker: 5 issues moved ✅ closed (PR #12, #13, #14, #15, P104
  superseded by PN17 + P15B)

Issues addressed:
- #16 (Cliff 1 mech B leaks past PN12 on inductor-compiled FFN) — partial:
  v0.20's revised TQ FA paths close the synthetic stress; PN25 (Sander's
  proper compile-path opaque-op fix) is on dev but explicitly opt-in pending
  worker-fork registration fix. Workarounds documented (tools-text fp8 path
  / --enforce-eager) until Sander ships PN25 default-on.
- #20 (launch.sh port + container-name mismatch) — already closed by
  77ca576 (post-issue-filing).
- #22 (cold-start caching) — closed by cache mounts above.

Remaining caveats:
- Cliff 2 (DeltaNet GDN forward, single prompt ≥50-60K) unchanged —
  architectural, applies to all single-card vLLM TQ3 paths. Mitigation:
  dual-turbo TP=2 (state splits across cards) or llama.cpp 262K.
- Default 48K narr_TPS (48.82) slightly under chart's 55 reference —
  bench variance + sample size n=3; not regression.
- Dual.yml / dual-dflash* not re-benched on v0.20; numbers carry forward
  from dev205 (fp8 paths were not TPS-changed by the migration).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-01 18:33:01 +00:00

10 KiB
Raw Blame History

Client examples

Copy-pasteable snippets for talking to the club-3090 endpoint. The default URL is http://localhost:8020; the served model name is qwen3.6-27b-autoround (vLLM) or qwen3.6-27b-autoround (llama.cpp via the --alias flag we set).

All examples assume:

  • Server running: bash scripts/launch.sh is up
  • API endpoint: http://localhost:8020 (override with OPENAI_BASE_URL env var or client-side base_url)

The endpoint is OpenAI-compatible — anything that speaks OpenAI's /v1/chat/completions API works without modification, just point base_url at the local endpoint.


max_tokens defaults — important if you've enabled thinking

Qwen3.6-27B is a thinking model. The <think>...</think> block before the answer routinely runs 2-4K tokens on medium reasoning, 4-8K on harder coding problems, and can exceed 16K on competition-grade problems. If max_tokens cuts the response off mid-think, the model never reaches the answer and the request looks like an "empty response" or truncated garbage. We hit this exact trap on our LiveCodeBench v6 baseline (docs/STRUCTURED_COT.md caveat section).

Use these defaults:

Scenario max_tokens
FREE thinking on (default long-text / long-vision composes) 8192 minimum. 16384 for hard reasoning / competition-grade problems.
FSM bounded thinking (bounded-thinking.yml) 4096 is fine — grammar caps the think block to a few hundred tokens of structured form.
enable_thinking: False Set as tight as the answer needs (50-200 typically).
Tool-using agents (multi-turn) 1024-2048 per turn. If a middle turn needs >2K to think, your prompt structure probably needs work.

The smoke-test examples below use max_tokens: 200 because they ask short questions where thinking + answer fits comfortably. Real workloads should follow the table above.


Quick curl sanity test

curl -sf http://localhost:8020/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.6-27b-autoround",
    "messages": [{"role": "user", "content": "Capital of France?"}],
    "max_tokens": 200
  }' | jq -r '.choices[0].message.content'

Expected response: a sentence containing Paris. The max_tokens: 200 headroom is intentional — Qwen3.6 thinks before answering by default, so even simple questions burn ~50–150 tokens inside <think>...</think> before reaching the answer. Set tighter (max_tokens: 30) only if you also pass chat_template_kwargs: {"enable_thinking": false} to skip the think block — that's what verify-full.sh does internally.


pip install openai

Basic chat

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8020/v1", api_key="not-needed")

resp = client.chat.completions.create(
    model="qwen3.6-27b-autoround",
    messages=[{"role": "user", "content": "Write a haiku about tensor cores."}],
    max_tokens=120,
    temperature=0.6,
    top_p=0.95,
)
print(resp.choices[0].message.content)

Streaming

stream = client.chat.completions.create(
    model="qwen3.6-27b-autoround",
    messages=[{"role": "user", "content": "Explain attention in 100 words."}],
    max_tokens=300,
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content or ""
    print(delta, end="", flush=True)
print()

Tool calling

Works on both engines (vLLM with --tool-call-parser qwen3_coder and llama.cpp with --jinja — both ship enabled in the default composes):

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather in a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }
]

resp = client.chat.completions.create(
    model="qwen3.6-27b-autoround",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
    tool_choice="auto",
    max_tokens=200,
)

msg = resp.choices[0].message
if msg.tool_calls:
    for tc in msg.tool_calls:
        print(f"Call {tc.function.name}({tc.function.arguments})")
else:
    print(msg.content)

Vision (image input)

vLLM and llama.cpp both auto-load the vision tower / mmproj when configured. Send images as base64 or URLs:

import base64
from pathlib import Path

img_b64 = base64.b64encode(Path("photo.png").read_bytes()).decode()

resp = client.chat.completions.create(
    model="qwen3.6-27b-autoround",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe this image."},
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{img_b64}"},
                },
            ],
        }
    ],
    max_tokens=200,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(resp.choices[0].message.content)

Reasoning mode (vLLM with Genesis only)

resp = client.chat.completions.create(
    model="qwen3.6-27b-autoround",
    messages=[{"role": "user", "content": "Solve: 7x + 14 = 49. Show your reasoning."}],
    max_tokens=2048,  # FREE thinking on; 2048 fits easy math comfortably. Bump to 8192 for harder reasoning.
    extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
msg = resp.choices[0].message
print("Reasoning:", getattr(msg, "reasoning_content", "") or "(empty)")
print("Answer:   ", msg.content)

Note: llama.cpp emits the <think>...</think> tokens inline rather than peeling them into a separate reasoning_content field. If you need that split, post-process client-side or stick with vLLM.


Python — requests (no SDK)

For environments where you can't install the openai package:

import requests, json

resp = requests.post(
    "http://localhost:8020/v1/chat/completions",
    headers={"Content-Type": "application/json"},
    json={
        "model": "qwen3.6-27b-autoround",
        "messages": [{"role": "user", "content": "What is 17 × 23?"}],
        "max_tokens": 50,
    },
    timeout=60,
)
print(resp.json()["choices"][0]["message"]["content"])

For streaming, use stream=True and parse SSE lines:

with requests.post(
    "http://localhost:8020/v1/chat/completions",
    headers={"Content-Type": "application/json"},
    json={"model": "qwen3.6-27b-autoround", "messages": [...], "stream": True, "max_tokens": 200},
    stream=True,
) as r:
    for line in r.iter_lines():
        if not line or not line.startswith(b"data: "):
            continue
        payload = line[6:]
        if payload == b"[DONE]":
            break
        chunk = json.loads(payload)
        delta = chunk["choices"][0]["delta"].get("content", "")
        print(delta, end="", flush=True)

TypeScript / Node — openai SDK

npm install openai
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:8020/v1",
  apiKey: "not-needed",
});

const resp = await client.chat.completions.create({
  model: "qwen3.6-27b-autoround",
  messages: [{ role: "user", content: "Quicksort in Rust, please." }],
  // FREE thinking is on by default. 4096 covers easy code-gen think+answer;
  // 8192 is the safe default for harder coding problems. 800 traps mid-think.
  max_tokens: 4096,
  temperature: 0.6,
  top_p: 0.95,
});

console.log(resp.choices[0].message.content);

Streaming:

const stream = await client.chat.completions.create({
  model: "qwen3.6-27b-autoround",
  messages: [{ role: "user", content: "..." }],
  max_tokens: 300,
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");

Connecting third-party clients

Open WebUI

Settings → Connections → Add OpenAI Connection:

  • Base URL: http://localhost:8020/v1 (or http://<host-ip>:8020/v1 from another machine on your LAN — see Security)
  • API Key: anything (e.g. sk-local) — the server doesn't check it
  • Model: qwen3.6-27b-autoround

Vision, tool calling, streaming all work through the WebUI's standard flows.

Cline / Roo (VS Code agentic coder)

In the Cline settings panel:

  • API Provider: OpenAI Compatible
  • Base URL: http://localhost:8020/v1
  • API Key: sk-local (any non-empty string)
  • Model ID: qwen3.6-27b-autoround

Cline sends large tool returns (file reads, web fetches) up to ~25K tokens. As of 2026-05-01 PM (vLLM v0.20 + Genesis v7.65 dev tip migration), both vllm/long-vision (198K + vision) and vllm/long-text (214K text-only) handle these cleanly. Both 33K-token AND 50K-token tool-prefill stress now PASS on master. The only remaining caveat: don't use vLLM single-card for one-shot prompts >50K (Cliff 2 — DeltaNet GDN forward — still applies; switch to llamacpp/default or dual-turbo.yml instead). See docs/SINGLE_CARD.md and the VRAM diagram.

Cursor

Settings → Models → Add OpenAI-compatible:

  • Override OpenAI Base URL: http://localhost:8020/v1
  • Verify config: click "Verify" — should list qwen3.6-27b-autoround
  • Model name: qwen3.6-27b-autoround

Cursor's "Apply" feature works against this model since tool-calling is supported.

LiteLLM proxy / aider / Continue.dev

All work the same way — OpenAI-compatible endpoint at http://localhost:8020/v1, any non-empty API key. Confirmed working with the default compose.


Security note: network binding

The default composes bind to 0.0.0.0:8020 so other machines on your LAN can connect. If you're on a shared / coffee-shop / coworking network, that exposes your model to anyone who can route to your machine.

To restrict to localhost-only:

# In any docker-compose.*.yml under ports:
ports:
  - "127.0.0.1:8020:8000"   # was: "8020:8000"

Or override at run-time with --host 127.0.0.1 (llama.cpp) / by editing the compose locally.


See also