Add docs/FAQ.md — common questions answered for tweet click-throughs
Topics: hardware (4090 / 5090 / NVLink / non-NVIDIA / Windows-WSL2),
engine choice (vLLM vs llama.cpp / why not Ollama-LMStudio / MTP not
EAGLE / why not GGUF on vLLM / why AutoRound), performance (TPS
expectations, ctx-load decode drop, prefill cliffs explained,
vllm#40914), setup (model paths, GPU index override, multi-variant
ports, Open WebUI), community (bench contributions, bug reports,
Genesis bumping).
Linked from top-level README. Designed to absorb repeat issue-tracker
questions; each answer is 2-4 sentences with links to deeper docs.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>