botty_next — visual test harness
Bootstrapped in commit 48d8445. A standalone, importable package (botty_next/) that
exercises the vision primitives — screen capture, template matching, OCR — in isolation from
the live bot, so they can be validated against fixtures in CI without D2R running.
It does not drive the game. Input is gated off by default and there is no live-input path yet
(see InputConfig below). Think of it as a test bench for the perception layer that the legacy
src/ bot will eventually be ported onto.
Layout
botty_next/
cli.py argparse entry point: config / detect / capture / ocr
capture/
window.py WindowRegion + find_window_region() (win32gui enumerate)
mss_backend.py MssCaptureBackend.grab() -> BGR ndarray; save_frame()
vision/
fixtures.py load_image / load_screenshot / load_template (cv2.imread)
template_matching.py match_template() -> MatchResult; save_match_debug()
ocr.py preprocess_for_ocr(), run_tesseract_ocr() -> OcrResult
config/
models.py pydantic config models + load_config()
default.yaml default profile
tests/ pytest suite for each module
debug/ debug-image output dir
CLI
python -m botty_next.cli <command> (entry: cli.main). Every command prints a JSON result and
returns a process exit code.
| Command | Args | Does | Exit code |
|---|---|---|---|
config validate |
-c/--config PATH |
Loads + validates a YAML profile, prints the resolved config | 0 |
detect template |
--image --template [--threshold 0.85] [--debug-output] |
Runs match_template, optionally writes an annotated debug image |
0 if passed, else 1 |
capture |
--output [--window-title] |
Grabs a frame (full monitor, or the matched window region) and saves it | 0 |
ocr |
--image [--lang eng] [--psm 6] [--tesseract-cmd] [--debug-output] |
OCRs an image; writes the preprocessed debug image first if requested | 0 ok / 2 if pytesseract missing |
Core logic
Capture (capture/)
find_window_region(title_contains)enumerates visible top-level windows viawin32gui, case-insensitively substring-matches the title, and returns the largest match as aWindowRegion(left, top, width, height, title). Raises if none found.WindowRegion.as_mss_monitor()adapts it to the dictmssexpects.MssCaptureBackend.grab(region)grabs that region (ormonitors[1]= primary monitor whenregion is None) and converts the raw BGRA to BGR so it matches OpenCV's convention.save_frame()creates parent dirs and writes viacv2.imwrite, raising on failure.
Template matching (vision/template_matching.py)
match_template(image, template, threshold=0.85, method=TM_CCOEFF_NORMED):- Validates non-empty inputs and that the template isn't larger than the image.
- Converts both to grayscale, runs
cv2.matchTemplate+cv2.minMaxLoc. - For
TM_SQDIFF*methods the min location wins andconfidence = 1 - min_val; for all other methods the max location wins andconfidence = max_val. This normalizes so "higher confidence = better" regardless of method. - Returns a frozen
MatchResult(confidence, bbox, passed, method, debug)wherepassed = confidence >= thresholdanddebugcarries the raw min/max values and shapes.
save_match_debug()draws the bbox green if passed, red if not, and writes it.
OCR (vision/ocr.py)
preprocess_for_ocr(image, scale=2.0): grayscale → 2× upscale (INTER_CUBIC) → Gaussian blur → adaptive Gaussian threshold (block 31, C 7). This is the single source of truth for OCR preprocessing — both the OCR run and the debug image use it.run_tesseract_ocr(...): lazily importspytesseract(raising aRuntimeErrorwith install guidance if absent), optionally setstesseract_cmd, runsimage_to_stringfor text andimage_to_datafor per-word confidences, and returnsOcrResult(text, confidence, bbox, debug). Confidence is the mean of word confidences (each normalized 0–1, negatives dropped).
Config (config/models.py)
Pydantic models with extra="forbid" (unknown keys are rejected). load_config(path) reads YAML
and validates it into BottyNextConfig:
CaptureConfig—backend(fixture/mss/dxcam, defaultfixture),monitor,fps_limit(1–240), optionalwindow_title.VisionConfig—template_threshold(0–1),debug_output_dir.InputConfig— safety gate:enabled=False,dry_run=Trueby default. A model validator onBottyNextConfigraises ifinput.enabledis true whiledry_runis false — i.e. live input is impossible until a future explicit safety gate is added.
Fixtures
fixtures/screenshots/sample_scene.ppm and fixtures/templates/sample_marker.ppm are committed
(PPM so they diff/version cleanly) and let the test suite run with no external assets.
fixtures/ocr_samples/.gitkeep reserves the OCR sample dir.
Tests
botty_next/tests/ has one module per concern (test_capture, test_cli, test_config,
test_fixtures, test_ocr, test_template_matching). They run against the committed fixtures,
so the harness is CI-safe without a display or a running game.