# botty_next — visual test harness Bootstrapped in commit `48d8445`. A standalone, importable package (`botty_next/`) that exercises the vision primitives — screen capture, template matching, OCR — in isolation from the live bot, so they can be validated against fixtures in CI without D2R running. It does **not** drive the game. Input is gated off by default and there is no live-input path yet (see `InputConfig` below). Think of it as a test bench for the perception layer that the legacy `src/` bot will eventually be ported onto. ## Layout ``` botty_next/ cli.py argparse entry point: config / detect / capture / ocr capture/ window.py WindowRegion + find_window_region() (win32gui enumerate) mss_backend.py MssCaptureBackend.grab() -> BGR ndarray; save_frame() vision/ fixtures.py load_image / load_screenshot / load_template (cv2.imread) template_matching.py match_template() -> MatchResult; save_match_debug() ocr.py preprocess_for_ocr(), run_tesseract_ocr() -> OcrResult config/ models.py pydantic config models + load_config() default.yaml default profile tests/ pytest suite for each module debug/ debug-image output dir ``` ## CLI `python -m botty_next.cli ` (entry: `cli.main`). Every command prints a JSON result and returns a process exit code. | Command | Args | Does | Exit code | |---------|------|------|-----------| | `config validate` | `-c/--config PATH` | Loads + validates a YAML profile, prints the resolved config | 0 | | `detect template` | `--image --template [--threshold 0.85] [--debug-output]` | Runs `match_template`, optionally writes an annotated debug image | 0 if `passed`, else 1 | | `capture` | `--output [--window-title]` | Grabs a frame (full monitor, or the matched window region) and saves it | 0 | | `ocr` | `--image [--lang eng] [--psm 6] [--tesseract-cmd] [--debug-output]` | OCRs an image; writes the preprocessed debug image first if requested | 0 ok / 2 if pytesseract missing | ## Core logic ### Capture (`capture/`) - `find_window_region(title_contains)` enumerates visible top-level windows via `win32gui`, case-insensitively substring-matches the title, and returns the **largest** match as a `WindowRegion(left, top, width, height, title)`. Raises if none found. - `WindowRegion.as_mss_monitor()` adapts it to the dict `mss` expects. - `MssCaptureBackend.grab(region)` grabs that region (or `monitors[1]` = primary monitor when `region is None`) and converts the raw BGRA to **BGR** so it matches OpenCV's convention. - `save_frame()` creates parent dirs and writes via `cv2.imwrite`, raising on failure. ### Template matching (`vision/template_matching.py`) - `match_template(image, template, threshold=0.85, method=TM_CCOEFF_NORMED)`: - Validates non-empty inputs and that the template isn't larger than the image. - Converts both to grayscale, runs `cv2.matchTemplate` + `cv2.minMaxLoc`. - For `TM_SQDIFF*` methods the **min** location wins and `confidence = 1 - min_val`; for all other methods the **max** location wins and `confidence = max_val`. This normalizes so "higher confidence = better" regardless of method. - Returns a frozen `MatchResult(confidence, bbox, passed, method, debug)` where `passed = confidence >= threshold` and `debug` carries the raw min/max values and shapes. - `save_match_debug()` draws the bbox green if passed, red if not, and writes it. ### OCR (`vision/ocr.py`) - `preprocess_for_ocr(image, scale=2.0)`: grayscale → 2× upscale (`INTER_CUBIC`) → Gaussian blur → adaptive Gaussian threshold (block 31, C 7). This is the single source of truth for OCR preprocessing — both the OCR run and the debug image use it. - `run_tesseract_ocr(...)`: lazily imports `pytesseract` (raising a `RuntimeError` with install guidance if absent), optionally sets `tesseract_cmd`, runs `image_to_string` for text and `image_to_data` for per-word confidences, and returns `OcrResult(text, confidence, bbox, debug)`. Confidence is the mean of word confidences (each normalized 0–1, negatives dropped). ### Config (`config/models.py`) Pydantic models with `extra="forbid"` (unknown keys are rejected). `load_config(path)` reads YAML and validates it into `BottyNextConfig`: - `CaptureConfig` — `backend` (`fixture`/`mss`/`dxcam`, default `fixture`), `monitor`, `fps_limit` (1–240), optional `window_title`. - `VisionConfig` — `template_threshold` (0–1), `debug_output_dir`. - `InputConfig` — **safety gate**: `enabled=False`, `dry_run=True` by default. A model validator on `BottyNextConfig` raises if `input.enabled` is true while `dry_run` is false — i.e. live input is impossible until a future explicit safety gate is added. ## Fixtures `fixtures/screenshots/sample_scene.ppm` and `fixtures/templates/sample_marker.ppm` are committed (PPM so they diff/version cleanly) and let the test suite run with no external assets. `fixtures/ocr_samples/.gitkeep` reserves the OCR sample dir. ## Tests `botty_next/tests/` has one module per concern (`test_capture`, `test_cli`, `test_config`, `test_fixtures`, `test_ocr`, `test_template_matching`). They run against the committed fixtures, so the harness is CI-safe without a display or a running game.