Both self-restart paths (Bot.restart_or_exit, GameController "could not
recover") spawn a replacement process that waited for the resume key like a
manual launch. With nobody at the keyboard every self-recovery became a
permanent stall - the "bot is stuck all the time" report on 2026-09-15.
- Replacement processes get BOTTY_AUTOSTART=1; main.py starts the bot when it
is set and D2R is running. A manual launch still waits for the key.
- The controller path now shares Bot's restart cap (reset on reaching town),
so an unrecoverable screen stops after 5 restarts instead of looping.
Co-Authored-By: Claude Opus 5 <[email protected]>
Many tests run real bot code paths with only the screen faked; those paths end in
win_input's SendInput/SetCursorPos, which act on the focused window. Running the
suite while the bot played on 2026-09-15 sent test clicks and keys into D2R (a
game started with no bot input, the inventory opened, the character got
stranded), and the bot was blamed for getting stuck.
conftest now wraps win_input.user32 so SendInput/SetCursorPos/keybd_event/
mouse_event report success and do nothing, and no-ops utils.misc window
SetWindowPos/SetForegroundWindow/ShowWindow. Verified: a test calling
win_input.mouse_move leaves the real cursor in place; with
BOTTY_TESTS_ALLOW_REAL_INPUT=1 the same test moves it.
Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-15 16:33: a game appeared while start_game was still waiting for PLAY.
It ran on D2R's remembered difficulty - NORMAL, left there by the baal_xp
leech - so the Hell A3 Travincal character spawned in Act 1. start_game
returned False, on_create_game ignored the result, and every A1 NPC/WP step
failed; two games hit the maintenance timeout and D2R went down.
- _wait_for_play_btn returns IN_GAME (not None) for that case; start_game
save+exits it and creates a game with the configured difficulty key.
Being in town on ENTRY is still accepted (baal_xp re-entry relies on it).
- on_create_game no longer continues after a failed start_game.
Co-Authored-By: Claude Opus 5 <[email protected]>
Resolves the 11 conflicting files blocking PR #41. Where both branches fixed
the same problem, the fixes are combined rather than one side dropped:
- a5.py: main's dual WP templates + require_visible pre-scan, bounded by
baalxp's town_world ROI (excludes the HUD band and party column)
- pindle.py: main's removal of the town-false-positive "already in temple"
shortcut (also dropped baalxp's post-nudge copy of it); kept baalxp's nudge,
click tuning and PORTAL_APPROACH node; main's detect-act retry
- town_manager.py: main's act-margin check in wait_for_town_spawn, then
baalxp's marker->location resolution; baalxp's in_act repair skip before
main's repair_npc destination table
- bot.py: baalxp's restructured leech cycle, with main's death recovery and
InGame miss tolerance ported into _baal_xp_hide_and_wait
- game_browser.py, skill_preflight.py, [hammerdin] binds: baalxp (lobby join
flow; binds verified against the live client 2026-09-04)
- params.ini: baalxp's "leech stays out of order=" rule with order=run_trav
- CLAUDE.md: both bug writeups kept; baalxp's 23-25 renumbered to 32-34 since
tests/docs reference main's numbering
- test_stealth_config.py: AFK-break invariant exempts the whole baal_xp block
instead of a 6-line window (the branch grew several handoffs)
Tests: 272 passed. Remaining failures are pre-existing on origin/main
(gems-tab transmute test, smoke_test/test_version_consistency collection).
Co-Authored-By: Claude Opus 5 <[email protected]>
Two bugs that combined to hang the bot for hours after the 5-fail breaker
fired (observed 2026-09-15: last log entry 08:49, bot still hung at 12:54):
1. go_to_hero_selection(): the inner 'while is_loading' loop had no timeout.
If D2R was stuck on a loading screen that never resolved, the bot spun
forever. Added a 30s bound to the inner wait so the outer 45s timeout
can actually fire.
2. _baal_xp_rearm + on_maintenance: the 5-fail breaker fired (reset the
counter, logged the warning) but on_maintenance immediately re-armed the
leech because _baal_xp_only() is always True. The breaker was a no-op —
the fail-rearm-fail loop continued indefinitely. Now the breaker sets a
flag that on_maintenance checks: instead of re-arming, it triggers
end_game so the controller can restart D2R with fresh config.
The state list was missing baal_xp (and cold_plains) even though the
route has been in bot.py since August. One short paragraph on how the
cycle works, since it deliberately bypasses the run_wrapper pattern.
The paladalla profile sets run_baal_xp=1 in [routes] but has no [baal_xp]
section, so the leech enable flag resolves to the params.ini default (1).
The 09-08 '5 cycles failed in a row' breaker was the join failing after the
fallback candidate was picked - the cycle then never re-armed and the bot
fell through to end_game. The rearm loop (98f9b73) handles that; this just
documents where the flag actually comes from so the next debug session does
not chase a phantom profile override.
The leech-only path in on_init fell through to the generic marker scan,
which triggered on_start_from_town and its 30s wait_for_town_spawn().
When no town marker was visible (character standing away from spawn,
or a stale template) the wait burned the full 30s, the controller's
cooperative_shutdown (8s, no force-kill) gave up on the bot thread,
D2R restarted, and the next on_init found the character already in a
game and left it again - a fresh game burned per cycle.
The leech does its own setup inside the cycle (corpse pickup, /nopickup,
pre-buff, hide spot), so the full on_start_from_town routine is dead
weight for it. Add a skip_to_baal_xp transition (initialization ->
baal_xp) that mirrors skip_to_level, and have on_init's baal-only
branch trigger it directly after a short (8s) spawn scan with act
detection as fallback.
Also shorten _recover_to_own_game's wait_for_town_spawn to the same
8s budget so a failed recovery doesn't stall the cycle the same way.
The stash taken before switching to main (stash@{1}) carried a re-verified
_try_join - row click retried up to 3x with a re-OCR of the row when the list
scrolled, JOIN GAME press retried up to 3x when the lobby swallowed it, and a
join_game() default of 900s - but the switch was stashed, never re-applied, and
the branch tip kept the single-click version from 09-02. A single missed click
(or a list that scrolled between scan and click) then failed the whole join,
which is exactly the 'no matching game found' the cycle kept hitting.
Also add the finally: start_detecting_window() the 09-08 version had dropped:
join_game() stops the window-detection thread at the top, and every return
path (success, failure, exception) has to re-arm it or screen tracking stays
dead for the rest of the cycle.
_baal_xp_rearm's 5-strike breaker calls _game_stats.reset_consecutive_fails();
the fixture's object.__new__(Bot) has no _game_stats, so the breaker test
crashed with AttributeError instead of asserting the flag stays down.
- leech cycle failures no longer feed the session consecutive-fail tally
(five bad joins disabled the route for the whole session and ended the
game); _baal_xp_rearm keeps its own local breaker and resets the
game-level counter so the controller restarts D2R and the leech retries
- on_start_from_town: a leech that re-enters its own game on a corpse
(died in the public game, or save-and-exit raced the death screen)
gets pre_buff + /nopickup here - without it the char stays dead and the
next cycle's save-and-exit has no living game to exit from
The leech cycle leaves the bot's own game for a public one and comes back.
The controller's max_game_length breaker counts real game time, so the
clock can already be over budget the moment the cycle ends - and it would
kill the game before the next cycle starts. Reset it in on_end_run for
baal_xp.
Also: on char-select failure in the join phases, go home before
end_run (the character is stranded on the character screen otherwise),
and grade the cycle by whether recovery to the own game succeeded -
a cycle that left early (low HP, timer) is a success, only a broken
cycle (join/leave/recovery failure) is a failure.
validate_build_skill_icons rejected any required check with a blank hotkey before
the icon was ever looked at, so moving Blessed Hammer to its real home (left slot,
no key) traded one false failure for another. SkillCheck gains a `permanent` flag
for skills that sit on a slot with no hotkey selecting them; the guard skips those
and the icon check inspects the slot as-is.
Preflight now passes clean: blessed_hammer 100% (left), concentration 86.5%,
redemption 92.0%, holy_shield 93.0%, teleport 92.4%.
Co-Authored-By: Claude Opus 5 <[email protected]>
Probed the live client by pressing F1-F8 and reading the right skill slot; D2R
renders the key label into the icon, so the crops are self-proving. Actual binds:
f2=Holy Shield, f3=Redemption, f4=a charged item skill, f5=TELEPORT,
f6=Concentration, f8=an aura with no template (Conviction). f1 and f7 are UNBOUND.
Against that, [hammerdin] was wrong three ways:
- conviction=f5 pointed at TELEPORT, so every attack-aura cast would have
teleported the character — the exact hazard the [fohdin] profile section already
warns about. Conviction also does nothing for magic-damage hammers (Bug 17), so
it is now unbound rather than remapped.
- concentration=f8 pointed at Conviction; Concentration is on f6. The build was
running with no Concentration at all — less hammer damage and no party aura for
the merc, which plausibly fed the Travincal chickens.
- blessed_hammer=f1 pointed at nothing. The hammer lives permanently on LEFT-click
and no hotkey moves it to the right slot, so the preflight's right-slot check
could never pass; it scored 44.8% all session while the hammer sat correctly on
the left the whole time. The check now inspects the left slot and presses no key,
matching the comment that was already above it.
The stored blessed_hammer.png was a right-slot capture with "F1" baked into the
image, which is why it scored 34.4% against the unlabelled left slot. Recaptured
from the live left slot.
Co-Authored-By: Claude Opus 5 <[email protected]>
take_rejuv_potion_health is 0.45 but chicken is 0.40, so a failed rejuv drink
ended the game at 41-45% HP — above the configured chicken line, and with health
potions still in the belt; in every observed case one had been drunk successfully
1-3 seconds earlier. Rejuvs cannot be bought, only dropped, so the column runs dry
routinely (rejuv=4, a fully empty column, recurs in the Consumables needs line).
This chickened 3 of the first 15 games of the 11:39 session — a 20% failure rate,
and the dominant failure mode once the town bugs were fixed.
Now a missing rejuv falls back to a health potion, bypassing lp_hp_potion_delay
(that cooldown avoids wasting potions; it should not withhold emergency healing),
and chickens only when the belt is genuinely dry. The real chicken threshold check
directly below is untouched and still fires at <=40% HP.
Co-Authored-By: Claude Opus 5 <[email protected]>
_baal_xp_reenter_own_game waited the full 30s default in wait_for_town_spawn when
no town marker was visible. The controller's cooperative_shutdown (8s, no force
kill) then gave up on the bot thread and restarted D2R, so on_init found the
character already in a game and left it again — a fresh game burned every cycle.
The spawn scan is the slow part, so it now runs with an 8s budget and falls back
to detect_current_act, which is what on_start_from_town does anyway.
The working-tree version of this change had commented out the docstring body but
deleted its closing triple-quote, so the docstring swallowed the whole function
and src/bot.py would not parse — the bot could not start at all. Restored.
Co-Authored-By: Claude Opus 5 <[email protected]>
on_maintenance: when run_baal_xp is the only route and nothing is armed,
re-arm the leech and go again instead of falling through to end_game.
The old path handed off to the controller, which restarted D2R into a
brand new game that the next on_init immediately left again - burning a
game and a cycle for nothing. The _BAAL_XP_MAX_CONSECUTIVE_FAILS breaker
in _baal_xp_rearm still hands off to end_game after enough real failures.
on_end_run: a pure leech ends its cycle on the character screen, not in
a town - _verify_town_location() would burn template searches for a town
that is not there. Re-enter the own game and go straight to maintenance.
params.ini: document that run_baal_xp must stay out of order= (the
profile layer re-arms it between cycles).
Measured over the 65-game session of 2026-09-04 (62 successes, 3 failures, all
three "Maintenance timeout ~300s before [gamble]"):
1. open_npc_menu's body+pose path is a guess, not a confirmation. For Ormus it
opened the dialogue 5/73 times (7%); the name-tag path hit 37/47 (79%). Body
score cannot separate them — failures averaged 0.46, the 5 successes were
0.40-0.46, and one false positive scored 0.79 — so there is nothing to
threshold on. Each guess also paid a retry click plus two 2.5s waits. 68 dead
clicks is ~8 minutes of town time per session and is what drove the timeouts.
Guesses are now budgeted (2 per call) and no longer retried, so the search
reaches the grid sweep, which does find the NPC, with budget left to do it.
2. TOWN_MARKERS carried A3_TOWN_0 and A3_TOWN_1 for Act 3. Against 8 Kurast Docks
screenshots they score 0.21 and 0.29 — the worst two of all 20 a3_town
templates, against a 0.68 match threshold. Act detection in A3 therefore could
not succeed: 35 "no town marker found" and 24 "_verify_town_location: act
detection failed" in one session, which is the Bug 9 desync trigger. Added
A3_TOWN_14 (0.95) and A3_TOWN_20 (0.74); both are act-unique, scoring <=0.35
against A4 and A5 frames, so they cannot misfire.
Co-Authored-By: Claude Opus 5 <[email protected]>
A charm pickit rejects gets keep=False, but protect_charms_from_sell blocks both
the sell and the drop, and the stash branch only transfers keep==True. The item
could therefore leave the inventory by no route at all: the same 8 charms were
re-hovered, re-OCR'd and re-rejected every game (~12s of town time each cycle),
and every new junk charm picked off the ground joined the set permanently.
_without_protected now flips a blocked item to keep=True so the stash step takes
it, mirroring the recovery inspect_inventory already does per item at the vendor.
bot.py recomputes keep_items after buy_consumables, since it is evaluated before
that call and would otherwise stay stale and skip the stash entirely.
The protection intent is preserved: a misread charm is still never sold or
dropped, it just goes to the stash rather than clogging the inventory.
Co-Authored-By: Claude Opus 5 <[email protected]>
The Travincal route leaves the character in A3 Kurast Docks, which has no repair
vendor, so TownManager.repair() fell through to "going to A5 for repair" and took
a waypoint round trip every runs_per_repair games. Larzuk detection then failed,
the A4 Halbu fallback took a second trip, and the character was left stranded in
the wrong act — which is what turned a repair into 5 minutes of maintenance and
open_wp failures in the following games.
repair_npc=in_act (also current_act/none/off) now means: repair only if the
current act has a vendor, otherwise skip and stay put. The paladalla profile
(untracked) is set to it.
Co-Authored-By: Claude Opus 5 <[email protected]>
assets/templates/a5_wp.png (commit 14f5876) was a crop of the blue mana globe,
not the Harrogath waypoint. It matched the HUD at 0.98-0.99 on every frame, so
A5.open_wp() teleported the character to the bottom-right of the screen and
clicked there; WaypointLabel never appeared and each anchor retry pushed the
character further out until it was stranded on top of the town wall. From there
Larzuk, Qual-Kehk and every town marker were off-screen, which cascaded into
repair/resurrect failures and 5 consecutive run_trav approach failures that
self-disabled all routes.
- restore the pre-14f5876 template (201x120, the real stone arch); it scores
~0.30 on the stranded-on-the-wall screenshots instead of 0.99
- scan only Config().ui_roi["town_world"], the already-defined ROI that excludes
the HUD band and the party/chat column, so a bad template can never again move
the character into the globes
- add an roi parameter to IChar.select_by_template() to support the above
Co-Authored-By: Claude Opus 5 <[email protected]>
Two config-layer gaps stopped the leech from ever completing a cycle:
- _do_runs built 'run_baal_xp' from routes.get('run_baal_xp'), which raised
KeyError when the key was absent (a profile-level [routes] section that
lists only its own runs, e.g. paladalla's run_pindle/run_trav/run_diablo,
does not define it). The exception escaped the Bot constructor and took
the whole bot down at startup. Now .get(key, False).
- _baal_xp_rearm only re-set the _do_runs flag. on_maintenance re-reads
Config().routes for its 'any run left?' check, so the leech fell through
to end_game after every cycle and the controller created a brand new game
purely to leave it again. Re-arm Config().routes too. The matching
run_baal_xp=1 entry (deliberately NOT in order= - the leech is
self-contained and must not sit between real runs) goes in the
gitignored profile ini.
Co-Authored-By: Claude Opus 5 <[email protected]>
* fix: raise maintenance budget for cross-act routes; document Trav findings
Travincal leaves the character in Act 3, but maintenance relocates it to
Act 4 (stash at a4_tyrael_stash, repair at a4_halbu) and the next run then
needs the A3 waypoint again -- two cross-act waypoint trips per game.
That travel is not attributed to any timed step, so a FAIL> trail can show
~38s of work inside a 250s maintenance window. The 240s budget (tuned when
Pindle kept all town business in one act) was tripped 5 times in one
session, each costing a whole game. Raised to 420s.
Also documents, in HANDOVER.md:
- The open_wp failures are NOT a stale template. Scored a5_wp.png against
all three real failure frames: 0.96-0.98 full-frame, but that match is
the belt/mana-orb HUD false positive the a5.py comment already documents.
Inside the real cut_skill_bar ROI it scores 0.436-0.480 against a 0.55
threshold at scattered positions -- noise. The waypoint is genuinely not
on screen; the character never reaches it. Recapturing would fix nothing.
Scoring a template without the ROI the code actually uses produces a
confident wrong answer.
- PR #39's pather abort is merged, live, and finally firing (1 abort in 21
games, after being 0-for-152 while it looked correct).
- Keeping Trav town business in Act 3 is the real fix for both the timeouts
and the waypoint failures. a3.py already reports can_buy_pots/can_heal/
can_stash as True, so the Act 4 trip is not a missing capability.
- FoH vs hammerdin on the council, and the [fohdin]/[paladin] section trap:
concentration and redemption lived in [fohdin], which hammerdin does not
read, so both would have gone silently unbound after a respec.
Co-Authored-By: Claude Opus 5 <[email protected]>
* docs: post-mortem for the hell Travincal / hammerdin session
Narrative write-up of 29 Aug 2026: run_trav on hell, FoHdin -> hammerdin.
Outcome: health chickens 3-in-8 -> 0-in-11, battle 64-68s -> 41-60s,
failure rate 12% -> 9%, zero deaths throughout.
The document covers the four times the obvious answer was wrong:
1. Battle time did not move after the respec, which looked like failure.
hammerdin.kill_council casts for a FIXED duration with no kill detection,
so the clock cannot report damage -- loot proved the respec worked.
2. a5_wp.png scored 0.96-0.98 full-frame on every failure frame, which reads
as "template fine, search broken". That match is the belt/mana-orb HUD
false positive the a5.py comment already documents. Inside the real
cut_skill_bar ROI it scores 0.436-0.480 against a 0.55 threshold at
wandering positions -- noise. The waypoint was never on screen.
3. Ormus: not masked (4-channel but fully opaque, so no mask is passed), and
not a threshold problem (in-ROI best 0.342, background level -- lowering
the bar would recreate the Bug 4 false-positive clicks). A3 town has
frames with ZERO pather landmarks over threshold, so the character never
arrives. One cause, three symptoms.
4. Reverses this session's own earlier recommendation to move Trav town
business into Act 3. That was reasoned from a3.py reporting can_buy_pots/
can_heal/can_stash as True -- a capability check, not evidence about
pathing. A3 is the worst-supported town in the project.
Also documents three settings that were live but inert: binds under [fohdin]
which hammerdin never reads, casting_frames=8 (unreachable -- a paladin's
fastest cast is 10 frames, so every cast was cut short), and repair_npc=
a5_larzuk, which Bug 12 set to keep repair in-act for Pindle but which forces
a cross-act trip on Trav (35.8s vs 9.0s).
And a caution on the key auto-detector: it reported the wrong stand_still
bind both before and after a live remap, consistent with a stale .keyo read.
Co-Authored-By: Claude Opus 5 <[email protected]>
* docs: correct stale hammerdin migration notes in HANDOVER (Codex review)
Addresses Codex feedback on PR #40: the handover still described the character
as needing conversion to hammerdin with both skills blank, while the same commit
verified type=hammerdin and concentration='f6' loading live.
The conflict was worse than reported. The stale "Two things only you can do"
section instructed the operator to rebind F7 -> Concentration, but F7 is
Battle Orders (CTA). Following it would have overwritten the CTA bind and cost a
large share of the life pool -- the exact problem this session was fixing.
- Replaces the migration paragraph with the verified end state, and records that
Blessed Hammer lives on LEFT-CLICK, not an F-key (_cast_hammers puts the aura
on the right slot, so a hammer hotkey would replace it).
- Notes that concentration/redemption had to move into [paladin]; under [fohdin]
a hammerdin never reads them and both would be silently unbound.
- Adds the full verified bind map so F7 cannot be reassigned by mistake.
- Rewrites the operator TODOs: fire resist is superseded (Zaka + Mara took
chickens 3-in-8 -> 0-in-11); the live items are lightning/poison resist and
the 75% FCR breakpoint, with the casting_frames mapping (75%=11, 60%=12,
<48%=13) so cast timing tracks any gear change.
Co-Authored-By: Claude Opus 5 <[email protected]>
* feat: baal_xp farm — game browser join, hide-and-wait state, config
- src/ui/game_browser.py: retry-wrapped join flow (Play -> Join Game tab ->
OCR game list -> click row -> loading/InGame confirm), wait_for_in_game
- src/bot.py: 'baal_xp' state + run_baal_xp/end_run transitions;
on_run_baal_xp: fast_save_and_exit -> hero select -> join public game ->
corpse/nopickup/pre_buff -> walk to hide spot -> wait loop (XP every 30s,
HP check, death recovery, InGame-miss re-anchor) -> leave -> recover to
own game; on_end_run short-circuit straight to maintenance
- src/config.py: [baal_xp] section (enabled, game_name_filter, max_wait_s,
xp_threshold, min_hp_pct, hide_x/y, join_timeout_s) with profile overrides
- config/params.ini: [baal_xp] section, order=run_baal_xp, run_trav
- profile paladalla: order=run_baal_xp, run_trav (profile routes override
base config — without this the route never activates)
- stealth: daily budget cap (persisted, restart-proof), buy at A4 Jamella
from A3 too (Ormus measured unreliable), session budget 20h
* fix: import mouse in bot.py (baal_xp hide-spot click)
* feat(stealth): hard per-day runtime cap that survives restarts, and quit D2R
session_budget_h is per PROCESS. Every restart re-rolls it, so a bot that gets
restarted -- by the user, or by the restart-on-crash path -- can run all day and
never trip it. That is exactly the signal a runtime cap is meant to remove. It
also silently failed in practice: params.ini still carried the temporary
levelling value session_budget_h = 20, which rolls 13-27h, so the live config
could not stop the bot within a day at all. Restored to 8.
Adds daily_budget_h, keyed on the CALENDAR DATE and persisted to
log/.daily_runtime.json, so restarts cannot hand out a fresh allowance.
Design points that matter:
- The rolled target is stored WITH the date. Re-rolling per process would make
restarting a way to draw a bigger budget; once a day's target is chosen it is
fixed until the date changes.
- Jitter subtracts ONLY. daily_budget_h is a ceiling, not an average: a
two-sided roll on 8h could hand out 9h, which is not what "cap it at 8 hours"
means. 8 with 0.12 jitter now runs 7.0-8.0h and never more. Verified over 300
rolls: min 7.04h, max exactly 8.00h.
- Time before the first tick in a process is not counted, so a crash
under-counts rather than over-counts -- the safe direction for a cap.
Also closes D2R when the cap trips (daily_budget_close_game, default on). The
check runs in on_end_game AFTER save-and-exit, so the character is at the menu
with nothing in progress and nothing is lost -- this is not the mid-game force
kill that leaves D2R unenterable. A bot parked at character select for 16h is
itself the pattern the cap exists to remove; a real player quits the game.
The stealth manifest reports daily_budget by checking the CALL SITE in bot.py,
not just the config value -- the lesson from chicken_variance and the AFK break
that was 0-for-225 while the manifest said "wired".
Co-Authored-By: Claude Opus 5 <[email protected]>
---------
Co-authored-by: alexpolo1 <[email protected]>
Co-authored-by: Claude Opus 5 <[email protected]>
Two failures observed 2026-09-02 that stopped the leech cycle completing.
min_players default 3 -> 1. A public Baal game spawns at 1-2 players and fills
within seconds, so a 3-player floor skipped every real game: all listed games
read 1-2p and were passed over, the fallback candidate's join never completed,
and the cycle failed until the controller's circuit breaker restarted D2R. The
band was always meant to be a preference rather than a gate — the in-game checks
(portal wait, party size, XP idle) are what decide whether a game is worth
staying in, so the floor was doing nothing except rejecting valid games.
LOBBY button detection hardened. Its template is the one that goes stale: after
the mod font was removed it scored 0.70 against a 0.8 threshold, so every
save-and-exit failed with "not at character select" while the screen was plainly
showing it. Now:
- retry the detect three times before giving up
- fall back to a coordinate click on the known LOBBY position when the template
never resolves, gated on at_character_select() so we only click blind when we
have independent confirmation of the screen
- retry open_lobby once on a transient miss
The "no character selected" path is preserved: LOBBY and PLAY are inert until a
character is highlighted, so that still reports the real cause instead of
burning the timeout on a misleading "lobby did not open".
Also refreshes config/fg_daily_estimates.json (generated data, updated by the
bot's own pricing pass).
Co-Authored-By: Claude Opus 5 <[email protected]>
Written so the bot can be run without me: control commands and their gotchas,
how to tell a normal break from a stuck bot, the break-length multiplier, the
health-check greps, the temporary settings to revert, and the two things only
the user can do (F7 rebind, fire resist).
Co-Authored-By: Claude Opus 5 <[email protected]>
Review catch. The previous commit checked the counter after
find_abs_node_pos, which is still too late.
The anti-stuck block sits BEFORE the scan in the loop body:
790 if _heading_rejects >= MAX: abort <- top-of-loop check
799 if not did_force_move and now - last_move > 3.1:
808 char.move(...) <- the wall-driving guess
826 node_pos_abs = self.find_abs_node_pos(...) <- 3rd rejection recorded
833 if _heading_rejects >= MAX: abort <- too late
With two rejections banked, the moment 3.1s elapses the anti-stuck block
force-moves along last_direction — driving a wall-wedged character further in —
before the third rejection has been recorded. The exact guess this guard exists
to prevent stayed reachable on the threshold iteration.
The scan and the abort decision now both run ahead of the anti-stuck block, so
the counter is current when that decision is made.
The ordering test could not catch this: it searched for "taking a random guess"
only in the source AFTER find_abs_node_pos, while the guess block sits before
that call, so the comparison was against nothing. It now locates the anti-stuck
block explicitly and asserts BOTH the scan and the abort precede it.
Verified by falsification: restoring the scan-after-guess order fails with
"the node scan must run BEFORE the anti-stuck force-move".
That is now four times in this codebase where a check was verified by where it
sat in the source rather than by whether it ran at the deciding moment.
Co-Authored-By: Claude Opus 5 <[email protected]>
Across 79 hell games: 4 random guesses, 0 aborts — including this sequence
inside a SINGLE traverse:
Traverse from a5_town_start to a5_nihlathak_portal
rejecting low-confidence A5_TOWN_1 (66.4%) for node 3
rejecting low-confidence A5_TOWN_1 (65.9%) for node 3
rejecting low-confidence A5_TOWN_1 (64.1%) for node 3
taking a random guess towards (-218, 23)
Wanted to select A5_RED_PORTAL, but could not find it
Three rejections is the threshold, so it should have aborted. The check existed,
was correctly indented inside the while loop, and sat before the anti-stuck
block — the placement I verified with a test when I added it. But the loop does
not reliably come back round to the top of the body after a rejection, so the
check was never evaluated at the moment the counter crossed.
That is why the earlier fix looked right and changed nothing: the ordering test
asserted where the check SAT in the source, not that it ever RAN.
Now checked immediately after find_abs_node_pos, in the same iteration the
counter trips, which removes the dependence on control flow entirely. The
original top-of-loop check is left in place as a second chance.
The character ended up outside the Harrogath battlements again, and the game
was lost to a 82s approach — the exact failure the abort was written to
prevent, still happening because the abort was inert.
Tests now assert the counter TRIPS at the threshold, not merely that the code is
ordered correctly.
Co-Authored-By: Claude Opus 5 <[email protected]>
Follow-up to the GameMenu guard, caught by its own instrumentation: 75 menu
escapes across 15 games, ~5 per game, all clustered around game end.
save_and_exit deliberately opens the in-game ESC menu, but callers only pause
the panel check AFTER it returns:
18:58:48.142 game | end | ok
18:58:48.334 In-game menu open - closing it (1/6) <- the guard
18:58:48.487 Clicking SAVE_AND_EXIT_NO_HIGHLIGHT <- the bot
18:58:48.736 In-game menu open - closing it (2/6) <- the guard again
18:58:49.098 Health Manager is now paused <- too late
Games still completed, so this was noise rather than breakage — but it is a
race, and the guard was pressing esc while the shutdown clicked the menu.
save_and_exit now pauses the panel check for the whole sequence and restores it
in a finally, so it cannot leak the paused state if save/exit raises.
The guard itself is working: 15 games, 0 failures, 0 portal failures, and the
loot-filter/Chronicle/Options incidents have not recurred.
MANA> instrumentation has also settled #23 — see the issue.
Co-Authored-By: Claude Opus 5 <[email protected]>
The bot was found sitting in OPTIONS -> VIDEO with a "settings have changed,
apply or discard?" modal. Discarding revealed the cause: the in-game ESC menu
carries these buttons at SCREEN CENTRE.
OPTIONS / SAVE AND EXIT / RETURN TO GAME / LOOT FILTER / CHRONICLE
A stray esc opens the menu and the bot's next movement click lands on one of
them. The HUD mask deliberately leaves screen centre clickable, so nothing
stops it.
That single mechanism explains three incidents previously treated as separate:
- the LOOT FILTER being toggled (blamed on vigor=f4, which was a real but
different bug)
- CHRONICLE blanking every template match for a whole run, costing a 66s
click_red_portal failure
- the video OPTIONS being opened and a setting changed, which could have
altered resolution and broken every template in the project
The menu has NO close button, so the CenterPanel guard (CLOSE_PANEL_2) cannot
see it. Bug 31 warned about precisely this: "just send esc is WORSE: with
nothing open, esc opens the GAME MENU, which LeftPanel/RightPanel do not
match".
SAVE_AND_EXIT_NO_HIGHLIGHT scores 1.000 on the menu frame and does not match a
normal town frame, so detection is unambiguous. Escaped without counting toward
a chicken, bounded like the waypoint and centred-panel cases.
Also adds MANA> threshold-crossing logging for #23. That issue measured "1 mana
potion per game, never 2" over 70 games but was undecidable, because mana is
only logged when a potion is DRUNK — a second dip that failed to trigger looks
identical to mana never dipping twice. Every crossing is now logged with the
gate state, edge-triggered so it fires once per crossing rather than per poll.
NOT fixed here: whatever sends the stray esc. This is the safety net; the source
is still unknown.
Co-Authored-By: Claude Opus 5 <[email protected]>
69 commits since v0.9.1. Most of what landed was found by measuring which
behaviours FIRED, not by reading code: AFK breaks had gone 0-for-225 against a
configured 5%/game, eight [stealth] settings were declared in params.ini and
never loaded into Config(), and several functions had no callers at all.
Nightmare failure rate ~11% -> 1.0% over 100 games; hell went from dying on
game 4 to 0.0% over 59.
Co-Authored-By: Claude Opus 5 <[email protected]>
walk() carried the identical unguarded click that move() had and was missed
because the tests were written against move() alone. This asserts the invariant
across the movement methods, so a future one is caught without anyone
remembering to extend the tests.
Deliberately NOT covered, because relocating these breaks what they do:
pick_up_item must click the item itself
_remap_skill_hotkey deliberately clicks the UI
cast_in_arc aims a cast direction, not a destination
Only clicks that choose a DESTINATION may be moved off the HUD. A first draft
of this test flagged all of the above and was wrong to; the distinction is
between "go here" and "hit that".
Verified by falsification: reverting walk() to randomize=5 fails with
"walk: mouse.move(x, y, randomize=5, ...)".
Co-Authored-By: Claude Opus 5 <[email protected]>
walk() carried the identical unguarded click that move() had:
x, y = convert_abs_to_monitor(pos_abs)
mouse.move(x, y, randomize=5, ...)
Same fix, same ordering: jitter first via _hud_safe_target, then click with
randomize=0. walk() is reached from bot.py's walk-back-to-town path and from
poison_necro, so it was a live second route to the same loot-filter toggle.
Found while verifying PR #35 had landed — grepping main for randomize= showed a
third mouse.move in i_char.py that the move()-scoped tests did not cover.
session_budget_h 10 -> 8 for a >=5h run. The value is rolled at 0.65-1.35x, so
setting 5 would AVERAGE five hours but could stop after 3.25. 8 gives a
5.2-10.8h window, which guarantees the five while keeping the variation.
Co-Authored-By: Claude Opus 5 <[email protected]>
Review catch on the previous commit — the guard was real but leaky.
get_closest_non_hud_pixel returns the NEAREST unmasked pixel, which by
construction sits exactly on the mask boundary. mouse.move(randomize=3) then
offsets each axis by randrange(-3, 3) = -3..+2, so the cursor can land back
inside the masked region before the right-click. The loot-filter click was made
intermittent, not fixed.
The order is now: jitter -> guard -> move with randomize=0. The human-like
offset is preserved; the guarantee is no longer given away. Same treatment for
the walk branch (randomize=5).
My test missed this because it only checked the guard's OUTPUT, never the point
finally clicked — it passed while the bug was live. The new test samples 400
jittered targets per filter button and asserts every FINAL point is unmasked,
plus a source check that no mouse.move in move() randomizes after the guard.
Verified by falsification: restoring guard-then-randomize makes the suite fail
with "move() still randomizes after the guard, which can re-enter the HUD".
Note on the test itself: an intermediate version compared string indexes over
the whole function source and reported the ordering backwards, because the
DOCSTRING mentions both names. It now compares code with the docstring and
comments stripped — the third time today a source-order assertion was fooled by
prose rather than code.
Co-Authored-By: Claude Opus 5 <[email protected]>
Reported after Larzuk trips: the bot pressed escape, then flipped the loot
filter. Larzuk stands on the left of Harrogath, so moves to and from him aim at
the bottom-left corner — where D2R puts the seven loot-filter category buttons
(screen x 395-560, y 692-712). A right-click there toggles a filter, which
changes what renders and therefore what every later template match can see.
IChar.move() applied NO HUD avoidance in either branch:
# teleport
mouse.move(pos_monitor[0], pos_monitor[1], randomize=3, ...)
mouse.click(button="right")
# walk
x, y = convert_abs_to_monitor(pos_abs)
mouse.move(x, y, randomize=5, ...)
The pather's anti-stuck path already called get_closest_non_hud_pixel; move()
never did, and assets/hud_mask.png covers those buttons correctly — the mask
was simply not consulted.
Latent for a long time and surfaced by Enigma. The walk branch shrinks its
target toward centre via adjust_factor, which mostly kept clicks off the HUD by
accident; the teleport branch clicks the raw target, so any low aim point lands
on the interface.
The escape that precedes it is unrelated and correct — common.close() dismissing
the repair panel.
Verified: filter-row targets are moved from y=700 to y=508, a centre target is
returned unchanged.
Co-Authored-By: Claude Opus 5 <[email protected]>
Two things that cost real time today and are invisible from the screen.
The bot sits at the D2R character-select menu during a NORMAL AFK break —
breaks happen between games, after save-and-exit. The stuck case looks
identical there, and `status` cannot separate them either, because it reports
the game controller rather than what the bot is doing. The tell is the log
filling with "select_char: Could not find online/offline tabs" and "Restarting
bot" every ~20s; a healthy break is simply quiet.
And the configured break length is not the real one. maybe_afk_break calls
wait(minutes*60, minutes*60*1.5) and wait() then applies its own jitter (up to
1.44x), so they compound:
planned 11.9m -> took=1167.7s (19.5m)
planned 20:56 -> took=1531.1s (25.5m)
afk_break_max_m = 12 therefore meant "up to ~26 minutes", and ~25m is what left
D2R unable to re-enter. Documented with the multiplier to apply before deciding
any break duration is safe.
Co-Authored-By: Claude Opus 5 <[email protected]>
afk_break_max_m is not the ceiling it looks like. maybe_afk_break calls
wait(minutes*60, minutes*60*1.5)
and wait() then applies its own jitter (up to 1.44x), so the two compound. A
configured 12 minutes can idle for roughly 26.
Measured today:
planned 11.9m -> took=1167.7s (19.5m)
planned 20:56 -> took=1531.1s (25.5m) [scheduled break]
A ~25m idle is what left D2R at character select on 2026-08-28, unable to
re-enter, with the bot spawning a replacement process every ~20s until four
instances were stacked. 19.5m resumed cleanly, so the tolerated limit sits
somewhere between.
12 -> 7 puts the real worst case at 7 x 1.5 x 1.44 = 15.1m, inside the range
proven to resume, while keeping the 2m floor and the variation intact.
Left the compounding itself alone deliberately: the double randomisation is
good cover, and only its unbounded tail was the problem.
Co-Authored-By: Claude Opus 5 <[email protected]>
restart_or_exit spawned a replacement process and exited, with no attempt limit
and no backoff:
subprocess.Popen([sys.executable, os.path.abspath(sys.argv[0])])
os._exit(0)
On 2026-08-28 a 25-minute scheduled break left D2R on a screen the bot could
not re-enter:
=== BOT START ===
select_char: Could not find online/offline tabs
Restarting bot — game kept running
Because the failure was persistent, this span up a new process roughly every 20
seconds. Instances stacked (4 observed) and then refused taskkill.
The counter has to survive the exec — each restart is a NEW PROCESS, so an
in-memory counter cannot bound the chain. It lives in log/.restart_count,
is checked BEFORE spawning a replacement, and after 5 consecutive restarts the
bot stops with a Discord alert instead of looping, telling the user to return
D2R to the main menu.
A 5s-per-attempt backoff (capped at 60s) stops a fast failure spinning CPU or
stacking processes faster than they exit.
The count is cleared on reaching town, not at game end: the loop failed at
select_char, well before town, so reaching town is what proves recovery — and a
healthy bot never accumulates toward the cap.
Verified by falsification: restoring the unbounded restart makes the suite fail
with "restart loop is unbounded".
Co-Authored-By: Claude Opus 5 <[email protected]>
Bug 31 fixed the keypress that opened CHRONICLE but explicitly left the real
gap open: the panel guard matches only LeftPanel / RightPanel, and a centred
panel matches neither. Nothing closed it, and a centred panel blanks EVERY
later template search.
Recurred 2026-08-28 with the merc key never pressed at all
(resurrect_merc | skip | merc alive), so something else opened it and it simply
stayed. The error screenshot shows Chronicle filling the screen while the bot
hunted A5_RED_PORTAL for 66s:
Traverse from a5_larzuk to a5_nihlathak_portal
Wanted to select A5_RED_PORTAL, but could not find it
random guess towards (96, -115)
Wanted to select A5_RED_PORTAL, but could not find it
Why the existing guard missed it, measured on that frame: the Chronicle's close
button scores 1.000 against CLOSE_PANEL_2 at (952, 56). That x IS inside
right_panel_header (830,0,455,56) — but the ROI is 56px tall and the button's
centre sits exactly at y=56, so the match falls outside. Both header ROIs
scored 0.409 and 0.373.
Adds a center_panel_header ROI (400,0,620,92) and a CenterPanel ScreenObject,
and escapes it WITHOUT counting toward a chicken — a centred panel is
self-inflicted UI state like the waypoint panel, and chickening on it throws
away a healthy game. Bounded by the same _MAX_WP_PANEL_ESCAPES so one that
genuinely will not close still falls through.
Verified on the failure frame: CenterPanel True, LeftPanel False, RightPanel
False.
Co-Authored-By: Claude Opus 5 <[email protected]>
afk_break fired 0 times in 9 games at a 50% test rate — odds of roughly 1 in
500, so not variance.
The four call sites added earlier all sit in on_end_run's town-return branches,
and with one route configured the bot never reaches them. After the run it goes
straight to end_game:
Loot from run_pindle: ...
TL> g8 r8 | game | end | ok
Clicking SAVE_AND_EXIT ... End game. Elapsed time: 69.80s
Starting game #9
No return_to_town or tp_town line appears anywhere in the log. Maintenance runs
at game START, not after the run, so those branches are dead code for this
configuration.
The roll now sits in on_end_game beside the scheduled-break check — the same
path that demonstrably works, since scheduled_break has fired and resumed.
Between games is also the right moment semantically: the game is closed, so
idling there is safe.
Worth recording why the guard missed it. The STEALTH> manifest reported
"afk_break 5% wired - 4 call sites" throughout, which was true and useless: a
call site EXISTING is not the same as a call site being REACHED. Static
reachability is not something the manifest can decide. The digest's "NEVER
FIRED" line is the check that actually catches this class, and it is why that
line exists.
Co-Authored-By: Claude Opus 5 <[email protected]>