Cherry-picks upstream agent-browser #1377 (chore: enforce pnpm minimum
release age), adapted for the fork:
- add .node-version (24); workflows read node-version-file instead of inline
- pin packageManager pnpm@11.1.3; drop hard-coded pnpm/action-setup versions
- pnpm-workspace.yaml: add minimumReleaseAge (48h supply-chain cooldown) +
allowBuilds allowlist, keeping our trimmed packages list (no packages/*, docs)
Deliberately dropped from upstream:
- engines.node >=24 / engines.pnpm >=11 — would impose a Node 24 floor on
end-users of the published agent-browser-stealth CLI (a compiled binary that
doesn't need it). packageManager + .node-version cover dev/CI pinning.
- docs/ and README hunks — those paths are removed/rewritten in this fork.
Two related bugs that conspired to ship stale linux binaries on
0.27.0-fork.5/.7/.8 (caught only by manually grepping the embedded
version string each release):
1. build:all-platforms used `(... & npm run build:linux & wait)`.
The bare `wait` waits for ALL children but exits with the LAST
waited child's status, not each individually. So if linux fell
over and windows succeeded last, the script reported success.
Worse, when both processes shared cli/target/ and fought over
cargo's filesystem locks, one would silently bail out and the
missing binary just stayed at the previous release's bytes.
Now serial: `npm run build:linux && npm run build:windows &&
npm run build:macos`. Costs ~3 extra minutes wall-clock vs.
parallel; trades latency for "every release ships what it says".
2. build:macos had the same `(... & ... & wait)` parallel pattern
for arm64 + x64 cross-compiles. Native cargo builds against the
same target/ dir share even more state than the docker'd Linux
build did, so the failure mode is the same. Now uses explicit
`PID1=$!; PID2=$!; wait $PID1 || exit 1; wait $PID2 || exit 1`
so both must succeed.
Companion to the docker-compose $$ fix in 947d150 (which fixed the
*inside-container* wait+cp eating shell vars). This one fixes the
*outer* npm-script layer.
fork.7 caught the X mask-overlay race correctly but reported it to
the user verbatim — every transient overlay (modal backdrop, focus
ring, click-outside mask, sticky banner) became an error the user
had to wrap in their own retry loop. Most of these clear within a
frame or two on their own.
Now `verify_click_target` retries the elementFromPoint probe a few
times (default 3 × 200ms = 600ms total grace period) before failing.
Real-world overlays that blink in for a render cycle clear during
the first retry; persistent overlays still surface as errors with
the same actionable message — just qualified with "still occluded
after N retries / Mms" so the user knows we tried.
Tunable:
AGENT_BROWSER_OCCLUSION_RETRIES (default 3, 0 disables)
AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS (default 200)
DOM.resolveNode is called once outside the loop — backendNodeId is
stable across renders, only the element under (x, y) changes when
overlays flicker. Each probe is still capped at 500ms so a stuck
Runtime.callFunctionOn can't stall a click for longer than the user
expects.
Real bug behind 0.27.0-fork.5 and fork.7 shipping stale linux binaries.
Docker compose interpolates \${VAR} (and \$VAR) at YAML parse time
against the host shell — including inside `command:` blocks. So:
PID1=\$! ← compose sees \$! → host has no `!` var → ""
wait \$PID1 ... ← compose sees \$PID1 → "" → becomes `wait `
SRC="...\$TARGET..." ← \$TARGET still works (set in `environment:`)
cp "\$SRC" "..." ← \$SRC eaten → empty → cp errors silently
Result: the per-PID error check I added in dbf272c never fired
because both lines were `wait` (no args) — which waits for ALL
children and exits with the LAST one's status, not each individually.
A failing arm64 build couldn't fail the script.
Fix: escape every script-local \$ as \$\$. Docker compose translates
\$\$ → literal \$ when materializing the command for the container,
and the in-container shell then expands \$VAR correctly.
Verified by `docker compose config` showing the resolved command
contains \$\$PID1 / \$\$SRC etc (which becomes \$PID1 / \$SRC in the
container's bash).
Closes the "modal silently closes when clicking 'Add post' on a thread"
bug. Verified root cause via instrumented page-side click logger:
click @e31 (aria-label="Add post" at button (1034, 285))
→ mouse event dispatched to (1045, 296)
→ document.elementFromPoint(1045, 296) returned:
DIV[testid="mask"], bounds (0,0,1746x934)
→ X interpreted as "click outside modal" → close + nav to /home
The cached coordinates were correct. Between snapshot and click, X
laid a transient full-viewport mask over the modal (their own
"click-outside-to-close" overlay). stealth dispatched the click
without checking what was actually at that pixel — the overlay
intercepted it.
Fix: just before returning (x, y) from resolve_element_center for
ref-based interactions, run a Runtime.callFunctionOn against the
ref's resolved element with `function(x, y) { return this.contains(
document.elementFromPoint(x, y)) || that.contains(this) ? null :
{...occluder details...}; }`. If the element at the point isn't us
(or our descendant — clicking the SVG icon inside a button is fine
— or our ancestor), we fail with a specific message:
Ref @e31 is occluded by DIV[testid=mask] at the click point.
A transient overlay (modal backdrop, mask, sticky banner, etc.)
appeared between snapshot and click. Wait for it to clear or
re-snapshot, then retry.
So instead of silently submitting an entire thread or nuking the
user's modal, agent gets a parseable error and can wait + retry.
Tight 500ms timeout per CDP call (matching the verify_ref_identity
defensive guard from fork.6) so a stuck DOM.resolveNode can't
re-introduce the multi-minute hang we just fixed. On any timeout
or error in the guard itself, fall through and let the click
proceed — strictly no worse than the unguarded code path.
Disable with AGENT_BROWSER_VERIFY_CLICK_TARGET=0.
Reported: a single `click @ref` could hang 5+ minutes, with multiple
queued click invocations adding up to 7+ minutes — worst case 30s
timeout × 3 CDP calls × N parallel processes:
- verify_ref_identity (Accessibility.getPartialAXTree) → default 30s
- resolveNode / getBoxModel → default 30s
- wait_for_paint_settled (Runtime.evaluate awaitPromise) → default 30s
The latter two are best-effort defenses added in fork.3-5 to fix SPA
race / DOM-reuse bugs. They should never block a real click for
30s — the unguarded code path was always faster than the guarded
path-that-hangs.
- verify_ref_identity capped at 1s (skips check on timeout)
- wait_for_paint_settled capped at 500ms (skips wait on timeout)
Both skip-on-timeout intentionally: the worst case is the click
behaves like fork.2 (race-prone but fast), which is strictly better
than the user pkilling stuck processes.
Also rewrites the misleading "Chrome 144+ chrome://inspect tip" in
the auto-connect failure message — the toggle exposes target
discovery only, not the /json/version HTTP API the auto-connect
flow expects (verified by user: lsof shows :9222 listening but
curl /json/version returns 404).
Two latent bugs in the release pipeline that conspired to ship a stale
linux-x64 binary in 0.27.0-fork.5 (only caught by manually grepping
the embedded version string):
1. build-linux ran x64 and arm64 in parallel and used a single
`wait $PID1 $PID2` to join them. That command waits for both, but
its exit code is the LAST waited pid only — so if x64 silently
broke and arm64 succeeded, the outer script exited 0 and shipped
whatever was already in /output from the previous release. Now we
wait on each pid individually and exit 1 on either failure.
2. build-single's cp used `agent-browser*` which globs to BOTH the
binary and its `.d` dependency file. When two sources are passed,
cp requires the destination to be a directory. We weren't, so cp
exited non-zero with "Not a directory" and the build script
shrugged it off because the next line was `chmod ... || true`.
Now we resolve a single explicit source path.
Two changes that pair with each other:
1. connect_auto_with_fresh_tab now does a Runtime.evaluate "1"
round-trip after creating the fresh tab. This catches the zombie
CDP socket case (process alive, websocket dead) where every step
up to that point reports success but the next user command would
silently no-op against a dead session. Failing here lets the
caller surface a proper "CDP session unresponsive" error instead
of returning Ok and letting `agent-browser open URL` exit 0 with
a still-blank tab.
2. handle_wait now recognizes @ref selectors (e.g. `wait @e8 --gone`).
It polls resolve_element_object_id, which already runs the
verify_ref_identity check from 007fd1b — so:
- `wait @e8` succeeds while the original element is
still mounted with its snapshot role+name
- `wait @e8 --gone` succeeds when the ref's identity changes
(modal closed, button re-textified, etc.)
This gives users the "assert modal still open" primitive that
prior versions could only approximate with screenshots.
In CDP-attach mode (the default since 0.24.0-fork.1), --headed has no
effect — the user's existing Chrome is already visible, and the
generic "use 'agent-browser close' first to restart" advice doesn't
help (the new daemon attaches right back). Explicitly say --headed is
moot and point to --launch as the actual escape hatch.
Other ignored flags (--profile, --proxy, etc.) keep the existing
"close + reopen" message because for those it IS the right advice.
Closes the "click @e20 hits the sibling element" bug. Real-world
example: snapshot shows @e20=[button "Add post"] next to
@e17=[button "Post all"]. By the time you click @e20, React has
re-rendered — and React often re-uses the same <button> DOM node
across renders, just updating its accessible name. The cached
backendNodeId still resolves to a real, well-positioned node, so
the click lands cleanly. It just lands on what is now the "Post all"
button, silently submitting the entire thread instead of adding a
draft row.
Before every ref-based interaction (click / fill / type / hover /
select / drag — anything routing through resolve_element_center or
resolve_element_object_id), call Accessibility.getPartialAXTree for
the cached backendNodeId and check role + name still match the
snapshot entry. On mismatch, abort with an error that names both
labels:
Ref @e20 no longer matches its snapshot. Was [button "Add post"],
now [button "Post all"].
...Take a fresh snapshot, then re-target.
If the node is gone (CDP fails / no AX node), we silently fall
through to the existing "find by role+name" recovery path, so this
guard never makes a working flow worse.
Adds one CDP roundtrip per ref interaction (~5–20ms). Disable with
AGENT_BROWSER_VERIFY_REF=0 if you control the page lifecycle and
need the latency back.
Pairs with the click paint-settle fix: even with that, a thread builder
that clicks "Add post" can race a misbehaving handler that closes the
parent modal instead of mounting the next textbox. To make that case
observable instead of silently corrupting the next inserttext, you can
now write:
click @add-post
wait .modal --gone --timeout 2000 # asserts modal stays mounted
inserttext "tweet 3"
If the modal vanished, `wait --gone` succeeds — flip the assertion to
`wait .modal` (default visible) to fail-fast on disappearance.
Implementation just sets `state: "detached"` (or "hidden") on the wait
command — daemon-side `wait_for_selector` already supported these
states; only the CLI parser was missing the user-facing flag.
Also accepts `--detached` as alias for `--gone` to match the daemon's
internal vocabulary.
Closes a real-world race that broke X multi-tweet thread composition
(and similar SPA flows): clicking "Add post" returned immediately,
inserttext fired before React had committed the new textarea, the
keystroke landed on the dialog wrapper, and X interpreted the stray
input as a request to dismiss the modal.
After mouseReleased we now wait for two requestAnimationFrame ticks
plus a microtask boundary (~33ms at 60fps, bounded). That's enough
for React/Vue/Svelte to commit any state update scheduled by the
click handler. Errors during the wait are swallowed — a click never
fails because of post-processing.
Opt out for perf-sensitive scripts that don't drive SPA UIs:
AGENT_BROWSER_CLICK_WAIT_STABLE=0
The previous lockfile had ~11k lines of transitive deps for
packages/dashboard which we deleted in 86c4cff. Re-running pnpm install
shrinks it to ~24 lines (just husky for git hooks).
Before: after `npm i -g` upgrade, the next agent-browser command would
detect daemon version mismatch, kill the old daemon, spawn a fresh one,
and connect to a brand-new about:blank tab. The user's previous
navigation state was silently lost — `get url` returned about:blank
even though the user's Chrome was still on the same page.
Now: before killing the old daemon, the CLI synchronously asks it for
its current URL via the existing socket. If non-empty and not
about:blank, it's persisted to a `.restore-url` sidecar in the socket
dir. After the new daemon spawns and auto-connects, it reads the
sidecar (read-and-delete), navigates the fresh tab to the saved URL,
and prints `⚠ Restored previous URL: <url>`.
Manual `agent-browser close` does NOT write the sidecar, so a clean
shutdown won't trigger surprise navigation. The sidecar is consumed on
read regardless of whether navigation succeeded, so a stale entry
can't haunt later auto-launches.
Before, `agent-browser find role button --name Submit` errored at the
daemon side with the cryptic `Unknown subaction: --name`. Now it errors
at parse time with the offending flag echoed back, the list of valid
actions (click, fill, check, hover, text), and a "Did you mean" hint
showing where to put the action verb.
Backwards compat: `find role button` (no flags, no action) still
defaults to click — only `--xxx` in action position errors.
npm 10+ strips bin paths starting with ./ as invalid, leaving the
package with no executable entries (so `npm i -g` doesn't put any
binary on PATH). Match the upstream form `bin/agent-browser.js`.
- Add fork binary names (agent-browser-stealth, abs) to allowed-tools
in all 6 SKILL.md files so installs into Claude Code / Cursor don't
prompt for permission on every command
- Document `npx skills add leeguooooo/agent-browser-stealth` in README
- Bump README upstream-base mention from v0.24.0 to v0.27.0
These directories are TypeScript-side tooling that the fork dropped at
v0.24.0 to keep the repo focused on the stealth CLI binary. Upstream
either kept evolving them (docs, packages/dashboard) or added new ones
(evals/) — they came back during the v0.27.0 rebase, so prune again.
Also include skill-data/ in package.json `files` so the specialized
skills (electron, slack, dogfood, etc.) that upstream relocated from
skills/ to skill-data/ still ship in the npm tarball.
Key insight: ANY JS-level modification to navigator.webdriver is detectable
by creepjs's lieProps system. The only undetectable approach is
Emulation.setAutomationOverride at the CDP protocol level, which tells
Chrome to natively return false for navigator.webdriver.
In CdpAttach mode, we now inject ZERO JavaScript patches — the browser's
real fingerprint is already perfect. Only the CDP protocol command is needed.
CreepJS results now match manual Chrome exactly:
- 0% headless (was 33%)
- 0% stealth (unchanged)
- 25% like headless (Chrome baseline, same as manual)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
CreepJS detects three things for webDriverIsOn:
1. Property deletion (navigator.webdriver === undefined)
2. Value check (!!navigator.webdriver)
3. Lie detection (descriptor tampering via lieProps)
Changed from delete/defineProperty-value approach to replacing the CDP
getter with a getter returning false, matching the native descriptor shape.
Note: 33% headless in CreepJS is a CDP-inherent signal (lieProps detects
the getter replacement). This cannot be eliminated at the JS layer since
CDP sets the webdriver getter before init scripts run. Real-world impact
is minimal — Cloudflare Turnstile passes successfully.
Also confirmed: Chrome's remote_debugging preference in Local State
persists across restarts, so users only need to enable CDP once via
chrome://inspect/#remote-debugging.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- CdpAttach mode: only removes navigator.webdriver (user's real Chrome
already has genuine fingerprint, heavy patches create detectable lies)
- FullLaunch mode: applies all 32 patches (new Chrome needs full coverage)
- Improved webdriver removal: uses Object.defineProperty to override CDP
getter on Navigator.prototype, not just delete
- CreepJS results: 0% stealth (was 20%), hasIframeProxy: gone
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Auto-connect is now ON by default (was opt-in via --auto-connect)
- Added --launch/--new flags to explicitly start a fresh browser
- CI environments (CI env var) automatically use --launch mode
- Friendly error message with platform-specific Chrome relaunch guide
- Mentions Chrome 144+ runtime CDP toggle (chrome://inspect)
- --cdp and --provider flags implicitly disable auto-connect
- AGENT_BROWSER_NO_AUTO_CONNECT=1 to disable, AGENT_BROWSER_FORCE_LAUNCH=1 to force
Track 3 of native-stealth migration.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Created cli/src/native/stealth.rs with stealth JS injection via CDP
- Extracted 32 patch IIFEs from TS stealth.ts into stealth_scripts.js
- Injected via Page.addScriptToEvaluateOnNewDocument on every launch/connect
- Added stealth Chrome args (disable AutomationControlled, use ANGLE GL)
- Auto-detects and cleans HeadlessChrome from User-Agent string
- Overrides navigator.userAgentData high-entropy hints
- Stealth enabled by default, disable with AGENT_BROWSER_STEALTH=0
Track 2 of native-stealth migration.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>