Three CLI gaps surfaced driving a Mercari signup→checkout flow:
- #24-B (correctness): a bare label like 'click 購入手続きへ' was fed straight to
document.querySelector as CSS and failed as an invalid selector, even though
snapshot listed the button by that exact name. build_find_element_js now tries
CSS first, then falls back to matching an interactive element by visible text
(exact then contains) — nested and non-ASCII labels resolve. 'text=<label>'
forces the text path. CSS still wins when it matches.
- #24-D: 'get text' with no selector now returns the whole page (body).
- #24-C: 'tab <ref> --activate' (alias --front) switches to the tab AND raises it
to the foreground — to surface a specific tab for the human.
Tests cover the text fallback / text= / xpath builder, body default, activate
flag. The core stale-sessionId-after-cross-process-nav bug is the #20/#23 class,
already fixed in ext 0.4.8 — needs that extension deployed.
Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled
Standalone product rename across the whole repo (issue: project identity):
- Binary/package/repo/skill/docs: agent-browser[-stealth] → chrome-use
(single binary name `chrome-use`; old aliases agent-browser/abs dropped).
- Version: 0.27.0-fork.51 → 1.0.0 (drop the upstream-fork counter).
- Native-messaging host: com.agent_browser.connect → com.leeguoo.chrome_use
(CLI + ab-connect extension in lockstep — this is a breaking handshake change,
extension bumped 0.4.2 → 0.5.0, needs a Web Store republish).
- Config dir: ~/.agent-browser → ~/.chrome-use.
- README/zh: reframed from "stealth fork of agent-browser" to a standalone
product with a small `originally based on vercel-labs/agent-browser` credit.
- Kept AGENT_BROWSER_* env vars working (63 vars across the codebase; renaming
them would break every existing script/skill for no user-facing gain).
Build green, 802 unit tests pass, fmt + clippy clean. Upstream attribution to
vercel-labs/agent-browser preserved.
- snapshot -c (compact) now always keeps lines with an interactive ARIA role
(button/link/textbox/combobox/option/…), not only `ref=`/`": "` lines — so a
clickable control can't vanish from compact output and leave the agent clicking
an empty ref (issue #2 P1). Additive: only ever keeps more. compact tests green.
- stale-ref error now leads with "take a fresh snapshot" and points to the `eval`
fallback for ref-churning SPAs, and demotes AGENT_BROWSER_VERIFY_REF=0 to a
flagged last resort instead of presenting it as the fix (issue #3 P1).
Completes the humanize suite:
- Clicks land on a jittered point inside the element's box (Fast/Human) instead
of its exact centre. `resolve_element_center` now also returns the element
width/height (box_model_dims); the CSS-selector path reports zero size → land
on centre (no jitter, no regression). Jitter is clamped to the inner box so the
click never misses.
- Wheel scrolls split into eased, jittered segments (humanize::scroll_segments,
unit-tested) instead of one instant jump.
- Drag follows the curved trajectory at Fast/Human (linear 10-step at Off).
Off is unchanged throughout. 9/9 unit tests; verified headless — jittered click
still lands (→ iana.org), segmented scroll moves the page.
Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled
- invalid CSS selector now errors "Invalid selector '<sel>': <reason>" instead of
the misleading "Element not found" — the coordinate path (resolve_by_selector)
now also inspects exception_details, matching resolve_element_object_id.
- `wait --url ""` is rejected at parse time ("needs a non-empty pattern") rather
than silently matching any URL. Unit test added.
Not changed: verb-less `find role X` defaulting to a click. That default is a
deliberate, tested decision (test_find_role_default_subaction_click_when_no_action);
changing it to locate-and-report is a design choice left to the maintainer.
Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled
- wait --url: the arg parser never read `--timeout`, so a non-matching pattern
waited the large default and wedged the daemon. Parse it. Also: matching was a
literal substring (`includes`) so globs never matched — convert `**`/`*`/`?`
globs to an anchored regex. And `poll_until_true` now bounds each probe with a
timeout and tolerates transient navigation errors, so a hung `Runtime.evaluate`
can never block past the deadline (un-wedges the daemon).
- find role <role> [--name]: the query was `[role="X"], X`, which matches a
literal <X> tag / explicit attribute but NOT implicit-role elements — so
`find role link` (<a href>) and `find role heading` (<h1>) never matched. Add a
proper ARIA-role → implicit-element map and broaden accessible-name matching
(aria-label/title/alt/value/text).
- click on a syntactically-invalid selector returned `✓ Done`: querySelector
throws, and Runtime.evaluate returned the thrown DOMException as an objectId
that was clicked as if it were the element. Check exception_details → error.
- output: a title-less page now prints `✓ <url>` instead of an empty title line.
- docs(skill): tab refs are `t2`, not `2` (SKILL.md, electron).
Verified live (isolated launch): wait --url glob matches instantly; non-matching
honors --timeout (2s) and leaves the daemon responsive; find role link/heading
match; invalid selector errors. Unit tests added for the glob + role map + parse.
The fork's CI had never been green. Pre-existing failures:
- version-sync: check-version-sync.js read packages/dashboard/package.json,
which doesn't exist in this fork (workspace is just "."). Drop the dashboard
comparison; check package.json vs cli/Cargo.toml only.
- Dashboard job: `pnpm install --filter dashboard` for a non-existent package.
Remove the job.
- Format check: repo was never `cargo fmt`-clean. Ran cargo fmt (mechanical).
- Clippy -D warnings (newly enforced on Rust 1.94 stable): manual_contains in
commands.rs (.iter().any()->.contains()), question_mark in element.rs
(if-let-Err -> ?), result_large_err on the tungstenite handshake callback in
connect.rs (allow — the Result type is fixed by the accept_hdr_async contract).
- rust-cross: lightpanda::waits_for_ready_without_logs spawns a real process +
binds a socket with timing assumptions; flaky in CI. Marked #[ignore].
Also: skill docs note fork.30's relay-preferred auto-connect (plain
`agent-browser open` is dialog-free once the ab-connect extension is loaded) and
the extension's new "agent-browser-stealth" display name.
Borrow Scrapling's adaptive element finding, adapted to this project's
in-session AX-ref model. When a saved @ref's node is gone (or its identity
no longer matches) and the role/name/nth re-query also fails, score the
current page's candidate elements against an AX fingerprint captured at
snapshot time and relocate to the best match.
- New `adaptive` module: pure, browser-free scoring (role, accessible name
via Levenshtein, AX properties, ancestor-role LCS, parent/sibling) plus
pick_best with a high absolute threshold (0.70) AND a clear margin (0.15)
over the runner-up — so ambiguous twins are refused rather than mis-clicked,
matching the existing "fail loudly over wrong click" posture.
- Fingerprint captured during the existing AX-tree snapshot walk — no extra
CDP round-trips. TreeNode is AX-only (no DOM tag/attrs), so we use AX role
as the type and a few discriminating AX properties (value/url/level/checked);
DOM id/class would have cost an N×describeNode storm per snapshot.
- Wired into both resolve_element_center and resolve_element_object_id: on a
verify-identity mismatch or a stale-node fallback miss, relocation is tried
before erroring. A confident match overrides the identity guard; otherwise
the original error is surfaced. Opt out with AGENT_BROWSER_ADAPTIVE_REF=0.
README documents the new tuning knobs. Adds 9 unit tests; full suite 760 passed.
fork.7 caught the X mask-overlay race correctly but reported it to
the user verbatim — every transient overlay (modal backdrop, focus
ring, click-outside mask, sticky banner) became an error the user
had to wrap in their own retry loop. Most of these clear within a
frame or two on their own.
Now `verify_click_target` retries the elementFromPoint probe a few
times (default 3 × 200ms = 600ms total grace period) before failing.
Real-world overlays that blink in for a render cycle clear during
the first retry; persistent overlays still surface as errors with
the same actionable message — just qualified with "still occluded
after N retries / Mms" so the user knows we tried.
Tunable:
AGENT_BROWSER_OCCLUSION_RETRIES (default 3, 0 disables)
AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS (default 200)
DOM.resolveNode is called once outside the loop — backendNodeId is
stable across renders, only the element under (x, y) changes when
overlays flicker. Each probe is still capped at 500ms so a stuck
Runtime.callFunctionOn can't stall a click for longer than the user
expects.
Closes the "modal silently closes when clicking 'Add post' on a thread"
bug. Verified root cause via instrumented page-side click logger:
click @e31 (aria-label="Add post" at button (1034, 285))
→ mouse event dispatched to (1045, 296)
→ document.elementFromPoint(1045, 296) returned:
DIV[testid="mask"], bounds (0,0,1746x934)
→ X interpreted as "click outside modal" → close + nav to /home
The cached coordinates were correct. Between snapshot and click, X
laid a transient full-viewport mask over the modal (their own
"click-outside-to-close" overlay). stealth dispatched the click
without checking what was actually at that pixel — the overlay
intercepted it.
Fix: just before returning (x, y) from resolve_element_center for
ref-based interactions, run a Runtime.callFunctionOn against the
ref's resolved element with `function(x, y) { return this.contains(
document.elementFromPoint(x, y)) || that.contains(this) ? null :
{...occluder details...}; }`. If the element at the point isn't us
(or our descendant — clicking the SVG icon inside a button is fine
— or our ancestor), we fail with a specific message:
Ref @e31 is occluded by DIV[testid=mask] at the click point.
A transient overlay (modal backdrop, mask, sticky banner, etc.)
appeared between snapshot and click. Wait for it to clear or
re-snapshot, then retry.
So instead of silently submitting an entire thread or nuking the
user's modal, agent gets a parseable error and can wait + retry.
Tight 500ms timeout per CDP call (matching the verify_ref_identity
defensive guard from fork.6) so a stuck DOM.resolveNode can't
re-introduce the multi-minute hang we just fixed. On any timeout
or error in the guard itself, fall through and let the click
proceed — strictly no worse than the unguarded code path.
Disable with AGENT_BROWSER_VERIFY_CLICK_TARGET=0.
Reported: a single `click @ref` could hang 5+ minutes, with multiple
queued click invocations adding up to 7+ minutes — worst case 30s
timeout × 3 CDP calls × N parallel processes:
- verify_ref_identity (Accessibility.getPartialAXTree) → default 30s
- resolveNode / getBoxModel → default 30s
- wait_for_paint_settled (Runtime.evaluate awaitPromise) → default 30s
The latter two are best-effort defenses added in fork.3-5 to fix SPA
race / DOM-reuse bugs. They should never block a real click for
30s — the unguarded code path was always faster than the guarded
path-that-hangs.
- verify_ref_identity capped at 1s (skips check on timeout)
- wait_for_paint_settled capped at 500ms (skips wait on timeout)
Both skip-on-timeout intentionally: the worst case is the click
behaves like fork.2 (race-prone but fast), which is strictly better
than the user pkilling stuck processes.
Also rewrites the misleading "Chrome 144+ chrome://inspect tip" in
the auto-connect failure message — the toggle exposes target
discovery only, not the /json/version HTTP API the auto-connect
flow expects (verified by user: lsof shows :9222 listening but
curl /json/version returns 404).
Closes the "click @e20 hits the sibling element" bug. Real-world
example: snapshot shows @e20=[button "Add post"] next to
@e17=[button "Post all"]. By the time you click @e20, React has
re-rendered — and React often re-uses the same <button> DOM node
across renders, just updating its accessible name. The cached
backendNodeId still resolves to a real, well-positioned node, so
the click lands cleanly. It just lands on what is now the "Post all"
button, silently submitting the entire thread instead of adding a
draft row.
Before every ref-based interaction (click / fill / type / hover /
select / drag — anything routing through resolve_element_center or
resolve_element_object_id), call Accessibility.getPartialAXTree for
the cached backendNodeId and check role + name still match the
snapshot entry. On mismatch, abort with an error that names both
labels:
Ref @e20 no longer matches its snapshot. Was [button "Add post"],
now [button "Post all"].
...Take a fresh snapshot, then re-target.
If the node is gone (CDP fails / no AX node), we silently fall
through to the existing "find by role+name" recovery path, so this
guard never makes a working flow worse.
Adds one CDP roundtrip per ref interaction (~5–20ms). Disable with
AGENT_BROWSER_VERIFY_REF=0 if you control the page lifecycle and
need the latency back.
* fix: rewrite getByRole to use CDP accessibility tree instead of CSS selectors
The old `handle_getbyrole` generated `querySelectorAll('[role="link"], link')`
which matched `<link>` stylesheet elements instead of `<a>` anchor tags.
This happened because ARIA role names were used directly as CSS tag selectors,
and several roles differ from their HTML element names (e.g. link → a,
heading → h1-h6, textbox → input/textarea).
The fix replaces the JS-based DOM query with the CDP `Accessibility.getFullAXTree`
API, where the browser engine correctly computes implicit ARIA roles per the
WAI-ARIA / HTML-AAM spec. This is the same approach already used by `snapshot.rs`
and `element.rs` in this codebase.
Changes:
- Rewrite `handle_getbyrole` to query the browser's accessibility tree via CDP
- Add `find_ax_node_by_role` helper for AX tree traversal with role/name/exact matching
- Use `DOM.resolveNode` + `Runtime.callFunctionOn` to bridge AX node → DOM marker
- Add iframe support via `resolve_ax_session` (missing in old implementation)
- Fix cleanup to use correct CDP session (old code used default session, breaking iframe cleanup)
- Export `extract_ax_string` as `pub(super)` for reuse
- Add 4 regression tests for `find_ax_node_by_role`
Fixes#1123
* style: apply cargo fmt
* chore: remove redundant comments
* refactor: replace marker attribute with temporary ref for element resolution
Eliminates 3 CDP round-trips (DOM.resolveNode, Runtime.callFunctionOn,
Runtime.evaluate cleanup) by registering a temporary ref in the ref_map.
execute_subaction resolves the element via backendNodeId directly.
No more DOM pollution with marker attributes.
* fix: ref counter collision, ref_map leak, and stale fallback name
- Increment next_ref_num after inserting temp ref to prevent id collision
- Remove temp ref after execute_subaction to prevent unbounded ref_map growth
- Return actual AX name from find_ax_node_by_role for accurate fallback resolution
- Add RefMap::remove method
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
* fix: support xpath= selector prefix in element resolution
Resolves#907. When a selector starts with "xpath=", use
document.evaluate() instead of document.querySelector() so that
XPath expressions like "xpath=//button" work correctly.
* test: replace overlapping test with edge case tests
Replace test_build_selector_js_xpath_strips_prefix (which overlapped
with the xpath test) with two edge case tests: empty xpath and
selector starting with "xpath" without "=" delimiter.
* fix: support xpath= selector in resolve_element_object_id
Apply the same xpath= handling to resolve_element_object_id, which is
used by type, fill, focus, hover, check, select, screenshot, drag,
and all other selector-based commands beyond basic click.
* refactor: extract build_find_element_js to deduplicate xpath/css logic
The xpath= vs querySelector branching was duplicated in both
build_selector_js and resolve_element_object_id. Extract the shared
logic into build_find_element_js and reuse it in both places.
* fix: support xpath= selector in get_element_count
Use ORDERED_NODE_SNAPSHOT_TYPE with snapshotLength for XPath counting,
matching the querySelectorAll().length behavior for CSS selectors.
* refactor: rename find to find_expr for clarity
* refactor: extract build_count_elements_js and add regression tests
Extract element counting JS generation into build_count_elements_js
helper (matching the pattern of build_find_element_js) and add tests
for both CSS and XPath counting paths.
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
* Add iframe support for CLI interactions and snapshots
This PR adds comprehensive iframe support to the agent browser CLI, allowing users to interact with elements inside iframes seamlessly.
## Problem
Users couldn't interact with elements inside iframes via the command line. The existing `frame` command was non-functional as it set `active_frame_id` but no other code read this value.
## Changes Made
### Enhanced Frame Context Tracking
- Added `frame_id` field to `RefEntry` to track which frame each element reference belongs to
- Updated `RefMap::add` and related methods to accept and store frame context
- Modified element resolution functions to use frame context from ref entries
### Improved Frame Command
- Fixed the existing `frame` command to actually work by threading `active_frame_id` through snapshot operations
- Added support for iframe element references (e.g., `frame @e2`) in addition to CSS selectors
- Enhanced frame detection to work with both named frames and iframe elements
### Updated Snapshot Behavior
- Modified `take_snapshot` to accept optional frame context parameter
- Updated all snapshot call sites to pass appropriate frame context
- Maintained backward compatibility while enabling frame-scoped operations
### Element Resolution Updates
- Updated `resolve_element_center` and `resolve_element_object_id` to use frame context from ref entries
- Modified `find_node_id_by_role_name` to support frame-specific element lookup
- Ensured all interaction functions work correctly within iframe contexts
## Implementation Details
- Frame context is now properly propagated through the entire element interaction pipeline
- The `frame` command can accept both CSS selectors and element references
- All existing functionality remains intact while adding iframe capabilities
- Added `Iframe` to interactive roles for better element discovery
Fixes#863
* docs: add iframe support documentation
Document the new iframe capabilities across all documentation surfaces:
- Auto-inlining of iframe content in snapshots
- Direct interaction with iframe element refs
- frame command support for element refs (@e3)
- Scoped snapshots via frame switching
* fix: pass active frame context to diff snapshots and fix nameless iframe lookup
- handle_diff_snapshot now respects active_frame_id instead of always
passing None, so diff snapshots work correctly inside iframes
- Nameless/id-less iframes now fall back to src URL (or null) instead of
the literal string 'frame' which never matched any frame in the tree
* fix: resolve iframe frame ID via DOM.describeNode and reduce code duplication
- handle_frame: Use DOM.describeNode + contentDocument.frameId to resolve
iframe frame IDs directly, fixing failures for nameless iframes that
lack name/id/src attributes
- element.rs: Deduplicate add() by delegating to add_with_frame()
- snapshot.rs: Guard against out-of-bounds insert_str when iframe marker
is on the last line without a trailing newline
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
The v0.20.0 migration from Playwright to the native Rust daemon introduced
two regressions in checkbox/radio handling:
1. `is_element_checked` only read `this.checked`, which is undefined on
non-input elements. Material Design and ARIA controls use wrapper divs
with `role="checkbox"` and `aria-checked`, or hide the native input
off-screen inside a label. The function now mirrors Playwright's
`getChecked()` with follow-label retargeting: native `.checked`,
`aria-checked` for ARIA roles, `label.control` traversal, and nested
input lookup.
2. `check`/`uncheck` accepted the coordinate-based CDP click result
without verifying the state actually changed. When the AX tree's
`backendDOMNodeId` points to a hidden off-screen input (common in
Material Design), `Input.dispatchMouseEvent` hits nothing. The actions
now re-check state after clicking and fall back to a JS `.click()` on
the resolved input — matching Playwright's `_setChecked` verify step.
Adds e2e regression test covering Material Design (hidden input + ripple
overlay), ARIA-only, and native checkbox patterns.
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: gracefully fall back to role/name lookup when backend_node_id is stale
When the DOM changes between snapshot and click (common with SPAs and
dynamic UIs), the stored backend_node_id becomes invalid. Previously,
DOM.getBoxModel and DOM.resolveNode failures propagated as hard errors,
bypassing the role/name fallback path entirely. Now these failures are
caught and the code falls through to a JS-based element lookup.
Also adds resolve_object_id_by_role_name so that resolve_element_object_id
has a fallback for ref-based lookups (previously it had none), and
improves the role matching JS to correctly map implicit ARIA roles
(e.g. <input type="submit"> → "button", <a href> → "link").
Closes#805
* test: add e2e regression test for stale ref click fallback (#805)
Verifies that clicking a ref whose backend_node_id has become stale
(because the DOM was replaced by JavaScript) falls back to role/name
lookup instead of failing with "Could not compute box model".
* fix: use accessibility tree for stale ref fallback instead of JS heuristic
Replace the hand-rolled JS role/name matching (getImplicitRole,
getAccessibleName) with a re-query of Accessibility.getFullAXTree —
the same data source that built the ref map during snapshot. This
guarantees role/name matching is identical to what was stored,
preventing silent wrong-element clicks from name computation
divergence (e.g. aria-labelledby, <label for>, alt text).
Matches v0.19.0 (Playwright) behavior where getByRole always
re-queried the live accessibility tree.
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>