Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled
Standalone product rename across the whole repo (issue: project identity):
- Binary/package/repo/skill/docs: agent-browser[-stealth] → chrome-use
(single binary name `chrome-use`; old aliases agent-browser/abs dropped).
- Version: 0.27.0-fork.51 → 1.0.0 (drop the upstream-fork counter).
- Native-messaging host: com.agent_browser.connect → com.leeguoo.chrome_use
(CLI + ab-connect extension in lockstep — this is a breaking handshake change,
extension bumped 0.4.2 → 0.5.0, needs a Web Store republish).
- Config dir: ~/.agent-browser → ~/.chrome-use.
- README/zh: reframed from "stealth fork of agent-browser" to a standalone
product with a small `originally based on vercel-labs/agent-browser` credit.
- Kept AGENT_BROWSER_* env vars working (63 vars across the codebase; renaming
them would break every existing script/skill for no user-facing gain).
Build green, 802 unit tests pass, fmt + clippy clean. Upstream attribution to
vercel-labs/agent-browser preserved.
e2e_save_state_cross_domain navigated to httpbin.org as "domain A", which is an
unreliable external service — when it was slow/unreachable in CI the page didn't
load on that origin, so its localStorage origin was missing from the saved state
and the test failed intermittently. Cookies/localStorage are set client-side via
CDP, so the page just needs to load reliably: use example.org (IANA-reserved,
like example.com) instead. Match full hostnames so the two example.* origins
don't alias. Verified locally: passes deterministically.
* feat(react): first-class React introspection, Web Vitals, and nextjs skill
Add React-general and web-universal features as first-class agent-browser verbs
(react tree/inspect/renders/suspense, vitals, pushstate). Genuinely Next.js-specific
workflows (PPR cookie protocol, /_next/mcp bridge, dev-server endpoints) ship as
a new `nextjs` skill that composes the primitives. No new runtime dependencies -
the React DevTools installHook.js is vendored (MIT) and include_str!'d into the
binary.
New commands:
react tree Full React component tree (depth id parent name)
react inspect <fiberId> Props, hooks, state, source for one fiber
react renders start|stop Fiber profiler with Insts/Mounts/Re-renders/Self/DOM
+ prev->next change details
react suspense Suspense boundaries + classifier (client-hook,
request-api, server-fetch, cache, stream, framework)
+ root-cause grouping + recommendations
vitals [url] LCP/CLS/TTFB/FCP/INP + React hydration phases
pushstate <url> Generic SPA client-side navigation
removeinitscript <id> Remove a script registered via addinitscript
New launch flags:
--init-script <path> Register init scripts before first navigation
(repeatable; env AGENT_BROWSER_INIT_SCRIPTS)
--enable <feature> Built-in init scripts; currently react-devtools
(repeatable; env AGENT_BROWSER_ENABLE)
Other primitives:
network route ... --resource-type <csv> Filter by CDP resource type
cookies set --curl <file> Auto-detects JSON/cURL/Cookie-header
* fixes
* fixes
* fixes
* fix(tabs): preserve refs across --tab peek and cover outer-tab-closed path
Follow-up to #1249 so `--tab <id>` is actually useful for agents:
- Save and restore the outer tab's `ref_map`, `iframe_sessions`, and
`active_frame_id` across a scoped command instead of clearing them.
`snapshot` → `--tab N <cmd>` → `click @e1` now keeps the outer tab's
refs intact. Scoped commands still see a clean slate so outer refs
can't resolve against the scoped tab's DOM.
- Close the coverage gap the Vercel review bot flagged on #1249: the
previous `e2e_tab_scoped_command_handles_outer_tab_closed` test used
`tab_close`, which is in the scoped-dispatch exclusion list, so it
never exercised the restore-skip branch it claimed to test. Renamed
to `e2e_tab_close_with_tab_id_closes_active_tab` with an honest
docstring, and added `e2e_tab_scoped_command_outer_tab_closed_mid_dispatch`
that actually hits the branch via `window.opener.close()` on a
script-opened intermediate tab.
- Add `e2e_tab_scoped_command_isolates_refs_from_outer_tab` pinning
that outer refs don't bleed into the scoped tab's DOM resolution.
- Rewrite `e2e_tab_scoped_command_clears_state_on_switch` as
`e2e_tab_scoped_command_preserves_outer_tab_state`, verifying the
restored @e1 still clicks end-to-end.
- Update the 52 `--help` entries for `--tab <id>` to describe peek /
restore semantics instead of a vague "Target specific tab ID".
- Update README, docs site, config schema, and the agent-facing
skills reference with working examples (refs survive the peek) and
a "when to use \`--tab <id>\` vs \`tab <id>\`" guide so agents pick
the right flag for their workflow.
* fix(tabs): use t<N> prefix for tab ids, add --label for named tabs
Follow-on to the tab work in #1249 and the prior commit, redesigning the
tab handle surface before release since nothing ships these features yet.
## Why
Incrementing integer tab ids (`1`, `2`, `3`) look indistinguishable from
positional indices in command output, LLM-generated scripts, and docs. In
the common single-agent case where position and id coincide, readers have
no visual cue for which mental model they're using. Positional indices
silently shift when unrelated tabs open/close, so misreading a handle as
an index is a correctness hazard.
## Changes
**Tab ids are now `t1`, `t2`, `t3` (strings).** Bare integer `tabId`
values are rejected with a teaching message rather than silently accepted.
The `t` prefix matches the `@e1` element-ref convention and makes ids
unmistakably non-positional at a glance.
**Labels.** Tabs can be created with a user-assigned label (e.g. `docs`,
`app`) via `tab new --label <name> [url]`. Labels are interchangeable
with `t<N>` ids everywhere a tab ref is accepted. They're never
auto-generated, never rewritten on navigation, and must be unique within
a session.
**Dashboard fix.** `packages/dashboard/src/types.ts` declared
`TabInfo.index: number` but the daemon has been sending `tabId` (not
`index`) since #892, making `tab.index` `undefined` and breaking the
dashboard's close/switch buttons silently. Updated the TS types and
usages to consume `tabId` (string) and optional `label`, restoring the
dashboard's tab interactions.
## Surface
- `cli/src/native/browser.rs`: `TabRef::parse` / `format_tab_id` /
`is_valid_label` / `PageInfo.label` / `BrowserManager::resolve_tab_ref`
/ `BrowserManager::has_label`. `tab_new` gains an optional label
argument with duplicate rejection. All JSON responses use the string
form and include the label.
- `cli/src/native/actions.rs`: scoped-command pre-dispatch and
`handle_tab_{switch,close,new}` parse string refs and resolve to
stable ids.
- `cli/src/{flags,commands,main,output}.rs`: `--tab` / config `tab`
are `String`; `tab` subcommand accepts `t<N>` or a label and supports
`tab new --label <name> [url]`. All 52 `--help` entries updated.
- `agent-browser.schema.json`: `tab` property type is now `string` with
a pattern matching `t<N>` or label form.
- `packages/dashboard`: `TabInfo.tabId: string` / `label?: string | null`;
`closeTabAtom`/`switchTabAtom` take `tabRef: string`; component props
updated.
- Docs: README, docs site (`commands/` and `configuration/`), and the
agent-facing skills reference rewritten with the new examples.
## Tests
- Added `TabRef::parse` / `format_tab_id` / `is_valid_label` unit tests
pinning the bare-integer rejection, the teaching error, label rules,
and round-tripping.
- Added `test_tab_switch_by_id` / `_by_label` / `test_tab_new_with_label`
/ `_with_label_and_url` / `_with_url_then_label` in `commands.rs`;
rewrote `test_tab_unknown_subcommand_errors` since labels make
`tab select` a legitimate ref.
- Added `e2e_tab_new_with_label_can_be_switched_and_peeked`,
`e2e_tab_new_with_duplicate_label_errors`,
`e2e_tab_scoped_command_rejects_bare_integer`.
- Migrated every existing tab e2e test (and one unit test) from
integer `tabId` to the string form.
`cargo fmt`, `cargo clippy -- -D warnings`, all 30 non-ignored tab unit
tests, all 13 tab e2e tests, and `tsc --noEmit` on the dashboard all
pass.
* refactor(tabs): drop --tab scoped peek flag; keep t<N> ids and labels
After fleshing out `--tab <id|label>` in the previous commits (scoped
pre/post-dispatch save/restore, ref preservation, outer-tab-closed edge
case, full e2e coverage), the machinery-to-value ratio makes the feature
hard to justify. Nixing it now while nothing has shipped.
## Why
- Every new daemon feature touching per-tab state has to reason about
scoped-dispatch interleaving. `ScopedRestore`, pre/post-dispatch hooks,
and the exclusion list add ongoing maintenance tax.
- Three separate PRs (#892, #1249, and this one pre-nix) were needed to
reach "works correctly." That's a smell.
- `tab <id|label>` switch + labels already cover the legible multi-tab
workflow case.
- `--tab` vs `tab <id>` have opposite lifecycle semantics but look
identical, teaching every agent two things where one would do.
- "Non-disruptive peek" isn't actually race-free: the daemon does swap
active tab during execution, so a concurrent client between pre- and
post-dispatch sees the scoped tab as active.
- Ref-based interaction with scoped tabs never worked ergonomically —
refs are per-tab, so `--tab N click @e1` requires `@e1` to already be
on tab N, which means a prior switch, which negates the peek.
- Adding a feature back is easy; removing shipped API is hard.
If per-tab caching (`HashMap<tab_id, RefMap>`) lands later, `--tab` can
be reintroduced essentially for free. That's the right time.
## Removed
- `--tab <id|label>` global flag (`cli/src/flags.rs`, `cli/src/main.rs`,
all 52 `--help` entries in `cli/src/output.rs`).
- `tab` property in `agent-browser.schema.json` and the config-options
row in `docs/src/app/configuration/page.mdx`.
- `ScopedRestore` struct, pre/post-dispatch save/restore in
`execute_command` (`cli/src/native/actions.rs`).
- `impl Default for RefMap` in `cli/src/native/element.rs` (only added
for `mem::take` in the scoped machinery).
- `e2e_tab_global_targeting`, `_snapshot`, `_snapshot_non_contiguous`,
`e2e_tab_scoped_command_preserves_outer_tab_state`,
`_isolates_refs_from_outer_tab`, `_restores_active_tab`,
`_outer_tab_closed_mid_dispatch`. 590 lines.
- The "When to use `--tab` vs `tab <id|label>`" sections in README,
docs site, and skills reference.
## Kept
- Stable tab ids (`t1`, `t2`, `t3`) with bare-integer rejection.
- User-assigned labels (`tab new --label docs [url]`), with duplicate
rejection and interchangeable use everywhere a tab ref is accepted.
- `BrowserManager::{active_tab_id, has_tab_id, resolve_tab_ref, has_label}`
accessors (still used by the remaining tab handlers).
- `TabRef::parse`, `format_tab_id`, `is_valid_label` and their unit
tests.
- Dashboard TS fix (`TabInfo.tabId` + `label`).
- `e2e_tab_close_with_tab_id_closes_active_tab` (renamed docstring to
drop the gone exclusion-list reference).
- `e2e_tab_new_with_label_can_be_switched_and_closed` (rewrite of the
previous `_and_peeked` test — now exercises only switch and close).
- `e2e_tab_switch_rejects_bare_integer` (rewrite targeting the
`tab_switch` daemon handler rather than the removed scoped path).
net: -900 lines across 12 files. `cargo fmt`, `cargo clippy -D warnings`,
all 25 non-ignored tab unit tests, all 6 tab e2e tests, and
`tsc --noEmit` on the dashboard all pass.
* fix(tabs): initialize tab_id on missing PageInfo sites
PR #892 added a required `tab_id: u32` field to `PageInfo` but missed two
initializer sites, which broke the build on the PR branch. CI never caught
this because the external-contributor workflow status was `action_required`
and never ran.
- `cli/src/native/browser.rs:395` — the `direct_page` branch of
`connect_cdp_inner` used by the cloud providers (Browserbase, Browserless,
Browser Use, Kernel, AgentCore). Use `assign_tab_id()` to get a fresh id.
- `cli/src/native/browser.rs:1580` — a unit test initializer. Use `tab_id: 1`
since the test doesn't exercise id assignment.
* feat(tabs): restore active tab and clear per-tab state for scoped --tab
Follow-up on PR #892's `--tab <id>` flag.
The original implementation called `tab_switch_by_id` directly from the
pre-dispatch block in `execute_command` but didn't touch the daemon's
per-tab state, and never restored the previously-active tab. Two concrete
issues this fixes:
1. `state.ref_map`, `state.iframe_sessions`, and `state.active_frame_id`
were left intact across the pre-dispatch switch, so `--tab N click @e1`
would try to resolve `@e1` against the scoped tab's DOM using a
backend-node id from the outer tab. In practice the click handler's
role+name fallback hid this as "element not found" errors, but on pages
where both tabs have similarly-labelled elements it could click the
wrong one.
2. The PR description promised scoped routing would "restore the previous
active tab", but the implementation permanently switched. `--tab 3
snapshot` would leave tab 3 as the active tab even after the command
returned, surprising subsequent non-scoped commands.
This change:
- Saves the current tab's stable `tab_id` (not its array index, which
would shift if the scoped command closed other tabs) before switching.
- Clears per-tab daemon state before the switch so refs/iframes/frame
context can't leak between tabs.
- After the action runs, restores the original active tab (also via
stable id) unless that tab was closed during the scoped command, in
which case we leave the scoped tab active.
- Adds `BrowserManager::active_tab_id()` and `has_tab_id()` accessors
to support the above without exposing the internal `pages` vector.
* test(tabs): regression tests for scoped --tab state clearing and restoration
Three new `#[ignore]` e2e tests pinning the fixed behavior:
- `e2e_tab_scoped_command_clears_state_on_switch` — populates `ref_map` on
tab 1, runs a `tabId: 2`-scoped command, asserts `ref_map`,
`iframe_sessions`, and `active_frame_id` are all cleared.
- `e2e_tab_scoped_command_restores_active_tab` — sets up two tabs, runs
a scoped command against the non-active one, asserts a subsequent
unscoped command reflects the originally-active tab.
- `e2e_tab_scoped_command_handles_outer_tab_closed` — runs a scoped
`tab_close` that kills the outer tab itself, asserts no error and the
scoped tab becomes active.
Also updates two misleading comments in the PR's existing
`e2e_tab_global_targeting*` tests to reflect restoration semantics; the
assertions themselves were already consistent with restoration.
* docs(tabs): document stable tab IDs and --tab scoped-command flag
Per AGENTS.md, changes that users or agents would need to know about must
land in every doc surface. Fills the gaps PR #892 left:
- `README.md` — new `--tab <id>` row in the Options table, rewrite the
tab command examples to use `<id>` instead of `<n>`, add a paragraph
explaining stable tab IDs and `--tab` peek semantics.
- `docs/src/app/commands/page.mdx` — same command-example rewrite plus a
new "Stable tab IDs and `--tab`" subsection.
- `docs/src/app/configuration/page.mdx` — add `tab` row to the config
options table so JSON config users can discover it.
- `agent-browser.schema.json` — add `tab` property with description,
matching the config schema.
- `skills/agent-browser/references/commands.md` — same command-example
rewrite plus a short paragraph for agents on when to use `--tab`.
Introduces stable per-tab IDs and a global `--tab <id>` flag for scoping individual commands to a specific tab.
Breaking change: response payloads for `tab_list`, `tab_new`, `tab_switch`, `tab_close`, and `window_new` now use `tabId` instead of `index`. `tab_close` returns `{tabId, closed: true}` instead of `{closed, activeIndex}`. `agent-browser tab <unknown>` now errors instead of silently listing tabs.
Follow-up PR to land immediately after this fixes a compile error on the provider direct-page path, clears per-tab daemon state around scoped switches, and implements active-tab restoration so `--tab N` is non-intrusive as intended.
* fix: load storage state at launch when --state / AGENT_BROWSER_STATE is set
The `--state` flag and `AGENT_BROWSER_STATE` env var were documented as
restoring saved browser state (cookies + localStorage) at launch, but
`load_state()` was never called after the browser started. The feature
has been broken since it was introduced.
Adds `try_load_storage_state()` and calls it from every early-return
path in `auto_launch()` (lazy launch triggered by commands like
`navigate`) and from `handle_launch()` (explicit `launch` command).
Also adds 4 e2e tests covering all state-persistence paths:
- Explicit launch with `storageState` field
- Auto-launch via `AGENT_BROWSER_STATE` env var
- Session-name auto-restore via `try_auto_restore_state`
- Explicit `state_load` command (baseline sanity check)
Fixes#1164.
* style: apply cargo fmt to e2e_tests.rs
Reformats a single long format\! call to satisfy CI's rustfmt check.
No behavior change.
* fix: call try_load_storage_state in all handle_launch branches
The CDP URL, CDP port, auto-connect, and provider early-return branches
were skipping storage state loading because try_load_storage_state was
only called in the normal BrowserManager::launch() path at the bottom
of handle_launch().
Also compute storage_state_owned once and reuse it across all branches
rather than borrowing storage_state (a &str tied to cmd) in a helper
that needs an owned Option<String>.
* Fix storage state reload on reused launches
* Fix storage-state launch errors
* Fix storage state replay ordering
* Align storage-state errors across launch paths
* Fix storageState launch cleanup
Chrome's `Page.startScreencast` `maxWidth`/`maxHeight` are upper bounds,
and early frames can arrive before the viewport resize fully takes effect.
Instead of asserting exact JPEG dimensions on the first frame, skip frames
with stale dimensions and wait for one that matches.
* fix: inherit viewport dimensions in recording context
When `record start` creates a new browser context, it now re-applies the
current viewport settings (from `set viewport` or `set device`) so the
recording resolution matches what the user configured instead of falling
back to the default 1280×720.
Closes#1207
* style: apply cargo fmt to e2e test
* chore: remove obvious comments
* chore: remove obvious comments from e2e test
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
* fix: use custom viewport dimensions in streaming frame metadata
CDP's Page.screencastFrame metadata returns physical device dimensions
instead of the emulated viewport, causing frame messages to report
incorrect deviceWidth/deviceHeight when a custom viewport is set.
Use the viewport dimensions captured at screencast start instead of
the CDP metadata values, since the screencast image is already captured
at the configured viewport size.
Closes#1031
* fix: resize browser content area on viewport change for correct
screencast dimensions
Emulation.setDeviceMetricsOverride only changes the CSS viewport, but
screencast captures the actual browser content area. This caused frame
images to have incorrect dimensions (e.g., 1000x451 instead of
1000x1000)
when a custom viewport was set.
- Call Browser.setContentsSize after setDeviceMetricsOverride so the
content area matches the emulated viewport
- Restart active screencast when viewport dimensions change so
maxWidth/maxHeight parameters are updated
- Skip redundant screencast restarts when dimensions are unchanged
- Extend E2E test to verify actual JPEG image dimensions, not just
metadata
* fix: pass viewport dimensions to --window-size at launch and log setContentsSize failures
- Add viewport_size to LaunchOptions so --window-size matches the
configured viewport from the start, reducing reliance on the
experimental Browser.setContentsSize CDP call at runtime
- Log Browser.setContentsSize failures instead of silently ignoring
them with let _ =
* fix: remove duplicate viewport change detection block (dead code from merge)
* fix: use log::debug! instead of eprintln! for setContentsSize failure
* revert: use eprintln! instead of log crate for setContentsSize failure
The daemon's stderr pipe is closed after startup, so log crate
subscribers cannot output during normal operation. eprintln! is
visible during startup and in tests, matching the existing convention.
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Windows: match "actively refused it" error message in
download_bytes_connection_refused test (os error 10061)
- E2E relaunch: use userAgent instead of extensions to trigger
relaunch, since extensions force headed mode which requires a
display server unavailable in CI
- E2E auth_login SPA: use addEventListener instead of inline
onsubmit for more reliable form submission prevention
* fix: support accessibility tree refs in upload command (#1107)
The upload command only accepted CSS selectors while click/fill supported
accessibility tree refs (e.g. e1, @e1, ref=e1). This resolves the API
inconsistency by reusing resolve_element_object_id for all selector types.
* style: apply cargo fmt
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
* v0.24.1
* fix: e2e test failures on CI
- e2e_relaunch_on_options_change: use headless for all launches;
the third launch only changes extensions, which is sufficient to
trigger the relaunch hash mismatch without needing an X display
- e2e_auth_login flake: reduce SPA render delay from 1200ms to 800ms
to add headroom within the 5s preferred selector window on slower
CI runners
* fix: relaunch browser when launch options change (#993)
When the daemon already held a running browser, handle_launch only
checked connection type and liveness to decide reuse. Config changes
like adding extensions to config.json were silently ignored.
Store a hash of the relaunch-relevant LaunchOptions fields and compare
on each launch command. If the hash differs the browser is closed and
relaunched with the new options.
* fmt
* fix
* fix
* fmt
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: include buttons bitmask in drag mouseMoved events
The drag handler was omitting the `buttons` field from every
`mouseMoved` event dispatched during the move phase. Without it the
browser sees `event.buttons === 0`, meaning no button is held, so
`dragstart`/`dragover`/`drop` never fire and the drop target never
receives the element.
Fix:
- Add `"buttons": 1` (left-button mask) to each `mouseMoved` sent
while the button is held.
- Add `"buttons": 1` to `mousePressed` and `"buttons": 0` to
`mouseReleased`, consistent with how `dispatch_click` handles the
same fields in interaction.rs.
- Correct the parity-test fixture for `drag`, which was supplying a
`selector` key instead of the `source` key that `handle_drag` reads.
- Add an e2e test (`e2e_drag_action_sends_buttons_during_move`) that
drives the high-level `drag` action against the existing
`html5_drag_probe` fixture and asserts that `mousemove` events carry
`buttons == 1` and that `dragstart` fires on the source element.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* style: fix rustfmt formatting in e2e drag test
---------
Co-authored-by: wangjingjing <wangjingjing.99@bytedance.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
Chrome returns loader_id: None for same-document navigations (e.g., hash
routing in SPAs). In these cases, Page.loadEventFired never fires, causing
wait_for_lifecycle to hang forever.
The fix checks nav_result.loader_id.is_some() before waiting for lifecycle
events. Also added regression test e2e_navigate_same_url_twice_should_not_hang.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix: detect externally opened tabs in --cdp mode (#1037)
Tabs opened outside of agent-browser (e.g. by the user or another CDP
client) were invisible to `tab list` because:
1. `Target.targetCreated` with chrome://newtab/ was filtered by
`is_internal_chrome_target`, and the subsequent `targetInfoChanged`
with the real URL could not update a target that was never tracked.
2. The background drain loop only ran when `request_tracking ||
har_recording` was active, so target events between commands were
silently dropped from the broadcast channel.
Fix: promote untracked targets in `targetInfoChanged` to new targets,
run the background drain unconditionally (guarded by browser presence),
and extract `apply_drained_events` to share target lifecycle processing
(attach, domain filter, iframe sessions) between execute_command and
the background drain.
* refactor: clean up HashSet import and remove call-site duplication
- Import HashSet alongside HashMap instead of using fully-qualified path
- Replace duplicated drain+apply sequence in execute_command with
drain_cdp_events_background call
* style: apply cargo fmt
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
The Rust rewrite of save_state only captured cookies and localStorage
for the current page's origin, silently dropping cross-domain data
(e.g. SSO/CAS auth cookies). This was a regression from the JS version.
Cookies: replace Network.getCookies with Network.getAllCookies to
return cookies from all domains the browser has visited.
localStorage: track visited origins in BrowserManager during navigation,
then collect their localStorage via a temporary CDP target with Fetch
interception (serves blank HTML to avoid real network requests).
Co-authored-by: hyunjinee <leehj0110@kakao.com>
* fix: add ref for cursor-interactive content roles
* fix: format
* feat: always include cursor-interactive elements in snapshot, -C is deprecated
* feat: process StaticText aggregation and deduplication
* update test
* clean up
* fix: escape text of elements in snapshot
* fix: redundant slicing
* fix: cargo fmt
* feat: deduplicate redundant StaticText
---------
Co-authored-by: 羲洋 <lipengyang.lpy@alibaba-inc.com>
Navigate with load, then wait for username/password/submit selectors using the default action timeout. This avoids networkidle hangs on pages with continuous background requests.
* fix: restore origin-scoped --headers persistence across commands
In the v0.20 Rust rewrite, headers passed via --headers on open were
only applied to that single navigation via Network.setExtraHTTPHeaders,
which did not persist them for subsequent commands. In v0.19
(Playwright-based), these headers persisted for all subsequent
same-origin requests.
This restores the v0.19 behavior using CDP Fetch interception:
- A background task processes Fetch.requestPaused events in real-time,
injecting origin-scoped headers into matching requests and continuing
non-matching requests unmodified. This avoids the deadlock that occurs
when Fetch interception pauses requests during Page.navigate or
Runtime.evaluate (which block waiting for completion).
- The same background task also handles domain filtering and route
interception, replacing the previous drain-between-commands approach
that couldn't process events during navigation or script evaluation.
Fixes:
- --headers persist for same-origin navigations without re-passing flag
- --headers persist for in-page fetch/XHR to the same origin
- --headers do not leak to cross-origin navigations or sub-resources
- `set headers` (global) is unaffected and stacks with --headers
- Domain filter Fetch interception no longer deadlocks during navigation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: apply cargo fmt formatting
* revert inaccurate comment change
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* feat: embed cursor-interactive elements into snapshot tree
* optimize format
* fix: address review feedback for e2e_snapshot_cursor_interactive unitest
---------
Co-authored-by: 羲洋 <lipengyang.lpy@alibaba-inc.com>
* fix: resolve snapshot -C and screenshot --annotate hang over WSS (#841)
Root cause: sequential CDP round-trips per element in
find_cursor_interactive_elements() and collect_annotations() caused
timeouts over high-latency WSS connections (~200ms × 200+ elements
exceeds the 30s CDP timeout).
Fix:
- snapshot -C: Replace per-element CDP calls with a single JS eval
that detects cursor:pointer/onclick/tabindex elements in-browser,
then batch-resolve via DOM.querySelectorAll + concurrent
DOM.describeNode calls using join_all
- screenshot --annotate: Replace sequential DOM.resolveNode +
getRect calls with concurrent join_all, matching v0.19.0's
Promise.all() pattern
Behavioral parity with v0.19.0 (Node.js/Playwright):
- cursor:pointer detection via getComputedStyle
- Inherited cursor:pointer dedup (skip children of pointer parents)
- interactiveTags and interactive ARIA roles exclusion
- Role differentiation: clickable vs focusable
- Text dedup against ARIA tree ref names and quoted strings
- Edge case: -i -C shows cursor elements even when ARIA tree is empty
Tests:
- 5 unit tests for build_dedup_set() helper
- 3 e2e regression tests: cursor-interactive detection, annotation
scaling to 50 elements, cursor scaling to 100 elements
* fix: add hidden/aria-hidden filtering, contentEditable support, and cleanup robustness
- Restore hidden/aria-hidden element filtering in cursor-interactive JS
(was present in old code, dropped during rewrite)
- Add contentEditable detection with 'editable' role and hint
- Replace fire-and-forget cleanup with warning on failure
- Simplify build_dedup_set to use ref_map only (eliminates fragile
tree-text quote parsing; ref_map already has all ref-bearing names)
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: use correct Windows virtual-key codes for punctuation in type command
The `type` command was dropping punctuation characters like `.`, `'`, and
`#` because `char_to_key_info()` used raw ASCII codes as the
`windowsVirtualKeyCode` in CDP `Input.dispatchKeyEvent` calls. For
punctuation the ASCII value collides with unrelated VK codes — most
critically '.' (ASCII 46) equals VK_DELETE (0x2E), causing Chrome to
interpret periods as Delete key presses.
Changes:
- Add `punctuation_key_info()` with correct VK_OEM_* codes matching
Playwright's USKeyboardLayout (e.g. Period=190, Slash=191, Semicolon=186)
- Fall back to `Input.insertText` for characters without a US keyboard
mapping (emoji, CJK, etc.), matching Playwright's `keyboard.type()`
- Update e2e test to use `type` instead of `fill` workaround for email
- Add unit tests verifying VK code parity with Playwright's layout
Fixes#833
* style: fix rustfmt formatting for InsertTextParams
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
The v0.20.0 migration from Playwright to the native Rust daemon introduced
two regressions in checkbox/radio handling:
1. `is_element_checked` only read `this.checked`, which is undefined on
non-input elements. Material Design and ARIA controls use wrapper divs
with `role="checkbox"` and `aria-checked`, or hide the native input
off-screen inside a label. The function now mirrors Playwright's
`getChecked()` with follow-label retargeting: native `.checked`,
`aria-checked` for ARIA roles, `label.control` traversal, and nested
input lookup.
2. `check`/`uncheck` accepted the coordinate-based CDP click result
without verifying the state actually changed. When the AX tree's
`backendDOMNodeId` points to a hidden off-screen input (common in
Material Design), `Input.dispatchMouseEvent` hits nothing. The actions
now re-check state after clicking and fall back to a JS `.click()` on
the resolved input — matching Playwright's `_setChecked` verify step.
Adds e2e regression test covering Material Design (hidden input + ripple
overlay), ARIA-only, and native checkbox patterns.
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: gracefully fall back to role/name lookup when backend_node_id is stale
When the DOM changes between snapshot and click (common with SPAs and
dynamic UIs), the stored backend_node_id becomes invalid. Previously,
DOM.getBoxModel and DOM.resolveNode failures propagated as hard errors,
bypassing the role/name fallback path entirely. Now these failures are
caught and the code falls through to a JS-based element lookup.
Also adds resolve_object_id_by_role_name so that resolve_element_object_id
has a fallback for ref-based lookups (previously it had none), and
improves the role matching JS to correctly map implicit ARIA roles
(e.g. <input type="submit"> → "button", <a href> → "link").
Closes#805
* test: add e2e regression test for stale ref click fallback (#805)
Verifies that clicking a ref whose backend_node_id has become stale
(because the DOM was replaced by JavaScript) falls back to role/name
lookup instead of failing with "Could not compute box model".
* fix: use accessibility tree for stale ref fallback instead of JS heuristic
Replace the hand-rolled JS role/name matching (getImplicitRole,
getAccessibleName) with a re-query of Accessibility.getFullAXTree —
the same data source that built the ref map during snapshot. This
guarantees role/name matching is identical to what was stored,
preventing silent wrong-element clicks from name computation
divergence (e.g. aria-labelledby, <label for>, alt text).
Matches v0.19.0 (Playwright) behavior where getByRole always
re-queried the live accessibility tree.
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>