#5: the CLI never pointed users at the issue tracker. `--help` and `--version`
now print the issues URL (https://github.com/leeguooooo/chrome-use/issues), so
agents hitting a rough edge know where to report it (the skill already nudges).
#1: navigating to a `data:` URL over the extension relay fails with a cryptic
`net::ERR_ABORTED` on an about:blank tab (chrome.debugger/chrome.tabs can't drive
a top-frame data: navigation; it works fine under --launch). Detect that case and
explain it — use a real http(s)://file:// URL or --launch — instead of leaking
the raw code. (Surfaced while stress-testing 5 concurrent agents.)
After strict isolation (v1.5.16) a session's tracked set is only its OWN tabs, so
in `adopt_newly_opened` (run after every click to follow a pop-up) EVERY foreign
tab looks "new" relative to the session's `before` set and got adopted — a
click-heavy flow on a busy shared Chrome pulled other agents'/the user's tabs
(github, Lark, iphone-use) into the session mid-flow.
A pop-up the agent itself opened can't be told apart from a foreign tab over the
relay (the synthesized targetInfo carries no opener/window/group), so on the relay
`adopt_newly_opened` now adopts nothing: the agent drives only tabs it explicitly
created; pop-ups (OAuth/login windows) are the user's. Launched browsers (every
tab ours) still follow pop-ups. Verified live: a fresh click-heavy relay session
stays clean (only its own tabs), and 5 concurrent agents churning tabs show zero
cross-agent drift.
Several agents (and other tools opening tabs) share one real Chrome via the
relay. Previously every session adopted ALL tabs from Target.getTargets, so
another agent's tab churn polluted the list, dropped the tab being driven, and
drifted commands onto the wrong page (the W-8BEN tax tab vanished mid-flow when a
concurrent iphone-use agent opened tabs).
A tab group belongs to exactly one agent. On the relay a session now tracks and
drives ONLY the tabs it created (its own colored group) plus pop-ups its own
clicks open — it never adopts the user's or other agents' tabs:
- discover_and_attach (relay): create the session's own tab and pin it; do not
adopt any existing foreign tab.
- resync_targets (relay): never adopt unknown targets; never prune the session's
tabs on a single getTargets snapshot (multi-agent churn / cross-process-nav
gaps routinely omit live tabs) — prune only after RELAY_PRUNE_MISSES
consecutive absent snapshots (debounced), pinned active always protected.
- adopt_newly_opened: a tab that appears right after this session's action is a
pop-up we opened — record it as owned.
Launched browsers (every tab ours) keep adopting all tabs. Adds debounced_prune_ids
+ unit tests for the churn tolerance.
replMode and awaitPromise are mutually exclusive in Chrome — under replMode a
returned promise serialises to {} instead of being awaited, which broke every
fetch/async eval (e2e_domain_filter, e2e_headers, e2e_react_tree all regressed).
Enable replMode only for synchronous scripts that declare a top-level let/const
(the #38 case); promise-returning scripts keep awaitPromise — restoring the
pre-#38 await behaviour while still fixing the let-redeclaration collision.
#38: `chrome-use eval` now runs with Runtime.evaluate replMode (like the DevTools
console) — top-level `let`/`const` no longer throw "already been declared" across
successive evals (independent `eval` steps in a `test` suite collided in the
page's shared lexical scope), and top-level await is allowed. Main-world and
completion-value semantics are unchanged.
#37: core skill gains a hard rule — snapshot-first, never screenshot+coordinates
to locate form fields/buttons; `snapshot -i` now pierces cross-origin iframes and
lists their elements by @ref; screenshots are for visual checks only, and a
full-page retina screenshot often exceeds an image reader's limits.
Follow-up to #36 — make the whole interaction surface reach cross-origin OOPIFs
and stop coordinate events drifting onto the user's foreground tab over the relay.
- hover/dblclick/drag now DOM-dispatch over the relay or into an iframe (like
click already did): a coordinate Input event isn't confined to the target tab
on a busy real Chrome and can't map an OOPIF element's box to a top-viewport
point. drag does an HTML5 DnD in the element's frame; cross-frame drag errors
loudly instead of drifting.
- snapshot recurses iframes to MAX_IFRAME_DEPTH (3) instead of one level, so refs
inside nested payment/checkout widgets get a frame_id and resolve into the
right frame.
- relay tab adoption merges several Target.getTargets snapshots — a single flaky
relay snapshot was dropping live tabs (a driven tab vanished after restart).
- `type --key-events` (alias --keys) sends real per-character keyDown/keyUp
instead of Input.insertText, so autocomplete/combobox widgets that ignore the
insertText input event fire (Google address postal lookup; commits Angular
reactive forms so Save enables).
- SKILL: hard "snapshot-first, never default to screenshot+coordinates" rule;
snapshot -i pierces cross-origin iframes since v1.5.12; cross-origin iframe
driving guidance (#37).
#35: `open` auto-reattaches when the bound relay tab is gone — drops the dead
page, opens a fresh tab in the session's group, and navigates it, instead of
only `tab new` recovering.
#36: scroll and click now reach content inside cross-origin OOPIFs:
- scroll dispatches a real wheel at a viewport point (default center, --at x,y,
or --frame n) so it scrolls the iframe under the pointer, which
window.scrollBy on the top document silently no-ops on.
- over the extension relay, clicks always use DOM-dispatch instead of
coordinate Input events — a coordinate event isn't confined to the target tab
on a busy real Chrome (it drifted onto the foreground tab) and an OOPIF
element's box can't be mapped to a top-viewport point.
#33: some sites prepend runs of ZWJ/word-joiner/invisible-times/BOM to
document.title (badging/anti-scrape); left in, they polluted 'tab list', broke
text matching, and wrecked column alignment. sanitize_title() now strips
zero-width/bidi-format chars at every title ingestion point + get_title().
#34: 'screenshot <selector>' (element capture) already worked but was
undocumented; added 'screenshot --clip x,y,w,h' for an explicit pixel region
(CDP captureScreenshot clip), documented both in --help. Verified live.
chrome.debugger forbids DOM.setFileInputFiles, so 'upload' used to hard-fail on
the relay and push users to a --launch/direct-CDP session. Now it falls back to
reconstructing the File entirely in the page (Playwright/Cypress-style: build a
File from the bytes, assign input.files = dataTransfer.files, fire input/change;
for drop/paste composers like X, dispatch synthetic paste+drop with the
DataTransfer). The bytes are streamed in <1 MiB base64 chunks because the relay
tunnels CDP through native messaging (1 MiB/message cap) — a whole image as one
arg closed the channel. Verified live over the relay: an 809 KB PNG lands intact
on a file input with change firing. No more direct-CDP needed for uploads.
Driving a busy real Chrome via the relay, a single Target.getTargets call
occasionally returns a different window's tabs ('tab list hops windows'). resync
pruned every tracked page absent from that snapshot — including the agent's
explicitly-adopted (pinned) tab — after which active-target resolution fell back
to active_page_index and eval/click/snapshot drifted onto a foreign tab
(about:blank / chrome-extension:// / the user's page), breaking any 3+ step flow.
prunable_target_ids() now protects the pinned active target from snapshot-based
pruning; a genuine close still arrives as Target.targetDestroyed (event drain) and
removes it properly. Unit-tested.
The consent modal only appears on a raw remote-debugging attach or a browser we
launched with a debug port — never on the ab-connect extension relay. Append one
line per connection to ~/.chrome-use/connect-mode.log (relay | raw-port-attach |
launched | remote-ws), flagging 'CONSENT-MODAL-RISK' when a raw-port/launch path
runs while the relay was available. Lets us tell a code regression from Chrome's
own extension-debugger consent UX when the modal reappears. Best-effort, never
fails a connection. Verified: normal 'open' logs mode=relay (consent-free).
#25 — fill() didn't fire the events framework inputs / site autocomplete need:
it set value directly (bypassing React's value-tracker) and typed via
Input.insertText, so controlled components and input/change/blur listeners (e.g.
Mercari's postal-code → 都道府県 lookup) never ran though the value showed. fill
now emulates a real edit: focus, set through the element's prototype value setter
(React _valueTracker registers), then dispatch input → input → change → blur/
focusout. SELECT and contenteditable handled too. type <sel> <text> remains for
per-keystroke sites. Verified live: an input wired with input/change/blur fired
'IICB' from one fill.
#26 (ergonomics):
- 'close <tab>' now closes just that tab and prints 'Tab [tN] closed'; bare
'close' still closes the browser. Previously 'close t12' ran a browser close
and alarmingly printed 'Browser closed'.
- new 'chrome-use current': prints the active tab's stable handle (tabId + CDP
targetId + url/title), refreshed live — so an agent holds the targetId (which
survives cross-process nav) instead of re-deriving 'which tab is live' from
'tabs' every step. The deeper tab-id churn is the #21/#23 stable-targetId story.
Tests cover fill events (live), close tab-vs-browser parse, and current.
A click on a target=_blank link / window.open opened a new tab, but the active
tab stayed put, so the post-click snapshot showed the OLD page — looking exactly
like the click failed. On the relay the new tab is discovered only via getTargets
(the relay doesn't push target events to the daemon), so it went unsurfaced.
handle_click now snapshots tracked targets before the click and, after, runs a
lightweight BrowserManager::adopt_newly_opened (one getTargets, attaches only the
new target — far cheaper than a full resync) to detect a freshly-opened tab. It's
reported as openedTab {tabId,url,title} in the response (and a '→ opened new tab
[tN] <url>' hint in text mode). Default keeps focus on the current tab (so
multi-tab flows aren't hijacked, per #7/#8.1); 'click <sel> --follow' switches to
the new tab. Verified live: clicking a _blank link prints
'→ opened new tab [t13] https://example.org/'.
Completes the #24 friction items (B/C/D shipped in 770708b).
Dogfooding a canvas game over the extension relay surfaced three issues:
1. (serious) A fresh relay session's first `open` navigated one of the USER's
existing tabs instead of opening its own — in testing it replaced a
half-filled form with the target site. On connect the daemon passively
attaches to the user's tabs and pinned one as active; navigate() then drove
it. Now: on the relay (agent_group set), if the active tab isn't one this
session created, navigate() opens its own tab in the session's group first.
Off the relay (a browser we launched) reusing the active tab stays correct.
Pure helper active_index_is_owned() + regression tests.
2. (discoverability) `keydown <key>` / `keyup <key>` (hold-to-move, essential
for games/shortcuts) already existed as commands+daemon handlers but were
absent from --help and the skill, so they were undiscoverable. Documented in
--help, the core skill, and the canvas-app hint.
3. (UX) Canvas/WebGL pages expose almost no a11y tree, so `snapshot` is empty
and agents get stuck hunting refs. snapshot now detects a viewport-dominating
canvas with a sparse tree and prints a hint pointing at the screenshot +
coordinate-click + keydown/keyup path.
Verified live over the relay: `open` now lands the game in its own new tab with
the user's tabs (incl. the Rakuten recovery form) untouched; the canvas hint
fires on the game page; `close` cleans up only the session's own tab.
Multi-session over one relayed Chrome had a tab-identity fracture: each daemon
discovered targets ONCE at connect and assigned its own t<N> indices, so a tab
filled in session A was unreachable from session B — B saw a disjoint/blank set
and rebinding via 'open' piled up duplicate tabs. A stranded, still-filled tab
could not be finished from any other session.
- 'tab list' now re-syncs the live target set on every call: adopts tabs other
sessions opened (or that re-attached after a cross-process nav), drops gone
ones (clears phantom rows), and refreshes url/title from each live tab via
Target.getTargetInfo (the relay only stamps target_info on attach, so it goes
stale/blank after navigation — which made rows indistinguishable).
- 'tab list --full' now prints each tab's stable CDP targetId. Unlike t<N>
(per-session, reassigned each connect), targetId is stable across every session
on the relayed Chrome.
- 'tab <targetId>' adopts a specific pre-existing tab — including another
session's — WITHOUT reloading, so a half-filled form survives. handle_tab_switch
resyncs first, then resolves a raw targetId before falling back to t<N>/label.
- 'open <url> --reuse-tab' (alias --reuse) switches to an existing tab already on
that URL (matched by origin+path, ignoring volatile query/fragment) instead of
spawning a duplicate.
Verified live over the extension relay: a fresh session's 'tab list --full' lists
the user's real tabs with correct titles + full URLs + targetIds, and
'tab <targetId>' lands on and reads the exact stranded Rakuten account-recovery
form from the report. Unit tests cover URL normalization + --reuse-tab parsing;
full suite green. Docs: --help Tabs section + core skill multi-session guidance.
When connected to the user's real Chrome via the extension relay, sequential
commands could land on the wrong tab: `get url` returned x.com/home, then with
no navigation in between `eval` executed against x.com/notifications — so it
read the wrong page and returned nothing.
Root cause: the session's anti-drift anchor is `active_target_id` (pinned by
stable target_id), documented to be set "on every explicit open". But `open`
runs through `navigate()`, which never pinned. On the relay path `open` reuses
an existing tab via `navigate` rather than `add_page` (the only "explicit" path
that pins), so `active_target_id` stayed `None` and the session rode the fragile
`active_page_index`. A later passive tab close/reorder (drained before every
command) then drifted `eval`/`get url`/`snapshot` onto a foreign tab.
Fix:
- `navigate()` now syncs the index to the resolved active page and pins it by
target_id after a successful navigation — restoring the "pin on explicit open"
invariant for the relay path.
- `ensure_page()` pins its freshly-created tab too (matches `add_page`).
- Extract the pin-vs-index resolution into a pure `resolve_active_index()` and
cover the invariant with unit tests (pin beats stale index; falls back when
the pin is gone; survives passive background-tab discovery).
cargo fmt + clippy -D warnings clean; full suite 816 passed.
DOM.setFileInputFiles is forbidden by Chrome's chrome.debugger API, so upload
always fails over the extension relay with an opaque -32000 "Not allowed". Map
it to a clear message: file upload needs a --launch/direct-CDP session, and
point at the cookies export|set --curl workaround. Note the limit in the core
skill upload line too.
Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled
Standalone product rename across the whole repo (issue: project identity):
- Binary/package/repo/skill/docs: agent-browser[-stealth] → chrome-use
(single binary name `chrome-use`; old aliases agent-browser/abs dropped).
- Version: 0.27.0-fork.51 → 1.0.0 (drop the upstream-fork counter).
- Native-messaging host: com.agent_browser.connect → com.leeguoo.chrome_use
(CLI + ab-connect extension in lockstep — this is a breaking handshake change,
extension bumped 0.4.2 → 0.5.0, needs a Web Store republish).
- Config dir: ~/.agent-browser → ~/.chrome-use.
- README/zh: reframed from "stealth fork of agent-browser" to a standalone
product with a small `originally based on vercel-labs/agent-browser` credit.
- Kept AGENT_BROWSER_* env vars working (63 vars across the codebase; renaming
them would break every existing script/skill for no user-facing gain).
Build green, 802 unit tests pass, fmt + clippy clean. Upstream attribution to
vercel-labs/agent-browser preserved.
`open` waits for the `load` event by default. SPAs whose `load` never fires
(a long-pending XHR or a stuck sub-resource holds it open) made `open`
hard-fail after the lifecycle timeout — even though the DOM was ready and
eval/screenshot worked immediately right after.
- Graceful degradation: if the lifecycle event times out but document.readyState
is interactive/complete, navigate returns success carrying a `warning` in the
response (the CLI prints it to stderr; --json keeps the field) instead of
erroring. Only a still-loading document is a real failure.
- `open/goto/navigate` now accept `--wait-until <load|domcontentloaded|
networkidle|none>` so SPAs can return as soon as the DOM is parsed. The URL
parser skips the --wait-until value so it isn't mistaken for the URL.
- WaitUntil::as_str() for the warning label; output.rs surfaces response warnings.
Verified live: --wait-until domcontentloaded returns immediately on a page whose
load never fires; default load on the same page now succeeds at the timeout with
a clear stderr warning instead of failing. Adds parse tests for both arg orders
+ bogus value.
remove_page_by_target_id left active_target_id dangling when the pinned
page itself was removed, so resolved_active_index silently fell back to
active_page_index — which after a passive about:blank discovery can point
at a blank tab. That matches issue #7's intermittent symptom: `wait` then
eval/snapshot landing on about:blank in a --launch session.
Re-pin to the surviving active page after removing the pinned target so
the pin is never left pointing at a target that no longer exists. Adds
pure regression tests for the re-anchor invariant (BrowserManager needs a
live CDP client, so the method can't be unit-constructed directly).
The session's active tab was a bare index into `pages`, which drifts when a
foreign/user/other-session tab is passively discovered, a tab closes, or the list
reorders — so `eval`/`screenshot`/`snapshot`/`click` could land on the wrong page.
With login state that's a safety bug (a fetch firing on the wrong origin), and it
made screenshot disagree with snapshot/eval.
Pin the intended tab by stable target_id (`active_target_id`), set on every
explicit open / tab new / tab switch / connect. `active_session_id` and
`active_target_id` resolve through it (falling back to the index only if the
pinned tab is gone), so all commands stick to the agent's tab regardless of
passive churn — and they all agree.
Verified (--cdp, multi-tab): a window.open foreign tab no longer drifts eval;
tab new / switch re-pin correctly.
When connected to the user's real Chrome, mgr.close() disconnected but never
closed the tabs the session opened — so every session (especially one that
failed before calling close, or a forgotten one) left its tabs piling up in the
user's browser. Idle-timeout and shutdown have the same exit path.
Track the target_ids this session creates via Target.createTarget in
`created_targets` (only ever our own tabs — never the user's existing tabs, which
the raw-CDP path attaches to, nor other sessions'). On close(), for the connected
path (not a launched browser, which Browser.close handles wholesale), close each
of those targets — the extension maps Target.closeTarget → chrome.tabs.remove.
Verified against a throwaway --cdp Chrome: open + 2 `tab new` → 3 pages; `close`
→ back to 1 (our 2 closed, the pre-existing tab untouched).
Driving the user's real Chrome should not yank their view around. Now the agent
operates entirely in the background:
- New tabs are created with `background: true` (CreateTargetParams) so opening
one never foregrounds it (the ab-connect extension already used active:false;
this covers the raw-CDP path too).
- Dropped the two AUTO `Page.bringToFront` calls (auto-connect fresh tab, and the
internal active-page switch). The explicit `bringToFront` command is untouched —
surfacing a tab stays opt-in.
- enable_domains now sets `Emulation.setFocusEmulationEnabled(true)` so a
backgrounded agent tab still renders (screenshots work), isn't render-throttled,
and reports document.hasFocus()/visibilityState='visible' — which also removes
the "tab is hidden the whole session" bot tell.
Verified headless: hasFocus=true/visible while backgrounded; click + screenshot
still work. Default behaviour, no flag.
The fork's CI had never been green. Pre-existing failures:
- version-sync: check-version-sync.js read packages/dashboard/package.json,
which doesn't exist in this fork (workspace is just "."). Drop the dashboard
comparison; check package.json vs cli/Cargo.toml only.
- Dashboard job: `pnpm install --filter dashboard` for a non-existent package.
Remove the job.
- Format check: repo was never `cargo fmt`-clean. Ran cargo fmt (mechanical).
- Clippy -D warnings (newly enforced on Rust 1.94 stable): manual_contains in
commands.rs (.iter().any()->.contains()), question_mark in element.rs
(if-let-Err -> ?), result_large_err on the tungstenite handshake callback in
connect.rs (allow — the Result type is fixed by the accept_hdr_async contract).
- rust-cross: lightpanda::waits_for_ready_without_logs spawns a real process +
binds a socket with timing assumptions; flaky in CI. Marked #[ignore].
Also: skill docs note fork.30's relay-preferred auto-connect (plain
`agent-browser open` is dialog-free once the ab-connect extension is loaded) and
the extension's new "agent-browser-stealth" display name.
After connect+grouping worked, follow-up eval/get-title/screenshot drifted to a
foreign tab: on a shared browser, Target.targetCreated events for tabs the user
or OTHER agent sessions open stream in and are drained on every command. The
drain path routed them through add_page(), which sets active_page_index to the
new page — so the session's active tab silently jumped to a foreign tab and its
commands landed there.
Add BrowserManager::add_background_page() (push without touching active, dedup by
target_id) and use it in the event-drain path. Explicit opens (tab new, the
add-and-switch paths) keep using add_page() and still focus the new tab.
Closes the last gap in concurrent multi-agent: each session now drives its OWN
tab regardless of other sessions'/the user's tab activity.
Shared browser, separate tab groups: when an agent drives the user's real Chrome
via ab-connect, every tab it opens lands in a Chrome tab group named after its
--session (stable color per name). Each agent's tabs stay visually separated from
other agents' and from the user's own (ungrouped) tabs. Visibility is NOT
restricted — all agents still see all tabs (per design).
- CreateTargetParams gains an optional non-CDP `agentGroup` hint (skip-if-none),
so a strict real-Chrome endpoint never receives it
- BrowserManager.agent_group(): Some(session) only when ws_url == the live
ab-connect relay URL (never on launched/direct CDP); DAEMON_SESSION set at
daemon start supplies the name; emitted at all createTarget sites (transient
storage target stays None)
- ab-connect: +tabGroups permission; Target.createTarget reads agentGroup and
groups the new tab (create/reuse by title, deterministic color), best-effort
- extension 0.3.0 -> 0.4.0; re-signed crx + zip (id unchanged)
Needs the v0.4.0 extension reloaded + a build with this change to take effect.
Borrow anti-detection hardening from Scrapling/patchright, preferring native
CDP/Chrome overrides over JS lies:
- Runtime.enable is now opt-in via AGENT_BROWSER_CAPTURE_CONSOLE (default off).
It was called on every session INCLUDING CdpAttach (the user's real Chrome),
leaking the patchright/rebrowser "runtime" CDP signal and undermining the
"real browser, no lies" guarantee. Runtime.evaluate/callFunctionOn and
runIfWaitingForDebugger work without it; only console/error capture needs it.
The console/errors commands now return a hint when capture is disabled.
- Timezone alignment via native Emulation.setTimezoneOverride, opt-in with
AGENT_BROWSER_TIMEZONE=<IANA>|auto (FullLaunch only). Intl and Date both
follow with no JS artifact.
- WebRTC IP-leak handling via the --force-webrtc-ip-handling-policy Chrome
flag: auto disable_non_proxied_udp when a proxy is set (so the real IP can't
leak past the proxy); AGENT_BROWSER_BLOCK_WEBRTC=1 hides the local IP when
there is no proxy; =0 opts out.
- Opt-in canvas/audio fingerprint noise via AGENT_BROWSER_HIDE_CANVAS=1
(FullLaunch only). Session-stable seed so reads stay consistent within a
session while differing from the headless-stable hash.
Adds 5 unit tests; full suite 751 passed, 0 failed.
The daemon re-validates the CDP connection before every browsing command via
is_connection_alive() (Browser.getVersion, 3s timeout). It treated any
timeout-or-error as "dead" and tore the connection down + reconnected.
For an externally-attached browser (the stealth fork's default — the user's
real Chrome), a timed-out probe is almost always Chrome being briefly busy or
showing the Chrome 136+ "Allow remote debugging?" consent modal, which blocks
CDP responses until the user clicks Allow. Tearing the already-consented
connection down forces a reconnect that re-pops the consent prompt — repeated
on every command this becomes an endless prompt loop, and the close +
multiple new /devtools/browser WS probes storm Chrome into a freeze.
Fix: distinguish the probe outcome.
- Responded -> alive
- TransportError -> dead (WS closed/reset; user closing Chrome lands here too,
so zombie-socket detection is preserved)
- TimedOut -> alive for an external attach (don't tear down a consented
connection on transient slowness); dead for a browser we
launched ourselves (a real hang worth reconnecting, and no
consent modal in play).
Extracted the verdict into a pure connection_alive_from_probe() with unit
tests covering all outcomes. No behavior change for locally-launched browsers.
* feat(react): first-class React introspection, Web Vitals, and nextjs skill
Add React-general and web-universal features as first-class agent-browser verbs
(react tree/inspect/renders/suspense, vitals, pushstate). Genuinely Next.js-specific
workflows (PPR cookie protocol, /_next/mcp bridge, dev-server endpoints) ship as
a new `nextjs` skill that composes the primitives. No new runtime dependencies -
the React DevTools installHook.js is vendored (MIT) and include_str!'d into the
binary.
New commands:
react tree Full React component tree (depth id parent name)
react inspect <fiberId> Props, hooks, state, source for one fiber
react renders start|stop Fiber profiler with Insts/Mounts/Re-renders/Self/DOM
+ prev->next change details
react suspense Suspense boundaries + classifier (client-hook,
request-api, server-fetch, cache, stream, framework)
+ root-cause grouping + recommendations
vitals [url] LCP/CLS/TTFB/FCP/INP + React hydration phases
pushstate <url> Generic SPA client-side navigation
removeinitscript <id> Remove a script registered via addinitscript
New launch flags:
--init-script <path> Register init scripts before first navigation
(repeatable; env AGENT_BROWSER_INIT_SCRIPTS)
--enable <feature> Built-in init scripts; currently react-devtools
(repeatable; env AGENT_BROWSER_ENABLE)
Other primitives:
network route ... --resource-type <csv> Filter by CDP resource type
cookies set --curl <file> Auto-detects JSON/cURL/Cookie-header
* fixes
* fixes
* fixes
* fix(tabs): preserve refs across --tab peek and cover outer-tab-closed path
Follow-up to #1249 so `--tab <id>` is actually useful for agents:
- Save and restore the outer tab's `ref_map`, `iframe_sessions`, and
`active_frame_id` across a scoped command instead of clearing them.
`snapshot` → `--tab N <cmd>` → `click @e1` now keeps the outer tab's
refs intact. Scoped commands still see a clean slate so outer refs
can't resolve against the scoped tab's DOM.
- Close the coverage gap the Vercel review bot flagged on #1249: the
previous `e2e_tab_scoped_command_handles_outer_tab_closed` test used
`tab_close`, which is in the scoped-dispatch exclusion list, so it
never exercised the restore-skip branch it claimed to test. Renamed
to `e2e_tab_close_with_tab_id_closes_active_tab` with an honest
docstring, and added `e2e_tab_scoped_command_outer_tab_closed_mid_dispatch`
that actually hits the branch via `window.opener.close()` on a
script-opened intermediate tab.
- Add `e2e_tab_scoped_command_isolates_refs_from_outer_tab` pinning
that outer refs don't bleed into the scoped tab's DOM resolution.
- Rewrite `e2e_tab_scoped_command_clears_state_on_switch` as
`e2e_tab_scoped_command_preserves_outer_tab_state`, verifying the
restored @e1 still clicks end-to-end.
- Update the 52 `--help` entries for `--tab <id>` to describe peek /
restore semantics instead of a vague "Target specific tab ID".
- Update README, docs site, config schema, and the agent-facing
skills reference with working examples (refs survive the peek) and
a "when to use \`--tab <id>\` vs \`tab <id>\`" guide so agents pick
the right flag for their workflow.
* fix(tabs): use t<N> prefix for tab ids, add --label for named tabs
Follow-on to the tab work in #1249 and the prior commit, redesigning the
tab handle surface before release since nothing ships these features yet.
## Why
Incrementing integer tab ids (`1`, `2`, `3`) look indistinguishable from
positional indices in command output, LLM-generated scripts, and docs. In
the common single-agent case where position and id coincide, readers have
no visual cue for which mental model they're using. Positional indices
silently shift when unrelated tabs open/close, so misreading a handle as
an index is a correctness hazard.
## Changes
**Tab ids are now `t1`, `t2`, `t3` (strings).** Bare integer `tabId`
values are rejected with a teaching message rather than silently accepted.
The `t` prefix matches the `@e1` element-ref convention and makes ids
unmistakably non-positional at a glance.
**Labels.** Tabs can be created with a user-assigned label (e.g. `docs`,
`app`) via `tab new --label <name> [url]`. Labels are interchangeable
with `t<N>` ids everywhere a tab ref is accepted. They're never
auto-generated, never rewritten on navigation, and must be unique within
a session.
**Dashboard fix.** `packages/dashboard/src/types.ts` declared
`TabInfo.index: number` but the daemon has been sending `tabId` (not
`index`) since #892, making `tab.index` `undefined` and breaking the
dashboard's close/switch buttons silently. Updated the TS types and
usages to consume `tabId` (string) and optional `label`, restoring the
dashboard's tab interactions.
## Surface
- `cli/src/native/browser.rs`: `TabRef::parse` / `format_tab_id` /
`is_valid_label` / `PageInfo.label` / `BrowserManager::resolve_tab_ref`
/ `BrowserManager::has_label`. `tab_new` gains an optional label
argument with duplicate rejection. All JSON responses use the string
form and include the label.
- `cli/src/native/actions.rs`: scoped-command pre-dispatch and
`handle_tab_{switch,close,new}` parse string refs and resolve to
stable ids.
- `cli/src/{flags,commands,main,output}.rs`: `--tab` / config `tab`
are `String`; `tab` subcommand accepts `t<N>` or a label and supports
`tab new --label <name> [url]`. All 52 `--help` entries updated.
- `agent-browser.schema.json`: `tab` property type is now `string` with
a pattern matching `t<N>` or label form.
- `packages/dashboard`: `TabInfo.tabId: string` / `label?: string | null`;
`closeTabAtom`/`switchTabAtom` take `tabRef: string`; component props
updated.
- Docs: README, docs site (`commands/` and `configuration/`), and the
agent-facing skills reference rewritten with the new examples.
## Tests
- Added `TabRef::parse` / `format_tab_id` / `is_valid_label` unit tests
pinning the bare-integer rejection, the teaching error, label rules,
and round-tripping.
- Added `test_tab_switch_by_id` / `_by_label` / `test_tab_new_with_label`
/ `_with_label_and_url` / `_with_url_then_label` in `commands.rs`;
rewrote `test_tab_unknown_subcommand_errors` since labels make
`tab select` a legitimate ref.
- Added `e2e_tab_new_with_label_can_be_switched_and_peeked`,
`e2e_tab_new_with_duplicate_label_errors`,
`e2e_tab_scoped_command_rejects_bare_integer`.
- Migrated every existing tab e2e test (and one unit test) from
integer `tabId` to the string form.
`cargo fmt`, `cargo clippy -- -D warnings`, all 30 non-ignored tab unit
tests, all 13 tab e2e tests, and `tsc --noEmit` on the dashboard all
pass.
* refactor(tabs): drop --tab scoped peek flag; keep t<N> ids and labels
After fleshing out `--tab <id|label>` in the previous commits (scoped
pre/post-dispatch save/restore, ref preservation, outer-tab-closed edge
case, full e2e coverage), the machinery-to-value ratio makes the feature
hard to justify. Nixing it now while nothing has shipped.
## Why
- Every new daemon feature touching per-tab state has to reason about
scoped-dispatch interleaving. `ScopedRestore`, pre/post-dispatch hooks,
and the exclusion list add ongoing maintenance tax.
- Three separate PRs (#892, #1249, and this one pre-nix) were needed to
reach "works correctly." That's a smell.
- `tab <id|label>` switch + labels already cover the legible multi-tab
workflow case.
- `--tab` vs `tab <id>` have opposite lifecycle semantics but look
identical, teaching every agent two things where one would do.
- "Non-disruptive peek" isn't actually race-free: the daemon does swap
active tab during execution, so a concurrent client between pre- and
post-dispatch sees the scoped tab as active.
- Ref-based interaction with scoped tabs never worked ergonomically —
refs are per-tab, so `--tab N click @e1` requires `@e1` to already be
on tab N, which means a prior switch, which negates the peek.
- Adding a feature back is easy; removing shipped API is hard.
If per-tab caching (`HashMap<tab_id, RefMap>`) lands later, `--tab` can
be reintroduced essentially for free. That's the right time.
## Removed
- `--tab <id|label>` global flag (`cli/src/flags.rs`, `cli/src/main.rs`,
all 52 `--help` entries in `cli/src/output.rs`).
- `tab` property in `agent-browser.schema.json` and the config-options
row in `docs/src/app/configuration/page.mdx`.
- `ScopedRestore` struct, pre/post-dispatch save/restore in
`execute_command` (`cli/src/native/actions.rs`).
- `impl Default for RefMap` in `cli/src/native/element.rs` (only added
for `mem::take` in the scoped machinery).
- `e2e_tab_global_targeting`, `_snapshot`, `_snapshot_non_contiguous`,
`e2e_tab_scoped_command_preserves_outer_tab_state`,
`_isolates_refs_from_outer_tab`, `_restores_active_tab`,
`_outer_tab_closed_mid_dispatch`. 590 lines.
- The "When to use `--tab` vs `tab <id|label>`" sections in README,
docs site, and skills reference.
## Kept
- Stable tab ids (`t1`, `t2`, `t3`) with bare-integer rejection.
- User-assigned labels (`tab new --label docs [url]`), with duplicate
rejection and interchangeable use everywhere a tab ref is accepted.
- `BrowserManager::{active_tab_id, has_tab_id, resolve_tab_ref, has_label}`
accessors (still used by the remaining tab handlers).
- `TabRef::parse`, `format_tab_id`, `is_valid_label` and their unit
tests.
- Dashboard TS fix (`TabInfo.tabId` + `label`).
- `e2e_tab_close_with_tab_id_closes_active_tab` (renamed docstring to
drop the gone exclusion-list reference).
- `e2e_tab_new_with_label_can_be_switched_and_closed` (rewrite of the
previous `_and_peeked` test — now exercises only switch and close).
- `e2e_tab_switch_rejects_bare_integer` (rewrite targeting the
`tab_switch` daemon handler rather than the removed scoped path).
net: -900 lines across 12 files. `cargo fmt`, `cargo clippy -D warnings`,
all 25 non-ignored tab unit tests, all 6 tab e2e tests, and
`tsc --noEmit` on the dashboard all pass.
* fix(tabs): initialize tab_id on missing PageInfo sites
PR #892 added a required `tab_id: u32` field to `PageInfo` but missed two
initializer sites, which broke the build on the PR branch. CI never caught
this because the external-contributor workflow status was `action_required`
and never ran.
- `cli/src/native/browser.rs:395` — the `direct_page` branch of
`connect_cdp_inner` used by the cloud providers (Browserbase, Browserless,
Browser Use, Kernel, AgentCore). Use `assign_tab_id()` to get a fresh id.
- `cli/src/native/browser.rs:1580` — a unit test initializer. Use `tab_id: 1`
since the test doesn't exercise id assignment.
* feat(tabs): restore active tab and clear per-tab state for scoped --tab
Follow-up on PR #892's `--tab <id>` flag.
The original implementation called `tab_switch_by_id` directly from the
pre-dispatch block in `execute_command` but didn't touch the daemon's
per-tab state, and never restored the previously-active tab. Two concrete
issues this fixes:
1. `state.ref_map`, `state.iframe_sessions`, and `state.active_frame_id`
were left intact across the pre-dispatch switch, so `--tab N click @e1`
would try to resolve `@e1` against the scoped tab's DOM using a
backend-node id from the outer tab. In practice the click handler's
role+name fallback hid this as "element not found" errors, but on pages
where both tabs have similarly-labelled elements it could click the
wrong one.
2. The PR description promised scoped routing would "restore the previous
active tab", but the implementation permanently switched. `--tab 3
snapshot` would leave tab 3 as the active tab even after the command
returned, surprising subsequent non-scoped commands.
This change:
- Saves the current tab's stable `tab_id` (not its array index, which
would shift if the scoped command closed other tabs) before switching.
- Clears per-tab daemon state before the switch so refs/iframes/frame
context can't leak between tabs.
- After the action runs, restores the original active tab (also via
stable id) unless that tab was closed during the scoped command, in
which case we leave the scoped tab active.
- Adds `BrowserManager::active_tab_id()` and `has_tab_id()` accessors
to support the above without exposing the internal `pages` vector.
* test(tabs): regression tests for scoped --tab state clearing and restoration
Three new `#[ignore]` e2e tests pinning the fixed behavior:
- `e2e_tab_scoped_command_clears_state_on_switch` — populates `ref_map` on
tab 1, runs a `tabId: 2`-scoped command, asserts `ref_map`,
`iframe_sessions`, and `active_frame_id` are all cleared.
- `e2e_tab_scoped_command_restores_active_tab` — sets up two tabs, runs
a scoped command against the non-active one, asserts a subsequent
unscoped command reflects the originally-active tab.
- `e2e_tab_scoped_command_handles_outer_tab_closed` — runs a scoped
`tab_close` that kills the outer tab itself, asserts no error and the
scoped tab becomes active.
Also updates two misleading comments in the PR's existing
`e2e_tab_global_targeting*` tests to reflect restoration semantics; the
assertions themselves were already consistent with restoration.
* docs(tabs): document stable tab IDs and --tab scoped-command flag
Per AGENTS.md, changes that users or agents would need to know about must
land in every doc surface. Fills the gaps PR #892 left:
- `README.md` — new `--tab <id>` row in the Options table, rewrite the
tab command examples to use `<id>` instead of `<n>`, add a paragraph
explaining stable tab IDs and `--tab` peek semantics.
- `docs/src/app/commands/page.mdx` — same command-example rewrite plus a
new "Stable tab IDs and `--tab`" subsection.
- `docs/src/app/configuration/page.mdx` — add `tab` row to the config
options table so JSON config users can discover it.
- `agent-browser.schema.json` — add `tab` property with description,
matching the config schema.
- `skills/agent-browser/references/commands.md` — same command-example
rewrite plus a short paragraph for agents on when to use `--tab`.
Introduces stable per-tab IDs and a global `--tab <id>` flag for scoping individual commands to a specific tab.
Breaking change: response payloads for `tab_list`, `tab_new`, `tab_switch`, `tab_close`, and `window_new` now use `tabId` instead of `index`. `tab_close` returns `{tabId, closed: true}` instead of `{closed, activeIndex}`. `agent-browser tab <unknown>` now errors instead of silently listing tabs.
Follow-up PR to land immediately after this fixes a compile error on the provider direct-page path, clears per-tab daemon state around scoped switches, and implements active-tab restoration so `--tab N` is non-intrusive as intended.
Fix idleTimeout description to document human-friendly formats (30s,
5m, 1h) alongside raw milliseconds. Add trailing newline. Serve the
schema from the docs app at agent-browser.dev/schema.json via a
prebuild copy step, and update all $schema URLs to use the stable
docs-hosted URL instead of raw GitHub.
Adjust tab-removal bookkeeping so closing or losing a page before the active tab keeps the session pointed at the same logical page instead of silently shifting to the next one.
Add regression coverage for earlier-tab removal, later-tab removal, last-tab clamping, and the empty-page case.
Signed-off-by: Asish Kumar <officialasishkumar@gmail.com>
* fix: use custom viewport dimensions in streaming frame metadata
CDP's Page.screencastFrame metadata returns physical device dimensions
instead of the emulated viewport, causing frame messages to report
incorrect deviceWidth/deviceHeight when a custom viewport is set.
Use the viewport dimensions captured at screencast start instead of
the CDP metadata values, since the screencast image is already captured
at the configured viewport size.
Closes#1031
* fix: resize browser content area on viewport change for correct
screencast dimensions
Emulation.setDeviceMetricsOverride only changes the CSS viewport, but
screencast captures the actual browser content area. This caused frame
images to have incorrect dimensions (e.g., 1000x451 instead of
1000x1000)
when a custom viewport was set.
- Call Browser.setContentsSize after setDeviceMetricsOverride so the
content area matches the emulated viewport
- Restart active screencast when viewport dimensions change so
maxWidth/maxHeight parameters are updated
- Skip redundant screencast restarts when dimensions are unchanged
- Extend E2E test to verify actual JPEG image dimensions, not just
metadata
* fix: pass viewport dimensions to --window-size at launch and log setContentsSize failures
- Add viewport_size to LaunchOptions so --window-size matches the
configured viewport from the start, reducing reliance on the
experimental Browser.setContentsSize CDP call at runtime
- Log Browser.setContentsSize failures instead of silently ignoring
them with let _ =
* fix: remove duplicate viewport change detection block (dead code from merge)
* fix: use log::debug! instead of eprintln! for setContentsSize failure
* revert: use eprintln! instead of log crate for setContentsSize failure
The daemon's stderr pipe is closed after startup, so log crate
subscribers cannot output during normal operation. eprintln! is
visible during startup and in tests, matching the existing convention.
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Security.setIgnoreCertificateErrors is session-scoped, so creating a new
BrowserContext for recording (Target.createBrowserContext) starts with the
default certificate validation enabled, ignoring the launch-time flag.
Store ignore_https_errors in BrowserManager alongside download_path, and
re-apply Security.setIgnoreCertificateErrors to the new session after
recording context creation — matching the existing pattern for download
behavior re-application.
Fixes#1172
Co-authored-by: wangjingjing <wangjingjing.99@bytedance.com>
* fix: support accessibility tree refs in upload command (#1107)
The upload command only accepted CSS selectors while click/fill supported
accessibility tree refs (e.g. e1, @e1, ref=e1). This resolves the API
inconsistency by reusing resolve_element_object_id for all selector types.
* style: apply cargo fmt
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
When connecting to a real, already-running browser (Chrome 144+) via CDP,
targets may be paused waiting for the debugger after attach. Without an
explicit Runtime.runIfWaitingForDebugger call, page-level commands hang
indefinitely even though the WebSocket connection is live.
Add Runtime.runIfWaitingForDebugger after Runtime.enable in all target
attachment paths: enable_domains (covers initial attach, tab_new,
tab_switch), enable_domains_direct (provider proxies), and the iframe
auto-attach handler. The call is placed before Network.enable to avoid
the documented deadlock when Network.enable precedes the resume. It is
a no-op for targets that are not paused.
Fixes#1130
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* feat: add AWS Bedrock AgentCore browser provider (native Rust)
- Add agentcore provider with SigV4 authentication
- AWS SDK deps are optional behind 'agentcore' feature flag
- Build with: cargo build --features agentcore
- Supports AGENTCORE_REGION, AGENTCORE_PROFILE_ID, AGENTCORE_BROWSER_ID env vars
- Returns session ID and Live View URL in launch response
- Add connect_cdp_with_headers for signed WebSocket connections
* test: add unit tests for AgentCore provider
* refactor: use lightweight manual SigV4 signing instead of AWS SDK
- Replace aws-sigv4/aws-config with manual HMAC-SHA256 signing
- Removes ~60s compile time and significant binary size
- Credentials read from AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY env vars
- Supports AWS_SESSION_TOKEN for temporary credentials
* fix: correct AgentCore API endpoints
- Host: bedrock-agentcore.{region}.amazonaws.com
- Start session: PUT /browsers/{id}/sessions/start
- Stop session: PUT /browsers/{id}/sessions/stop
- Add urlencoding for browser ID in path
- Add AWS_DEFAULT_REGION fallback
* fix: use profileConfiguration.profileIdentifier for AgentCore profile
The AWS Bedrock AgentCore API expects profile configuration in the format:
{
"profileConfiguration": {
"profileIdentifier": "<profile-id>"
}
}
Not the flat "profileId" field that was previously used.
* feat: support AWS credential provider chain via AWS CLI
- Try env vars first (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY)
- Fall back to 'aws configure export-credentials --format env'
- Honor AWS_PROFILE environment variable
- Works with SSO, IAM roles, credential files, etc.
---------
Co-authored-by: Chris Tate <chris@ctate.dev>
Add provider icons and session creation from the dashboard UI.
Sessions can now be created with cloud providers (Browserbase,
Browserless, Browser Use, Kernel) in addition to local engines.
CLI changes:
- Track provider via .provider files alongside .engine files
- Add WaitUntil::None variant to skip lifecycle event waits for providers
- Auto-set waitUntil=none when --provider is used with navigate
- Fix Browser Use: use direct WSS connection (wss://connect.browser-use.com)
- Add connect_cdp_direct for providers with page-level CDP proxies
- Fix resolve_cdp_url to convert https:// provider URLs to wss://
- Treat empty CDP session_id as None (omit from protocol messages)
- Fix Browserbase: send explicit JSON body + Content-Type header
- Increase CDP connect timeout to 25s for remote providers
- Clean up .provider files on session close
Dashboard changes:
- Show provider or engine icon per session in sidebar
- New session dialog with unified engine/provider selector grid
- Async session creation with loading state and error display
- Kill zombie daemons on provider connection failure
- Parse CLI JSON error output for user-friendly messages
- Default new session URL to https://agent-browser.dev
Chrome returns loader_id: None for same-document navigations (e.g., hash
routing in SPAs). In these cases, Page.loadEventFired never fires, causing
wait_for_lifecycle to hang forever.
The fix checks nav_result.loader_id.is_some() before waiting for lifecycle
events. Also added regression test e2e_navigate_same_url_twice_should_not_hang.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
The Rust rewrite of save_state only captured cookies and localStorage
for the current page's origin, silently dropping cross-domain data
(e.g. SSO/CAS auth cookies). This was a regression from the JS version.
Cookies: replace Network.getCookies with Network.getAllCookies to
return cookies from all domains the browser has visited.
localStorage: track visited origins in BrowserManager during navigation,
then collect their localStorage via a temporary CDP target with Fetch
interception (serves blank HTML to avoid real network requests).
Co-authored-by: hyunjinee <leehj0110@kakao.com>
When Chrome crashes (e.g. SIGTRAP from CHECK() assertion), the daemon
now:
1. Reaps the zombie immediately via a SIGCHLD handler in the event loop
that calls waitpid(-1, WNOHANG)
2. Detects the crash instantly on the next command via a non-blocking
try_wait() check (has_process_exited), avoiding the 3-second CDP
timeout that is_connection_alive() would incur
3. Auto-relaunches Chrome transparently for the caller
Fixes#1017
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: re-apply download behavior on recording context
record start creates a new browser context via Target.createBrowserContext.
Browser.setDownloadBehavior called at launch only applies to the default
context, so downloads in the recording context are silently dropped.
Fix:
1. Store download_path on BrowserManager (from LaunchOptions)
2. After creating the recording context, call Browser.setDownloadBehavior
with the new browserContextId
This ensures downloads work during recording.
Fixes#1018
* fix: add download_path to third BrowserManager constructor (auto_connect_cdp)
When --cdp is given an HTTP/HTTPS URL (e.g. http://host:5095?mode=Hello),
resolve_cdp_url extracts host and port for CDP discovery but discards the
query string. The discovered WebSocket URL therefore never includes the
user's original query parameters, breaking relay servers that depend on
them.
Thread the original query string through discover_cdp_url and append it
to the final WebSocket URL after host/port rewriting. WebSocket URLs
(ws://, wss://) are already passed through unchanged and are unaffected.
Fixes#977
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
Navigate with load, then wait for username/password/submit selectors using the default action timeout. This avoids networkidle hangs on pages with continuous background requests.