Compare commits

...
Author SHA1 Message Date
leeguooooo 6f4e63ba91 chore(release): bump to 0.27.0-fork.9 — upstream sync + CDP consent fix
Upstream cherry-picks (onto v0.27.0 base):
- security: same-origin stream command relay (#1355)
- feat: hide scrollbars in headless screenshots (#1396)
- chore: pnpm minimum release age + node pinning (#1377, fork-adapted)

Fork fixes:
- fix(connect): stop remote-debugging consent storm — is_connection_alive no
  longer tears down an externally-attached browser on a transient liveness
  timeout (was an endless prompt loop / browser freeze)
- fix(connect): single consenting WebSocket — drop the throwaway verify probe
  so the user's one "Allow remote debugging?" click sticks to the real
  connection
2026-06-01 12:26:16 +09:00
leeguooooo 98622a7415 fix(connect): single consenting WebSocket — drop throwaway verify probe
auto-connect resolved the DevToolsActivePort URL by first opening a
verification WebSocket (verify_ws_endpoint: connect, Browser.getVersion,
close) and only then opening the real connection. On Chrome 136+ the
"Allow remote debugging?" consent is granted per-connection, so the user's
single Allow click was consumed by the throwaway probe and the real
connection (opened afterwards) asked again — surfacing as repeated prompts
or a hung command after the user had already clicked Allow.

resolve_cdp_from_active_port now gates the direct DevToolsActivePort URL on
a consent-free TCP liveness check (tcp_port_alive) instead of a WebSocket
probe, so the real connection is the single WebSocket the user consents to.
A bare TCP connect does not trigger the consent flow (that fires on the CDP
upgrade), and the real connect_async has no client-side timeout, so it waits
for the user to click Allow at their own pace. verify_ws_endpoint removed;
discovery-order tests updated, plus a guard test that resolution opens no
WebSocket.

Verified live: single prompt on a real Chrome attach, then open + eval +
scroll x2 + eval with zero re-prompts and no freeze.
2026-06-01 12:20:22 +09:00
leeguooooo 3d032f9e88 fix(connect): stop remote-debugging consent storm on transient liveness timeout
The daemon re-validates the CDP connection before every browsing command via
is_connection_alive() (Browser.getVersion, 3s timeout). It treated any
timeout-or-error as "dead" and tore the connection down + reconnected.

For an externally-attached browser (the stealth fork's default — the user's
real Chrome), a timed-out probe is almost always Chrome being briefly busy or
showing the Chrome 136+ "Allow remote debugging?" consent modal, which blocks
CDP responses until the user clicks Allow. Tearing the already-consented
connection down forces a reconnect that re-pops the consent prompt — repeated
on every command this becomes an endless prompt loop, and the close +
multiple new /devtools/browser WS probes storm Chrome into a freeze.

Fix: distinguish the probe outcome.
- Responded      -> alive
- TransportError -> dead (WS closed/reset; user closing Chrome lands here too,
                    so zombie-socket detection is preserved)
- TimedOut       -> alive for an external attach (don't tear down a consented
                    connection on transient slowness); dead for a browser we
                    launched ourselves (a real hang worth reconnecting, and no
                    consent modal in play).

Extracted the verdict into a pure connection_alive_from_probe() with unit
tests covering all outcomes. No behavior change for locally-launched browsers.
2026-06-01 11:36:00 +09:00
leeguooooo d027659571 feat(screenshot): hide scrollbars in headless screenshots (cherry-pick b4f2f37)
Cherry-picks upstream agent-browser #1396. Adds a configurable
--hide-scrollbars flag (AGENT_BROWSER_HIDE_SCROLLBARS env, hideScrollbars
config key, default true) that appends Chrome's --hide-scrollbars launch arg
for headless (non-extension) launches so native scrollbars aren't painted into
screenshots. Plumbed through flags.rs, connection.rs, main.rs, native/actions.rs
and native/cdp/chrome.rs; help text in output.rs + skill-data.

Fork adaptation:
- the arg lands in the headless && !has_extensions block, separate from the
  stealth base args — no interaction with anti-detection.
- dropped upstream docs/, agent-browser.schema.json and README hunks (removed
  or rewritten in this fork).

Verified: cargo check --tests passes.
2026-06-01 10:35:20 +09:00
leeguooooo 44b6218ef9 chore(ci): adopt upstream pnpm release-age + node pinning (cherry-pick 4ad2848)
Cherry-picks upstream agent-browser #1377 (chore: enforce pnpm minimum
release age), adapted for the fork:

- add .node-version (24); workflows read node-version-file instead of inline
- pin packageManager pnpm@11.1.3; drop hard-coded pnpm/action-setup versions
- pnpm-workspace.yaml: add minimumReleaseAge (48h supply-chain cooldown) +
  allowBuilds allowlist, keeping our trimmed packages list (no packages/*, docs)

Deliberately dropped from upstream:
- engines.node >=24 / engines.pnpm >=11 — would impose a Node 24 floor on
  end-users of the published agent-browser-stealth CLI (a compiled binary that
  doesn't need it). packageManager + .node-version cover dev/CI pinning.
- docs/ and README hunks — those paths are removed/rewritten in this fork.
2026-06-01 10:34:27 +09:00
Chris TateandMuhtasham e93acc68f8 Require same-origin stream commands (#1355)
* Require same-origin stream commands

Protect the per-session command relay from browser-originated cross-origin requests while preserving same-origin dashboard access.

Co-authored-by: Muhtasham <20128202+Muhtasham@users.noreply.github.com>

* Harden stream command origin checks

Require command relay requests to come from loopback same-origin metadata and prevent request bodies from spoofing security headers.

Co-authored-by: Muhtasham <20128202+Muhtasham@users.noreply.github.com>

---------

Co-authored-by: Muhtasham <20128202+Muhtasham@users.noreply.github.com>
2026-06-01 10:32:44 +09:00
leeguooooo d2a33cc005 fix(scripts): serialize all-platforms build + per-pid wait checks
Two related bugs that conspired to ship stale linux binaries on
0.27.0-fork.5/.7/.8 (caught only by manually grepping the embedded
version string each release):

1. build:all-platforms used `(... & npm run build:linux & wait)`.
   The bare `wait` waits for ALL children but exits with the LAST
   waited child's status, not each individually. So if linux fell
   over and windows succeeded last, the script reported success.
   Worse, when both processes shared cli/target/ and fought over
   cargo's filesystem locks, one would silently bail out and the
   missing binary just stayed at the previous release's bytes.

   Now serial: `npm run build:linux && npm run build:windows &&
   npm run build:macos`. Costs ~3 extra minutes wall-clock vs.
   parallel; trades latency for "every release ships what it says".

2. build:macos had the same `(... & ... & wait)` parallel pattern
   for arm64 + x64 cross-compiles. Native cargo builds against the
   same target/ dir share even more state than the docker'd Linux
   build did, so the failure mode is the same. Now uses explicit
   `PID1=$!; PID2=$!; wait $PID1 || exit 1; wait $PID2 || exit 1`
   so both must succeed.

Companion to the docker-compose $$ fix in 947d150 (which fixed the
*inside-container* wait+cp eating shell vars). This one fixes the
*outer* npm-script layer.
2026-05-09 12:58:39 +09:00
leeguooooo c26afbaba6 chore(release): bump to 0.27.0-fork.8 — auto-retry transient occlusion 2026-05-09 12:39:14 +09:00
leeguooooo ffb386e3af feat(click): auto-retry on transient occlusion before erroring
fork.7 caught the X mask-overlay race correctly but reported it to
the user verbatim — every transient overlay (modal backdrop, focus
ring, click-outside mask, sticky banner) became an error the user
had to wrap in their own retry loop. Most of these clear within a
frame or two on their own.

Now `verify_click_target` retries the elementFromPoint probe a few
times (default 3 × 200ms = 600ms total grace period) before failing.
Real-world overlays that blink in for a render cycle clear during
the first retry; persistent overlays still surface as errors with
the same actionable message — just qualified with "still occluded
after N retries / Mms" so the user knows we tried.

Tunable:
  AGENT_BROWSER_OCCLUSION_RETRIES         (default 3, 0 disables)
  AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS  (default 200)

DOM.resolveNode is called once outside the loop — backendNodeId is
stable across renders, only the element under (x, y) changes when
overlays flicker. Each probe is still capped at 500ms so a stuck
Runtime.callFunctionOn can't stall a click for longer than the user
expects.
2026-05-09 12:39:03 +09:00
leeguooooo 947d150561 fix(docker): escape \$ as \$\$ so docker compose doesn't eat shell vars
Real bug behind 0.27.0-fork.5 and fork.7 shipping stale linux binaries.
Docker compose interpolates \${VAR} (and \$VAR) at YAML parse time
against the host shell — including inside `command:` blocks. So:

  PID1=\$!                ← compose sees \$! → host has no `!` var → ""
  wait \$PID1 ...         ← compose sees \$PID1 → "" → becomes `wait `
  SRC="...\$TARGET..."    ← \$TARGET still works (set in `environment:`)
  cp "\$SRC" "..."        ← \$SRC eaten → empty → cp errors silently

Result: the per-PID error check I added in dbf272c never fired
because both lines were `wait` (no args) — which waits for ALL
children and exits with the LAST one's status, not each individually.
A failing arm64 build couldn't fail the script.

Fix: escape every script-local \$ as \$\$. Docker compose translates
\$\$ → literal \$ when materializing the command for the container,
and the in-container shell then expands \$VAR correctly.

Verified by `docker compose config` showing the resolved command
contains \$\$PID1 / \$\$SRC etc (which becomes \$PID1 / \$SRC in the
container's bash).
2026-05-09 11:07:01 +09:00
leeguooooo 06a29251a2 chore(release): bump to 0.27.0-fork.7 — click occlusion guard 2026-05-09 10:49:39 +09:00
leeguooooo 0eacec9b9f fix(click): occlusion check via document.elementFromPoint before dispatch
Closes the "modal silently closes when clicking 'Add post' on a thread"
bug. Verified root cause via instrumented page-side click logger:

  click @e31 (aria-label="Add post" at button (1034, 285))
  → mouse event dispatched to (1045, 296)
  → document.elementFromPoint(1045, 296) returned:
       DIV[testid="mask"], bounds (0,0,1746x934)
  → X interpreted as "click outside modal" → close + nav to /home

The cached coordinates were correct. Between snapshot and click, X
laid a transient full-viewport mask over the modal (their own
"click-outside-to-close" overlay). stealth dispatched the click
without checking what was actually at that pixel — the overlay
intercepted it.

Fix: just before returning (x, y) from resolve_element_center for
ref-based interactions, run a Runtime.callFunctionOn against the
ref's resolved element with `function(x, y) { return this.contains(
document.elementFromPoint(x, y)) || that.contains(this) ? null :
{...occluder details...}; }`. If the element at the point isn't us
(or our descendant — clicking the SVG icon inside a button is fine
— or our ancestor), we fail with a specific message:

  Ref @e31 is occluded by DIV[testid=mask] at the click point.
  A transient overlay (modal backdrop, mask, sticky banner, etc.)
  appeared between snapshot and click. Wait for it to clear or
  re-snapshot, then retry.

So instead of silently submitting an entire thread or nuking the
user's modal, agent gets a parseable error and can wait + retry.

Tight 500ms timeout per CDP call (matching the verify_ref_identity
defensive guard from fork.6) so a stuck DOM.resolveNode can't
re-introduce the multi-minute hang we just fixed. On any timeout
or error in the guard itself, fall through and let the click
proceed — strictly no worse than the unguarded code path.

Disable with AGENT_BROWSER_VERIFY_CLICK_TARGET=0.
2026-05-09 10:49:27 +09:00
leeguooooo 7159012173 chore(release): bump to 0.27.0-fork.6 — defensive-guard timeouts + accurate CDP tip 2026-05-09 10:04:25 +09:00
leeguooooo 1b3d41e579 fix(timeout): cap defensive CDP guards so click can't hang multi-minute
Reported: a single `click @ref` could hang 5+ minutes, with multiple
queued click invocations adding up to 7+ minutes — worst case 30s
timeout × 3 CDP calls × N parallel processes:

  - verify_ref_identity (Accessibility.getPartialAXTree)  →  default 30s
  - resolveNode / getBoxModel                              →  default 30s
  - wait_for_paint_settled (Runtime.evaluate awaitPromise) →  default 30s

The latter two are best-effort defenses added in fork.3-5 to fix SPA
race / DOM-reuse bugs. They should never block a real click for
30s — the unguarded code path was always faster than the guarded
path-that-hangs.

  - verify_ref_identity   capped at 1s   (skips check on timeout)
  - wait_for_paint_settled capped at 500ms (skips wait on timeout)

Both skip-on-timeout intentionally: the worst case is the click
behaves like fork.2 (race-prone but fast), which is strictly better
than the user pkilling stuck processes.

Also rewrites the misleading "Chrome 144+ chrome://inspect tip" in
the auto-connect failure message — the toggle exposes target
discovery only, not the /json/version HTTP API the auto-connect
flow expects (verified by user: lsof shows :9222 listening but
curl /json/version returns 404).
2026-05-09 10:04:14 +09:00
leeguooooo dbf272ced7 fix(docker): catch parallel-build failures + stop using glob in cp
Two latent bugs in the release pipeline that conspired to ship a stale
linux-x64 binary in 0.27.0-fork.5 (only caught by manually grepping
the embedded version string):

1. build-linux ran x64 and arm64 in parallel and used a single
   `wait $PID1 $PID2` to join them. That command waits for both, but
   its exit code is the LAST waited pid only — so if x64 silently
   broke and arm64 succeeded, the outer script exited 0 and shipped
   whatever was already in /output from the previous release. Now we
   wait on each pid individually and exit 1 on either failure.

2. build-single's cp used `agent-browser*` which globs to BOTH the
   binary and its `.d` dependency file. When two sources are passed,
   cp requires the destination to be a directory. We weren't, so cp
   exited non-zero with "Not a directory" and the build script
   shrugged it off because the next line was `chmod ... || true`.
   Now we resolve a single explicit source path.
2026-05-09 04:30:20 +09:00
leeguooooo 64140879d5 chore(release): bump to 0.27.0-fork.5 — attach-mode UX + zombie-CDP probe + wait @ref 2026-05-09 04:11:09 +09:00
leeguooooo d3bfd76c96 fix(connect): liveness probe + wait @ref support
Two changes that pair with each other:

1. connect_auto_with_fresh_tab now does a Runtime.evaluate "1"
   round-trip after creating the fresh tab. This catches the zombie
   CDP socket case (process alive, websocket dead) where every step
   up to that point reports success but the next user command would
   silently no-op against a dead session. Failing here lets the
   caller surface a proper "CDP session unresponsive" error instead
   of returning Ok and letting `agent-browser open URL` exit 0 with
   a still-blank tab.

2. handle_wait now recognizes @ref selectors (e.g. `wait @e8 --gone`).
   It polls resolve_element_object_id, which already runs the
   verify_ref_identity check from 007fd1b — so:
     - `wait @e8`             succeeds while the original element is
                              still mounted with its snapshot role+name
     - `wait @e8 --gone`      succeeds when the ref's identity changes
                              (modal closed, button re-textified, etc.)
   This gives users the "assert modal still open" primitive that
   prior versions could only approximate with screenshots.
2026-05-09 04:10:48 +09:00
leeguooooo 47dfe760be fix(cli): better message when only --headed is ignored in attach mode
In CDP-attach mode (the default since 0.24.0-fork.1), --headed has no
effect — the user's existing Chrome is already visible, and the
generic "use 'agent-browser close' first to restart" advice doesn't
help (the new daemon attaches right back). Explicitly say --headed is
moot and point to --launch as the actual escape hatch.

Other ignored flags (--profile, --proxy, etc.) keep the existing
"close + reopen" message because for those it IS the right advice.
2026-05-09 04:10:46 +09:00
leeguooooo 0db6604105 chore(release): bump to 0.27.0-fork.4 — ref identity guard 2026-05-09 03:26:48 +09:00
leeguooooo 007fd1b27f fix(refs): verify identity before using cached backendNodeId
Closes the "click @e20 hits the sibling element" bug. Real-world
example: snapshot shows @e20=[button "Add post"] next to
@e17=[button "Post all"]. By the time you click @e20, React has
re-rendered — and React often re-uses the same <button> DOM node
across renders, just updating its accessible name. The cached
backendNodeId still resolves to a real, well-positioned node, so
the click lands cleanly. It just lands on what is now the "Post all"
button, silently submitting the entire thread instead of adding a
draft row.

Before every ref-based interaction (click / fill / type / hover /
select / drag — anything routing through resolve_element_center or
resolve_element_object_id), call Accessibility.getPartialAXTree for
the cached backendNodeId and check role + name still match the
snapshot entry. On mismatch, abort with an error that names both
labels:

  Ref @e20 no longer matches its snapshot. Was [button "Add post"],
  now [button "Post all"].
  ...Take a fresh snapshot, then re-target.

If the node is gone (CDP fails / no AX node), we silently fall
through to the existing "find by role+name" recovery path, so this
guard never makes a working flow worse.

Adds one CDP roundtrip per ref interaction (~5–20ms). Disable with
AGENT_BROWSER_VERIFY_REF=0 if you control the page lifecycle and
need the latency back.
2026-05-09 03:26:20 +09:00
leeguooooo 3d1132af90 chore(release): bump to 0.27.0-fork.3 — click paint-settle + wait --gone 2026-05-09 02:45:45 +09:00
leeguooooo 90ba44cd38 feat(wait): add --gone / --hidden flags so users can fail fast on closed UIs
Pairs with the click paint-settle fix: even with that, a thread builder
that clicks "Add post" can race a misbehaving handler that closes the
parent modal instead of mounting the next textbox. To make that case
observable instead of silently corrupting the next inserttext, you can
now write:

  click @add-post
  wait .modal --gone --timeout 2000   # asserts modal stays mounted
  inserttext "tweet 3"

If the modal vanished, `wait --gone` succeeds — flip the assertion to
`wait .modal` (default visible) to fail-fast on disappearance.

Implementation just sets `state: "detached"` (or "hidden") on the wait
command — daemon-side `wait_for_selector` already supported these
states; only the CLI parser was missing the user-facing flag.

Also accepts `--detached` as alias for `--gone` to match the daemon's
internal vocabulary.
2026-05-09 02:45:34 +09:00
leeguooooo 52f8ead0f2 fix(click): wait for paint to settle so SPA renders complete before next command
Closes a real-world race that broke X multi-tweet thread composition
(and similar SPA flows): clicking "Add post" returned immediately,
inserttext fired before React had committed the new textarea, the
keystroke landed on the dialog wrapper, and X interpreted the stray
input as a request to dismiss the modal.

After mouseReleased we now wait for two requestAnimationFrame ticks
plus a microtask boundary (~33ms at 60fps, bounded). That's enough
for React/Vue/Svelte to commit any state update scheduled by the
click handler. Errors during the wait are swallowed — a click never
fails because of post-processing.

Opt out for perf-sensitive scripts that don't drive SPA UIs:
  AGENT_BROWSER_CLICK_WAIT_STABLE=0
2026-05-09 02:45:21 +09:00
leeguooooo ffa5bd63f6 chore(release): bump to 0.27.0-fork.2 — find error UX + URL preservation 2026-05-09 01:41:25 +09:00
leeguooooo 926f08203c chore: regenerate pnpm-lock.yaml after dashboard removal
The previous lockfile had ~11k lines of transitive deps for
packages/dashboard which we deleted in 86c4cff. Re-running pnpm install
shrinks it to ~24 lines (just husky for git hooks).
2026-05-09 01:41:12 +09:00
leeguooooo 2b1a3c308a feat(daemon): preserve URL across version-mismatch restart
Before: after `npm i -g` upgrade, the next agent-browser command would
detect daemon version mismatch, kill the old daemon, spawn a fresh one,
and connect to a brand-new about:blank tab. The user's previous
navigation state was silently lost — `get url` returned about:blank
even though the user's Chrome was still on the same page.

Now: before killing the old daemon, the CLI synchronously asks it for
its current URL via the existing socket. If non-empty and not
about:blank, it's persisted to a `.restore-url` sidecar in the socket
dir. After the new daemon spawns and auto-connects, it reads the
sidecar (read-and-delete), navigates the fresh tab to the saved URL,
and prints `⚠ Restored previous URL: <url>`.

Manual `agent-browser close` does NOT write the sidecar, so a clean
shutdown won't trigger surprise navigation. The sidecar is consumed on
read regardless of whether navigation succeeded, so a stale entry
can't haunt later auto-launches.
2026-05-09 01:41:07 +09:00
leeguooooo 6c556e519d feat(parse): friendly error when find has --flag where action verb expected
Before, `agent-browser find role button --name Submit` errored at the
daemon side with the cryptic `Unknown subaction: --name`. Now it errors
at parse time with the offending flag echoed back, the list of valid
actions (click, fill, check, hover, text), and a "Did you mean" hint
showing where to put the action verb.

Backwards compat: `find role button` (no flags, no action) still
defaults to click — only `--xxx` in action position errors.
2026-05-09 01:40:56 +09:00
leeguooooo 9e48b0757c fix(package): drop ./ prefix from bin entries
npm 10+ strips bin paths starting with ./ as invalid, leaving the
package with no executable entries (so `npm i -g` doesn't put any
binary on PATH). Match the upstream form `bin/agent-browser.js`.
2026-05-09 00:52:36 +09:00
leeguooooo e46232c496 chore(release): bump to 0.27.0-fork.1 on upstream v0.27.0 base 2026-05-09 00:28:02 +09:00
leeguooooo a3d4711c61 feat(skills): support npx skills add via skills.sh
- Add fork binary names (agent-browser-stealth, abs) to allowed-tools
  in all 6 SKILL.md files so installs into Claude Code / Cursor don't
  prompt for permission on every command
- Document `npx skills add leeguooooo/agent-browser-stealth` in README
- Bump README upstream-base mention from v0.24.0 to v0.27.0
2026-05-09 00:26:33 +09:00
leeguooooo 86c4cff26e chore(fork): drop upstream-only docs/, evals/, packages/dashboard, schema
These directories are TypeScript-side tooling that the fork dropped at
v0.24.0 to keep the repo focused on the stealth CLI binary. Upstream
either kept evolving them (docs, packages/dashboard) or added new ones
(evals/) — they came back during the v0.27.0 rebase, so prune again.

Also include skill-data/ in package.json `files` so the specialized
skills (electron, slack, dogfood, etc.) that upstream relocated from
skills/ to skill-data/ still ship in the npm tarball.
2026-05-09 00:24:27 +09:00
leeguoooooandClaude Opus 4.6 9202c1c919 fix(ci): add missing force_launch field in test Flags constructors
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-08 23:54:41 +09:00
leeguoooooandClaude Opus 4.6 9b56c07e33 docs: rewrite README to focus on fork differences
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:56 +09:00
leeguoooooandClaude Opus 4.6 6488aae458 chore(release): bump to 0.24.0-fork.2, publish as latest tag
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:50 +09:00
leeguoooooandClaude Opus 4.6 016d60f293 fix(docker): update Rust to 1.94 for cross-compilation builds
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:50 +09:00
leeguoooooandClaude Opus 4.6 76cfe75636 fix(stealth): achieve 0% headless via CDP-native automation override
Key insight: ANY JS-level modification to navigator.webdriver is detectable
by creepjs's lieProps system. The only undetectable approach is
Emulation.setAutomationOverride at the CDP protocol level, which tells
Chrome to natively return false for navigator.webdriver.

In CdpAttach mode, we now inject ZERO JavaScript patches — the browser's
real fingerprint is already perfect. Only the CDP protocol command is needed.

CreepJS results now match manual Chrome exactly:
- 0% headless (was 33%)
- 0% stealth (unchanged)
- 25% like headless (Chrome baseline, same as manual)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:50 +09:00
leeguoooooandClaude Opus 4.6 320bb61de3 fix(stealth): use getter-based webdriver override to match native Chrome shape
CreepJS detects three things for webDriverIsOn:
1. Property deletion (navigator.webdriver === undefined)
2. Value check (!!navigator.webdriver)
3. Lie detection (descriptor tampering via lieProps)

Changed from delete/defineProperty-value approach to replacing the CDP
getter with a getter returning false, matching the native descriptor shape.

Note: 33% headless in CreepJS is a CDP-inherent signal (lieProps detects
the getter replacement). This cannot be eliminated at the JS layer since
CDP sets the webdriver getter before init scripts run. Real-world impact
is minimal — Cloudflare Turnstile passes successfully.

Also confirmed: Chrome's remote_debugging preference in Local State
persists across restarts, so users only need to enable CDP once via
chrome://inspect/#remote-debugging.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:50 +09:00
leeguoooooandClaude Opus 4.6 81cdd3b216 fix(stealth): split minimal/full mode to eliminate detection lies on real Chrome
- CdpAttach mode: only removes navigator.webdriver (user's real Chrome
  already has genuine fingerprint, heavy patches create detectable lies)
- FullLaunch mode: applies all 32 patches (new Chrome needs full coverage)
- Improved webdriver removal: uses Object.defineProperty to override CDP
  getter on Navigator.prototype, not just delete
- CreepJS results: 0% stealth (was 20%), hasIframeProxy: gone

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:50 +09:00
leeguoooooandClaude Opus 4.6 7ee3d5fb94 feat(connect): make auto-connect to user's Chrome the default behavior
- Auto-connect is now ON by default (was opt-in via --auto-connect)
- Added --launch/--new flags to explicitly start a fresh browser
- CI environments (CI env var) automatically use --launch mode
- Friendly error message with platform-specific Chrome relaunch guide
- Mentions Chrome 144+ runtime CDP toggle (chrome://inspect)
- --cdp and --provider flags implicitly disable auto-connect
- AGENT_BROWSER_NO_AUTO_CONNECT=1 to disable, AGENT_BROWSER_FORCE_LAUNCH=1 to force

Track 3 of native-stealth migration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:48:50 +09:00
leeguoooooandClaude Opus 4.6 77616a209c feat(stealth): inject anti-detection patches in native Rust architecture
- Created cli/src/native/stealth.rs with stealth JS injection via CDP
- Extracted 32 patch IIFEs from TS stealth.ts into stealth_scripts.js
- Injected via Page.addScriptToEvaluateOnNewDocument on every launch/connect
- Added stealth Chrome args (disable AutomationControlled, use ANGLE GL)
- Auto-detects and cleans HeadlessChrome from User-Agent string
- Overrides navigator.userAgentData high-entropy hints
- Stealth enabled by default, disable with AGENT_BROWSER_STEALTH=0

Track 2 of native-stealth migration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:47:50 +09:00
leeguoooooandClaude Opus 4.6 6addc80aa1 feat(rebase): fork base on upstream v0.24.0 native architecture
- Rebased onto upstream/main (v0.24.0, full Rust native)
- Renamed package to agent-browser-stealth, version 0.24.0-fork.1
- Preserved fork-specific: abs alias, extensions/tab-group-cdp, .husky hooks
- Removed upstream-only: docs/, packages/dashboard, examples/, benchmarks/
- Simplified pnpm workspace to root-only
- Added [[bin]] section to keep binary name as "agent-browser"

Track 1 of native-stealth migration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-08 23:46:53 +09:00
Chris Tate 82eadcee41 Fix trusted publishing: add Release environment and per-job permissions (#1333) 2026-05-07 10:45:00 -05:00
Chris Tate c830d1b67d Prepare v0.27.0 release (#1332) 2026-05-07 10:15:30 -05:00
Thomas Kosiewski d33bdb36f3 Make dashboard work from proxied origins via same-origin proxy (#1111)
* Restore dashboard session proxy routes

Change-Id: I36ffc3727ce44100121bc94a81510a5f009ee0bc
Signed-off-by: Thomas Kosiewski <tk@coder.com>

* Port dashboard frontend and docs

Change-Id: I80356f64d618dab9d07b610ba67def14539f98ac
Signed-off-by: Thomas Kosiewski <tk@coder.com>

* docs: restore dashboard note in skill

Change-Id: Id0913c64e7a6f2cbbfc429ef03b34dae185d8487
Signed-off-by: Thomas Kosiewski <tk@coder.com>

* fix: tighten dashboard proxy same-origin checks

Change-Id: I792bc859a24cd47314bd46c94344ef3dfb7d6db5
Signed-off-by: Thomas Kosiewski <tk@coder.com>

---------

Signed-off-by: Thomas Kosiewski <tk@coder.com>
2026-05-07 09:08:12 -05:00
Chris Tate 3bb1d43f8b fix(doctor): make generated ids unique per call (#1330) 2026-05-06 10:48:19 -05:00
Andrew Qu 918d407411 Update README.md (#1328) 2026-05-05 16:55:19 -05:00
Walter KormanandClaude Opus 4.6 7ada3384e2 feat(docs): add AI Gateway app attribution headers (#1305)
Pass http-referer and x-title headers to streamText so Vercel can
identify agent-browser on AI Gateway pages.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-29 08:39:12 -07:00
Chris Tate 57405f9361 feat(react): React introspection, Web Vitals, and SPA primitives (#1257)
* feat(react): first-class React introspection, Web Vitals, and nextjs skill

Add React-general and web-universal features as first-class agent-browser verbs
(react tree/inspect/renders/suspense, vitals, pushstate). Genuinely Next.js-specific
workflows (PPR cookie protocol, /_next/mcp bridge, dev-server endpoints) ship as
a new `nextjs` skill that composes the primitives. No new runtime dependencies -
the React DevTools installHook.js is vendored (MIT) and include_str!'d into the
binary.

New commands:
  react tree                  Full React component tree (depth id parent name)
  react inspect <fiberId>     Props, hooks, state, source for one fiber
  react renders start|stop    Fiber profiler with Insts/Mounts/Re-renders/Self/DOM
                              + prev->next change details
  react suspense              Suspense boundaries + classifier (client-hook,
                              request-api, server-fetch, cache, stream, framework)
                              + root-cause grouping + recommendations
  vitals [url]                LCP/CLS/TTFB/FCP/INP + React hydration phases
  pushstate <url>             Generic SPA client-side navigation
  removeinitscript <id>       Remove a script registered via addinitscript

New launch flags:
  --init-script <path>        Register init scripts before first navigation
                              (repeatable; env AGENT_BROWSER_INIT_SCRIPTS)
  --enable <feature>          Built-in init scripts; currently react-devtools
                              (repeatable; env AGENT_BROWSER_ENABLE)

Other primitives:
  network route ... --resource-type <csv>  Filter by CDP resource type
  cookies set --curl <file>                Auto-detects JSON/cURL/Cookie-header

* fixes

* fixes

* fixes
2026-04-20 16:12:47 -05:00
Chris Tate cff12598bf adds trusted publishing (#1273)
* adds trusted publishing

* rename
2026-04-20 00:24:06 -05:00
Chris Tate 717d1b09e1 v0.26.0 (#1255) 2026-04-16 18:33:23 -05:00
Chris Tate 14ece9b3ad feat: add doctor command for diagnosing installs and cleaning stale daemon state (#1254)
* feat: add `doctor` command for install diagnostics and cleanup

Adds `agent-browser doctor`, a one-shot diagnostic that checks
environment, Chrome install, daemon state, config, encryption key,
providers, network reachability, and a live headless launch test.
Auto-cleans stale `.sock` / `.pid` / `.version` / `.stream` sidecar
files on every run. Destructive repairs (reinstall Chrome, purge old
state, close version-mismatched daemons, generate missing encryption
key) are gated behind `--fix`. Supports `--offline`, `--quick`, and
`--json`.

* fixes
2026-04-16 18:20:41 -05:00
Chris Tate 4cc6ca40b7 feat(skills): rename "agent-browser" skill to "core"; make CLI-served main skill actually useful (#1253)
Before this change, the main skill served by the CLI (`agent-browser
skills get agent-browser`) was a ~40-line discovery stub whose content
was essentially "run `agent-browser skills get <name>` before doing
anything." Agents already inside the CLI got no signal from it — the
content they needed to actually use the tool lived only in the `--full`
references.

Split the two jobs apart:

- **`skill-data/core/`** (new) — the runtime usage guide. 420-line
  `SKILL.md` covering the snapshot-and-ref loop, common workflows
  (login, extract, screenshot, multi-tab, sessions, iframes, dialogs),
  waiting strategies, element selection strategies, troubleshooting,
  and when to load a specialized skill. Supplementary `references/` and
  `templates/` (moved from `skills/agent-browser/`) provide the full
  command reference under `--full`.
- **`skills/agent-browser/SKILL.md`** — still the discovery stub that
  `npx skills add` installs, now marked `hidden: true` so it stays out
  of `skills list` inside the CLI. Body is a clean pointer to
  `agent-browser skills get core` and the specialized skills.

The `hidden: true` frontmatter flag is a new, general mechanism: skills
marked hidden are omitted from `skills list` and `skills get --all` but
can still be fetched by explicit name. This keeps the stub reachable
for anyone who installed via `npx skills add` without polluting the
CLI-side skill listing.

## Behavior

```
$ agent-browser skills list
  agentcore       Run agent-browser on AWS Bedrock AgentCore cloud browsers...
  core            Core agent-browser usage guide. Read this before running...
  dogfood         Systematically explore and test a web application...
  electron        Automate Electron desktop apps (VS Code, Slack, Discord...)
  slack           Interact with Slack workspaces using browser automation...
  vercel-sandbox  Run agent-browser + Chrome inside Vercel Sandbox microVMs...

$ agent-browser skills get core          # the actual usage guide
# ~420 lines of workflows, patterns, troubleshooting

$ agent-browser skills get agent-browser # still works if called explicitly
# the thin stub, now pointing at `core`
```

External `npx skills add vercel-labs/agent-browser` behavior is
unchanged: it finds and installs the thin `agent-browser` stub, which
tells the agent to run `agent-browser skills get core` for real
content. Version drift protection is preserved — the stub is the only
thing that gets copied; the real content is always runtime-fetched.

## Updated

- `cli/src/skills.rs` — `SkillInfo.hidden: bool`, parsed from
  frontmatter; `run_list` and `run_get --all` filter it. 3 new unit
  tests for the frontmatter parser.
- `cli/src/output.rs` — top-level `--help` and `skills` subcommand help
  reference `skills get core` / `skills get core --full`.
- `AGENTS.md` — "update these files for user-facing features" now
  points at `skill-data/core/` instead of the stub, with a note that
  the stub is not the right place for feature content.
- `README.md`, `docs/src/app/skills/page.mdx` — describe the new
  split and `skills get core --full` as the recommended entry point.
- `evals/cases/{command-usage,skill-selection}.ts` — expect
  `skills get core` in agent output instead of `skills get
  agent-browser`. Eval lib still reads `skills/agent-browser/SKILL.md`
  (simulating what an agent sees after `npx skills add`).

All 11 skills unit tests pass. `cargo clippy -- -D warnings` and
`cargo fmt --check` clean. Verified end-to-end: `skills list` shows
`core` + specialized (no stub), `skills get core` returns the new
content, `skills get agent-browser` still returns the stub on explicit
request.
2026-04-16 14:36:59 -05:00
Chris Tate 1afcaa0e84 docs(help): promote skills to the top of --help so agents discover them first (#1251)
The `Skills:` section was buried between `Setup:` and `Snapshot Options:` in
the top-level `--help`, where an agent skimming the output would pass over it
on the way to flag docs. Move it to a prominent "Start here (for AI agents)"
block directly below `Usage:` so it's the first thing an agent sees, and
reframe the copy so it conveys what skills *are* (workflow patterns, ref
usage, copy-paste examples) rather than just listing subcommand flags.

Skills are the intended entry point for agents. They ship with the CLI,
always version-match the installed binary, and cover both `agent-browser`
core usage and specialized workflows (Electron, Slack, exploratory testing,
cloud browser providers). Surfacing them up front prevents agents from
guessing commands out of flag docs when a hand-written workflow guide is
one command away.

No functional change. Only the ordering and wording of `--help` output.
2026-04-16 14:33:55 -05:00
Chris Tate 585d93a02b feat(tabs): t<N> prefix for tab ids; --label for named tabs; drop --tab peek flag (#1250)
* fix(tabs): preserve refs across --tab peek and cover outer-tab-closed path

Follow-up to #1249 so `--tab <id>` is actually useful for agents:

- Save and restore the outer tab's `ref_map`, `iframe_sessions`, and
  `active_frame_id` across a scoped command instead of clearing them.
  `snapshot` → `--tab N <cmd>` → `click @e1` now keeps the outer tab's
  refs intact. Scoped commands still see a clean slate so outer refs
  can't resolve against the scoped tab's DOM.
- Close the coverage gap the Vercel review bot flagged on #1249: the
  previous `e2e_tab_scoped_command_handles_outer_tab_closed` test used
  `tab_close`, which is in the scoped-dispatch exclusion list, so it
  never exercised the restore-skip branch it claimed to test. Renamed
  to `e2e_tab_close_with_tab_id_closes_active_tab` with an honest
  docstring, and added `e2e_tab_scoped_command_outer_tab_closed_mid_dispatch`
  that actually hits the branch via `window.opener.close()` on a
  script-opened intermediate tab.
- Add `e2e_tab_scoped_command_isolates_refs_from_outer_tab` pinning
  that outer refs don't bleed into the scoped tab's DOM resolution.
- Rewrite `e2e_tab_scoped_command_clears_state_on_switch` as
  `e2e_tab_scoped_command_preserves_outer_tab_state`, verifying the
  restored @e1 still clicks end-to-end.
- Update the 52 `--help` entries for `--tab <id>` to describe peek /
  restore semantics instead of a vague "Target specific tab ID".
- Update README, docs site, config schema, and the agent-facing
  skills reference with working examples (refs survive the peek) and
  a "when to use \`--tab <id>\` vs \`tab <id>\`" guide so agents pick
  the right flag for their workflow.

* fix(tabs): use t<N> prefix for tab ids, add --label for named tabs

Follow-on to the tab work in #1249 and the prior commit, redesigning the
tab handle surface before release since nothing ships these features yet.

## Why

Incrementing integer tab ids (`1`, `2`, `3`) look indistinguishable from
positional indices in command output, LLM-generated scripts, and docs. In
the common single-agent case where position and id coincide, readers have
no visual cue for which mental model they're using. Positional indices
silently shift when unrelated tabs open/close, so misreading a handle as
an index is a correctness hazard.

## Changes

**Tab ids are now `t1`, `t2`, `t3` (strings).** Bare integer `tabId`
values are rejected with a teaching message rather than silently accepted.
The `t` prefix matches the `@e1` element-ref convention and makes ids
unmistakably non-positional at a glance.

**Labels.** Tabs can be created with a user-assigned label (e.g. `docs`,
`app`) via `tab new --label <name> [url]`. Labels are interchangeable
with `t<N>` ids everywhere a tab ref is accepted. They're never
auto-generated, never rewritten on navigation, and must be unique within
a session.

**Dashboard fix.** `packages/dashboard/src/types.ts` declared
`TabInfo.index: number` but the daemon has been sending `tabId` (not
`index`) since #892, making `tab.index` `undefined` and breaking the
dashboard's close/switch buttons silently. Updated the TS types and
usages to consume `tabId` (string) and optional `label`, restoring the
dashboard's tab interactions.

## Surface

- `cli/src/native/browser.rs`: `TabRef::parse` / `format_tab_id` /
  `is_valid_label` / `PageInfo.label` / `BrowserManager::resolve_tab_ref`
  / `BrowserManager::has_label`. `tab_new` gains an optional label
  argument with duplicate rejection. All JSON responses use the string
  form and include the label.
- `cli/src/native/actions.rs`: scoped-command pre-dispatch and
  `handle_tab_{switch,close,new}` parse string refs and resolve to
  stable ids.
- `cli/src/{flags,commands,main,output}.rs`: `--tab` / config `tab`
  are `String`; `tab` subcommand accepts `t<N>` or a label and supports
  `tab new --label <name> [url]`. All 52 `--help` entries updated.
- `agent-browser.schema.json`: `tab` property type is now `string` with
  a pattern matching `t<N>` or label form.
- `packages/dashboard`: `TabInfo.tabId: string` / `label?: string | null`;
  `closeTabAtom`/`switchTabAtom` take `tabRef: string`; component props
  updated.
- Docs: README, docs site (`commands/` and `configuration/`), and the
  agent-facing skills reference rewritten with the new examples.

## Tests

- Added `TabRef::parse` / `format_tab_id` / `is_valid_label` unit tests
  pinning the bare-integer rejection, the teaching error, label rules,
  and round-tripping.
- Added `test_tab_switch_by_id` / `_by_label` / `test_tab_new_with_label`
  / `_with_label_and_url` / `_with_url_then_label` in `commands.rs`;
  rewrote `test_tab_unknown_subcommand_errors` since labels make
  `tab select` a legitimate ref.
- Added `e2e_tab_new_with_label_can_be_switched_and_peeked`,
  `e2e_tab_new_with_duplicate_label_errors`,
  `e2e_tab_scoped_command_rejects_bare_integer`.
- Migrated every existing tab e2e test (and one unit test) from
  integer `tabId` to the string form.

`cargo fmt`, `cargo clippy -- -D warnings`, all 30 non-ignored tab unit
tests, all 13 tab e2e tests, and `tsc --noEmit` on the dashboard all
pass.

* refactor(tabs): drop --tab scoped peek flag; keep t<N> ids and labels

After fleshing out `--tab <id|label>` in the previous commits (scoped
pre/post-dispatch save/restore, ref preservation, outer-tab-closed edge
case, full e2e coverage), the machinery-to-value ratio makes the feature
hard to justify. Nixing it now while nothing has shipped.

## Why

- Every new daemon feature touching per-tab state has to reason about
  scoped-dispatch interleaving. `ScopedRestore`, pre/post-dispatch hooks,
  and the exclusion list add ongoing maintenance tax.
- Three separate PRs (#892, #1249, and this one pre-nix) were needed to
  reach "works correctly." That's a smell.
- `tab <id|label>` switch + labels already cover the legible multi-tab
  workflow case.
- `--tab` vs `tab <id>` have opposite lifecycle semantics but look
  identical, teaching every agent two things where one would do.
- "Non-disruptive peek" isn't actually race-free: the daemon does swap
  active tab during execution, so a concurrent client between pre- and
  post-dispatch sees the scoped tab as active.
- Ref-based interaction with scoped tabs never worked ergonomically —
  refs are per-tab, so `--tab N click @e1` requires `@e1` to already be
  on tab N, which means a prior switch, which negates the peek.
- Adding a feature back is easy; removing shipped API is hard.

If per-tab caching (`HashMap<tab_id, RefMap>`) lands later, `--tab` can
be reintroduced essentially for free. That's the right time.

## Removed

- `--tab <id|label>` global flag (`cli/src/flags.rs`, `cli/src/main.rs`,
  all 52 `--help` entries in `cli/src/output.rs`).
- `tab` property in `agent-browser.schema.json` and the config-options
  row in `docs/src/app/configuration/page.mdx`.
- `ScopedRestore` struct, pre/post-dispatch save/restore in
  `execute_command` (`cli/src/native/actions.rs`).
- `impl Default for RefMap` in `cli/src/native/element.rs` (only added
  for `mem::take` in the scoped machinery).
- `e2e_tab_global_targeting`, `_snapshot`, `_snapshot_non_contiguous`,
  `e2e_tab_scoped_command_preserves_outer_tab_state`,
  `_isolates_refs_from_outer_tab`, `_restores_active_tab`,
  `_outer_tab_closed_mid_dispatch`. 590 lines.
- The "When to use `--tab` vs `tab <id|label>`" sections in README,
  docs site, and skills reference.

## Kept

- Stable tab ids (`t1`, `t2`, `t3`) with bare-integer rejection.
- User-assigned labels (`tab new --label docs [url]`), with duplicate
  rejection and interchangeable use everywhere a tab ref is accepted.
- `BrowserManager::{active_tab_id, has_tab_id, resolve_tab_ref, has_label}`
  accessors (still used by the remaining tab handlers).
- `TabRef::parse`, `format_tab_id`, `is_valid_label` and their unit
  tests.
- Dashboard TS fix (`TabInfo.tabId` + `label`).
- `e2e_tab_close_with_tab_id_closes_active_tab` (renamed docstring to
  drop the gone exclusion-list reference).
- `e2e_tab_new_with_label_can_be_switched_and_closed` (rewrite of the
  previous `_and_peeked` test — now exercises only switch and close).
- `e2e_tab_switch_rejects_bare_integer` (rewrite targeting the
  `tab_switch` daemon handler rather than the removed scoped path).

net: -900 lines across 12 files. `cargo fmt`, `cargo clippy -D warnings`,
all 25 non-ignored tab unit tests, all 6 tab e2e tests, and
`tsc --noEmit` on the dashboard all pass.
2026-04-16 14:33:43 -05:00
Chris Tate c201623710 fix(tabs): correct --tab scoped commands and un-break provider direct-page path (#1249)
* fix(tabs): initialize tab_id on missing PageInfo sites

PR #892 added a required `tab_id: u32` field to `PageInfo` but missed two
initializer sites, which broke the build on the PR branch. CI never caught
this because the external-contributor workflow status was `action_required`
and never ran.

- `cli/src/native/browser.rs:395` — the `direct_page` branch of
  `connect_cdp_inner` used by the cloud providers (Browserbase, Browserless,
  Browser Use, Kernel, AgentCore). Use `assign_tab_id()` to get a fresh id.
- `cli/src/native/browser.rs:1580` — a unit test initializer. Use `tab_id: 1`
  since the test doesn't exercise id assignment.

* feat(tabs): restore active tab and clear per-tab state for scoped --tab

Follow-up on PR #892's `--tab <id>` flag.

The original implementation called `tab_switch_by_id` directly from the
pre-dispatch block in `execute_command` but didn't touch the daemon's
per-tab state, and never restored the previously-active tab. Two concrete
issues this fixes:

1. `state.ref_map`, `state.iframe_sessions`, and `state.active_frame_id`
   were left intact across the pre-dispatch switch, so `--tab N click @e1`
   would try to resolve `@e1` against the scoped tab's DOM using a
   backend-node id from the outer tab. In practice the click handler's
   role+name fallback hid this as "element not found" errors, but on pages
   where both tabs have similarly-labelled elements it could click the
   wrong one.

2. The PR description promised scoped routing would "restore the previous
   active tab", but the implementation permanently switched. `--tab 3
   snapshot` would leave tab 3 as the active tab even after the command
   returned, surprising subsequent non-scoped commands.

This change:

- Saves the current tab's stable `tab_id` (not its array index, which
  would shift if the scoped command closed other tabs) before switching.
- Clears per-tab daemon state before the switch so refs/iframes/frame
  context can't leak between tabs.
- After the action runs, restores the original active tab (also via
  stable id) unless that tab was closed during the scoped command, in
  which case we leave the scoped tab active.
- Adds `BrowserManager::active_tab_id()` and `has_tab_id()` accessors
  to support the above without exposing the internal `pages` vector.

* test(tabs): regression tests for scoped --tab state clearing and restoration

Three new `#[ignore]` e2e tests pinning the fixed behavior:

- `e2e_tab_scoped_command_clears_state_on_switch` — populates `ref_map` on
  tab 1, runs a `tabId: 2`-scoped command, asserts `ref_map`,
  `iframe_sessions`, and `active_frame_id` are all cleared.
- `e2e_tab_scoped_command_restores_active_tab` — sets up two tabs, runs
  a scoped command against the non-active one, asserts a subsequent
  unscoped command reflects the originally-active tab.
- `e2e_tab_scoped_command_handles_outer_tab_closed` — runs a scoped
  `tab_close` that kills the outer tab itself, asserts no error and the
  scoped tab becomes active.

Also updates two misleading comments in the PR's existing
`e2e_tab_global_targeting*` tests to reflect restoration semantics; the
assertions themselves were already consistent with restoration.

* docs(tabs): document stable tab IDs and --tab scoped-command flag

Per AGENTS.md, changes that users or agents would need to know about must
land in every doc surface. Fills the gaps PR #892 left:

- `README.md` — new `--tab <id>` row in the Options table, rewrite the
  tab command examples to use `<id>` instead of `<n>`, add a paragraph
  explaining stable tab IDs and `--tab` peek semantics.
- `docs/src/app/commands/page.mdx` — same command-example rewrite plus a
  new "Stable tab IDs and `--tab`" subsection.
- `docs/src/app/configuration/page.mdx` — add `tab` row to the config
  options table so JSON config users can discover it.
- `agent-browser.schema.json` — add `tab` property with description,
  matching the config schema.
- `skills/agent-browser/references/commands.md` — same command-example
  rewrite plus a short paragraph for agents on when to use `--tab`.
2026-04-16 12:34:14 -05:00
Daniel Hails 67dc631977 Consistent Tab IDs & Global Tag Targeting (#892)
Introduces stable per-tab IDs and a global `--tab <id>` flag for scoping individual commands to a specific tab.

Breaking change: response payloads for `tab_list`, `tab_new`, `tab_switch`, `tab_close`, and `window_new` now use `tabId` instead of `index`. `tab_close` returns `{tabId, closed: true}` instead of `{closed, activeIndex}`. `agent-browser tab <unknown>` now errors instead of silently listing tabs.

Follow-up PR to land immediately after this fixes a compile error on the provider direct-page path, clears per-tab daemon state around scoped switches, and implements active-tab restoration so `--tab N` is non-intrusive as intended.
2026-04-16 12:02:55 -05:00
Chris Tate c691b269cb fix: improve config schema and serve from docs site (#1248)
Fix idleTimeout description to document human-friendly formats (30s,
5m, 1h) alongside raw milliseconds. Add trailing newline. Serve the
schema from the docs app at agent-browser.dev/schema.json via a
prebuild copy step, and update all $schema URLs to use the stable
docs-hosted URL instead of raw GitHub.
2026-04-16 10:42:44 -05:00
Michaelandvercel[bot] <35613825+vercel[bot]@users.noreply.github.com> 4f9edf9337 feat: add JSON Schema for agent-browser config files (#1242)
* feat: add JSON Schema for agent-browser config files

Adds agent-browser.schema.json describing all config options with
types and descriptions. Enables IDE autocomplete and validation when
referenced via $schema in agent-browser.json or
~/.agent-browser/config.json.

README and docs site updated to document the schema reference.

* fix(schema): use integer type for maxOutput to match usize deserialization

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>

---------

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
2026-04-16 08:44:01 -05:00
Tom Dale 19808d08f8 fix: load storage state at launch when --state / AGENT_BROWSER_STATE is set (#1241)
* fix: load storage state at launch when --state / AGENT_BROWSER_STATE is set

The `--state` flag and `AGENT_BROWSER_STATE` env var were documented as
restoring saved browser state (cookies + localStorage) at launch, but
`load_state()` was never called after the browser started. The feature
has been broken since it was introduced.

Adds `try_load_storage_state()` and calls it from every early-return
path in `auto_launch()` (lazy launch triggered by commands like
`navigate`) and from `handle_launch()` (explicit `launch` command).

Also adds 4 e2e tests covering all state-persistence paths:
- Explicit launch with `storageState` field
- Auto-launch via `AGENT_BROWSER_STATE` env var
- Session-name auto-restore via `try_auto_restore_state`
- Explicit `state_load` command (baseline sanity check)

Fixes #1164.

* style: apply cargo fmt to e2e_tests.rs

Reformats a single long format\! call to satisfy CI's rustfmt check.
No behavior change.

* fix: call try_load_storage_state in all handle_launch branches

The CDP URL, CDP port, auto-connect, and provider early-return branches
were skipping storage state loading because try_load_storage_state was
only called in the normal BrowserManager::launch() path at the bottom
of handle_launch().

Also compute storage_state_owned once and reuse it across all branches
rather than borrowing storage_state (a &str tied to cmd) in a helper
that needs an owned Option<String>.

* Fix storage state reload on reused launches

* Fix storage-state launch errors

* Fix storage state replay ordering

* Align storage-state errors across launch paths

* Fix storageState launch cleanup
2026-04-16 08:38:54 -05:00
Chris Tate a884960806 Prepare v0.25.5 (#1246)
* fix(test): tolerate stale screencast frames in viewport e2e test

Chrome's `Page.startScreencast` `maxWidth`/`maxHeight` are upper bounds,
and early frames can arrive before the viewport resize fully takes effect.
Instead of asserting exact JPEG dimensions on the first frame, skip frames
with stale dimensions and wait for one that matches.

* Prepare v0.25.5
2026-04-16 01:19:52 -05:00
Chris Tate dba382350b fix(test): tolerate stale screencast frames in viewport e2e test (#1245)
Chrome's `Page.startScreencast` `maxWidth`/`maxHeight` are upper bounds,
and early frames can arrive before the viewport resize fully takes effect.
Instead of asserting exact JPEG dimensions on the first frame, skip frames
with stale dimensions and wait for one that matches.
2026-04-16 00:54:29 -05:00
Chris Tate 2e99293e80 fix(ci): install ffmpeg for e2e recording test (#1244)
The `e2e_recording_inherits_viewport` test added in #1208 requires
ffmpeg on the CI runner. Without it, `recording_start` fails with
"ffmpeg not found".
2026-04-16 00:20:04 -05:00
jin.2andhyunjinee b02e485a37 fix: prefer DevToolsActivePort websocket path over HTTP discovery in --auto-connect (#1218)
* fix: prefer DevToolsActivePort websocket path over HTTP discovery in --auto-connect

Reverses the discovery order in `auto_connect_cdp()` so the exact
WebSocket path from DevToolsActivePort is tried first, falling back
to legacy HTTP endpoints (`/json/version`, `/json/list`) only when
the direct path fails. This eliminates the duplicate remote-debugging
permission prompts caused by unnecessary HTTP probes on Chrome M144+.

Also adds `verify_ws_endpoint()` to validate the WebSocket URL is a
live CDP server before returning it, preventing stale URLs from being
handed to callers.

Fixes #1210
Fixes #1206

* chore: remove unrelated issue references from test comment

* style: apply rustfmt

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
2026-04-15 17:50:11 -05:00
jin.2andhyunjinee db29d5fead fix: inherit viewport dimensions in recording context (#1208)
* fix: inherit viewport dimensions in recording context

When `record start` creates a new browser context, it now re-applies the
current viewport settings (from `set viewport` or `set device`) so the
recording resolution matches what the user configured instead of falling
back to the default 1280×720.

Closes #1207

* style: apply cargo fmt to e2e test

* chore: remove obvious comments

* chore: remove obvious comments from e2e test

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
2026-04-13 23:40:26 -05:00
Chris Tate ddf6d6a2af fix: print data for get box and get styles in text mode (#1231) (#1233)
The text-mode output formatter had branches for most `get` subcommand
response shapes but was missing handlers for `boundingbox` and `styles`.
Both commands fell through to the default "Done" message instead of
printing the returned data.

Closes #1231
2026-04-13 23:39:11 -05:00
Asish Kumar 50323499c8 fix: preserve the active page when removing earlier tabs (#1220)
Adjust tab-removal bookkeeping so closing or losing a page before the active tab keeps the session pointed at the same logical page instead of silently shifting to the next one.

Add regression coverage for earlier-tab removal, later-tab removal, last-tab clamping, and the empty-page case.

Signed-off-by: Asish Kumar <officialasishkumar@gmail.com>
2026-04-13 16:50:46 -05:00
Chris Tate 2114bdf847 Prepare v0.25.4 release (#1228) 2026-04-12 13:44:15 -05:00
Chris Tate 7c2ff0a2a6 Move specialized skills to skill-data/ so npx skills add only finds one (#1227)
The skills CLI metadata.internal flag was never implemented (PRs #587
and #652 were both closed). All 6 skills were showing in the installer.

Move the 5 specialized skills (dogfood, electron, slack, vercel-sandbox,
agentcore) from skills/ to skill-data/, which the skills CLI does not
search. The bootstrap skill stays in skills/ for discovery. The Rust CLI
searches both directories so agent-browser skills list/get still serves
all 6.
2026-04-12 13:13:04 -05:00
Chris Tate 71343069d2 Add agent-browser skills command with evals (#1225)
* Add `agent-browser skills` command

Adds a `skills` CLI command that serves bundled skill content at runtime,
always matching the installed CLI version. This solves the problem of
agents relying on stale cached SKILL.md files after CLI upgrades.

The `npx skills add vercel-labs/agent-browser` flow now installs a single
thin discovery skill with trigger words for all use cases (browser
automation, dogfooding, Electron apps, Slack, etc.) that directs agents
to `agent-browser skills get <name>` for current instructions. The other
five skills (dogfood, electron, slack, vercel-sandbox, agentcore) are
marked `metadata.internal: true` so they are not installed by default but
remain accessible via the CLI command.

Subcommands:
  skills [list]              List available skills
  skills get <name> [--full] Get skill content (with optional references)
  skills get --all           Get all skill content
  skills path [name]         Print skill directory path

* Fix skills command robustness: UTF-8 safety, flag handling, path output

- Make truncate_description UTF-8-safe using char_indices() instead of
  byte-indexed slicing that panics on multi-byte codepoints
- Pass get_all as a bool parameter to run_get instead of embedding
  --all as a sentinel string in the names list
- Canonicalize skills_dir path so `skills path` output is clean
- Warn on unrecognized flags in `skills get` instead of silently
  ignoring them

* Add evals framework and strengthen SKILL.md for better agent compliance

Strengthen SKILL.md loading instructions to require `skills get` before
running commands, and trim skill descriptions to prevent agents from
guessing at command syntax. Add TypeScript/Bun eval framework that tests
skill-loading, skill-selection, and command-usage via Claude CLI with
Vercel AI Gateway. Evals pass 20/20 (100%), up from 85% baseline.

* Fix formatting in skills.rs

* Add Codex provider to evals framework

Add multi-provider support with a shared Provider interface. Codex
provider spawns `codex exec --json`, parses JSONL output, and writes
~/.codex/config.toml for AI Gateway routing. Use `--provider codex`
to run evals with Codex (default model: openai/o3). First run scores
19/20 (95%) with 100% on skill-loading and skill-selection.

* Use scoped temp dir for Codex config instead of overwriting ~/.codex
2026-04-12 12:55:46 -05:00
Chris Tate fa043a496f fetch GitHub star count dynamically in docs header (#1202)
* fetch GitHub star count dynamically in docs header

Replace the hardcoded "27k" star count with a live fetch from the
GitHub API, revalidated every 24 hours via Next.js fetch caching.
Gracefully hides the count if the API is unreachable.

* remove GITHUB_TOKEN usage from star count fetch
2026-04-09 02:07:32 -05:00
Marshall Sun e4e2fe8633 fix(skill): correct duplicate Option numbering in auth section (#1161) 2026-04-07 01:29:01 -05:00
2164e71c30 fix: use custom viewport dimensions in streaming frame metadata and image resolution (#1033)
* fix: use custom viewport dimensions in streaming frame metadata

  CDP's Page.screencastFrame metadata returns physical device dimensions
  instead of the emulated viewport, causing frame messages to report
  incorrect deviceWidth/deviceHeight when a custom viewport is set.

  Use the viewport dimensions captured at screencast start instead of
  the CDP metadata values, since the screencast image is already captured
  at the configured viewport size.

  Closes #1031

* fix: resize browser content area on viewport change for correct
  screencast dimensions

  Emulation.setDeviceMetricsOverride only changes the CSS viewport, but
  screencast captures the actual browser content area. This caused frame
  images to have incorrect dimensions (e.g., 1000x451 instead of
  1000x1000)
  when a custom viewport was set.

  - Call Browser.setContentsSize after setDeviceMetricsOverride so the
    content area matches the emulated viewport
  - Restart active screencast when viewport dimensions change so
    maxWidth/maxHeight parameters are updated
  - Skip redundant screencast restarts when dimensions are unchanged
  - Extend E2E test to verify actual JPEG image dimensions, not just
    metadata

* fix: pass viewport dimensions to --window-size at launch and log setContentsSize failures

- Add viewport_size to LaunchOptions so --window-size matches the
  configured viewport from the start, reducing reliance on the
  experimental Browser.setContentsSize CDP call at runtime
- Log Browser.setContentsSize failures instead of silently ignoring
  them with let _ =

* fix: remove duplicate viewport change detection block (dead code from merge)

* fix: use log::debug! instead of eprintln! for setContentsSize failure

* revert: use eprintln! instead of log crate for setContentsSize failure

The daemon's stderr pipe is closed after startup, so log crate
subscribers cannot output during normal operation. eprintln! is
visible during startup and in tests, matching the existing convention.

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-07 01:28:36 -05:00
juniper929andwangjingjing 6520e4123c fix: re-apply ignore_https_errors to recording context (#1178)
Security.setIgnoreCertificateErrors is session-scoped, so creating a new
BrowserContext for recording (Target.createBrowserContext) starts with the
default certificate validation enabled, ignoring the launch-time flag.

Store ignore_https_errors in BrowserManager alongside download_path, and
re-apply Security.setIgnoreCertificateErrors to the new session after
recording context creation — matching the existing pattern for download
behavior re-application.

Fixes #1172

Co-authored-by: wangjingjing <wangjingjing.99@bytedance.com>
2026-04-07 01:26:07 -05:00
Chris Tate 6d05a9485d v0.25.3 (#1176) 2026-04-06 21:04:38 -05:00
jin.2andhyunjinee 1a6ea17ed0 fix: promote hidden radio/checkbox inputs in snapshot refs (#1085)
* fix: promote hidden radio/checkbox inputs in snapshot refs (#1024)

When a <label> wraps a display:none <input type="radio">, Chrome
excludes the input from the accessibility tree entirely. The label
appears as role="LabelText" with an empty name, making it impossible
for AI agents to identify radio buttons via data.refs.

Detect hidden radio/checkbox inputs during cursor-interactive scanning
and promote their parent LabelText/generic nodes to the correct role
with proper name and checked state.

- Add HiddenInputKind enum to validate input types at parse boundary
- Extend cursor-interactive JS to detect hidden inputs inside elements
- Extract promote_hidden_inputs() for testable role promotion logic
- Add unit tests for promotion, name preservation, and skip conditions

* style: apply cargo fmt

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
2026-04-06 20:53:13 -05:00
Chris Tate c4e0f9d367 anchors (#1175) 2026-04-06 19:59:03 -05:00
Chris Tate b75fba130b v0.25.2 (#1174) 2026-04-06 18:56:05 -05:00
Chris Tate eb15cc0894 fix: remove PR_SET_PDEATHSIG that kills Chrome after ~10s idle (#1157) (#1173)
v0.24.1 introduced `prctl(PR_SET_PDEATHSIG, SIGKILL)` in #1137 to kill Chrome
when the daemon dies. However, `PR_SET_PDEATHSIG` tracks the **thread** that
called `fork()`, not the process (`prctl(2)` documents this). Chrome is spawned
via `tokio::task::spawn_blocking`, whose threads are reaped after ~10 seconds of
idle time. When the blocking thread exits, the kernel sends SIGKILL to Chrome
even though the daemon is still alive.

Symptoms reported in #1157:
- `tab list` shows `about:blank` after a few seconds
- `snapshot` returns an empty page
- All Chrome processes exit ~9 seconds after launch
- Any workflow involving navigation or waiting breaks

The fix removes `PR_SET_PDEATHSIG` from the Chrome `pre_exec` hook. Orphan
cleanup is already handled by the process-group kill (`kill(-pgid, SIGKILL)`) in
`ChromeProcess::kill()`, which runs via daemon signal handlers, `close_notify`,
idle timeout, and `Drop`.

Fixes #1157
2026-04-06 18:44:56 -05:00
Chris Tate 7b3f826cbb v0.25.1 (#1170) 2026-04-06 10:53:37 -05:00
Chris Tate 1f8757b215 embed dashboard (#1169)
* embed dashboard

* docs

* fmt
2026-04-06 10:45:08 -05:00
Chris Tate 3896ed0d9d fix: recover GitHub release when npm published but release creation failed (#1168)
check-release now detects when the npm version matches but the GitHub
release is missing. build-binaries and github-release run in that case
so binaries, dashboard, and release notes are created without requiring
a version bump.
2026-04-06 10:22:11 -05:00
Chris Tate 92d730e5fd fix dashboard build (#1167) 2026-04-06 10:05:35 -05:00
Chris Tate 77805ff4bc v0.25.0 (#1166) 2026-04-06 09:50:13 -05:00
Chris Tate c3bbb15c5f fix: CI test failures on Windows and E2E (#1165)
- Windows: match "actively refused it" error message in
  download_bytes_connection_refused test (os error 10061)
- E2E relaunch: use userAgent instead of extensions to trigger
  relaunch, since extensions force headed mode which requires a
  display server unavailable in CI
- E2E auth_login SPA: use addEventListener instead of inline
  onsubmit for more reliable form submission prevention
2026-04-06 09:42:30 -05:00
Chris Tate 131f229971 chat (#1163)
* chat

* docs

* fmt
2026-04-06 09:21:11 -05:00
Chris Tate 317e6869b6 Add AI chat to dashboard, refactor stream module, snapshot --urls, batch argument mode (#1160)
* chat

* refactor

* fixes

* fixes

* fixes

* fixes

* improvements

* download chat

* batch

* fixes

* fixes

* fixes

* fmt

* fixes

* fixes

* fixes

* fmt
2026-04-06 08:10:43 -05:00
jin.2andhyunjinee fcb6615f5a fix: support accessibility tree refs in upload command (#1156)
* fix: support accessibility tree refs in upload command (#1107)

The upload command only accepted CSS selectors while click/fill supported
accessibility tree refs (e.g. e1, @e1, ref=e1). This resolves the API
inconsistency by reusing resolve_element_object_id for all selector types.

* style: apply cargo fmt

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
2026-04-05 15:38:49 -05:00
Chris Tateandctate c47756be9b fix(cli): honor AGENT_BROWSER_DEFAULT_TIMEOUT env var for wait commands (#1153)
* fix(cli): honor AGENT_BROWSER_DEFAULT_TIMEOUT env var for wait commands

The `AGENT_BROWSER_DEFAULT_TIMEOUT` environment variable was being ignored by CLI wait commands, causing them to use hardcoded 30-second timeouts instead of the configured default.

## Changes Made

- **Centralized timeout injection**: Modified `parse_command()` to automatically inject `flags.default_timeout` into any wait-family command that doesn't already have an explicit `--timeout` flag
- **Environment variable parsing**: Added `default_timeout` field to `Flags` struct that reads from `AGENT_BROWSER_DEFAULT_TIMEOUT` env var
- **Daemon propagation**: Updated daemon spawning to pass through the default timeout via environment variables
- **Unified timeout handling**: Added `timeout_ms()` helper method in `DaemonState` that all wait handlers now use instead of scattered `unwrap_or()` calls
- **Comprehensive test coverage**: Added 10 regression tests covering all wait command variants and edge cases

## Implementation Details

The fix uses a two-stage approach:
1. CLI parses the env var and injects timeout values into command JSON for any `wait*` action
2. Daemon reads the env var and provides a centralized fallback via `timeout_ms()` helper

This ensures new wait variants automatically inherit the default timeout without requiring per-variant wiring.

Fixes #1147

* fix: preserve 30s default timeout for backward compatibility

The default_timeout_ms fallback was set to 25_000ms, which silently
changes the existing 30_000ms behavior for users who haven't set
AGENT_BROWSER_DEFAULT_TIMEOUT. Restore the original 30s default.

---------

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-05 14:15:00 -05:00
Chris Tate 44f37c92d3 fix(cli): improve dashboard download error handling and retry logic (#1154)
This PR fixes dashboard installation failures by improving HTTP error handling and adding retry logic for network issues.

## Problem
Users were experiencing dashboard installation failures with cryptic error messages like "error sending request for url" when network issues occurred or when GitHub releases were temporarily unavailable.

## Changes
- **Enhanced HTTP client**: Added proper User-Agent, timeouts (120s total, 30s connect), and better error formatting
- **Retry logic**: Added exponential backoff retry (up to 3 attempts) for connection errors and server errors (5xx)
- **Better error messages**: Improved error formatting with full error chain context
- **Comprehensive tests**: Added unit tests for various failure scenarios (404, connection errors, partial downloads)

## Implementation Details
- Replaced direct `reqwest::get()` calls with a configured HTTP client
- Added `format_reqwest_error()` to provide detailed error context
- Implemented retry logic in `download_bytes()` with exponential backoff
- Added extensive test coverage including mock HTTP server scenarios

Fixes #1146
2026-04-05 10:07:08 -05:00
jin.2andhyunjinee 9f51879012 fix: rewrite getByRole to use CDP accessibility tree with ref-based element resolution (#1145)
* fix: rewrite getByRole to use CDP accessibility tree instead of CSS selectors

The old `handle_getbyrole` generated `querySelectorAll('[role="link"], link')`
which matched `<link>` stylesheet elements instead of `<a>` anchor tags.
This happened because ARIA role names were used directly as CSS tag selectors,
and several roles differ from their HTML element names (e.g. link → a,
heading → h1-h6, textbox → input/textarea).

The fix replaces the JS-based DOM query with the CDP `Accessibility.getFullAXTree`
API, where the browser engine correctly computes implicit ARIA roles per the
WAI-ARIA / HTML-AAM spec. This is the same approach already used by `snapshot.rs`
and `element.rs` in this codebase.

Changes:
- Rewrite `handle_getbyrole` to query the browser's accessibility tree via CDP
- Add `find_ax_node_by_role` helper for AX tree traversal with role/name/exact matching
- Use `DOM.resolveNode` + `Runtime.callFunctionOn` to bridge AX node → DOM marker
- Add iframe support via `resolve_ax_session` (missing in old implementation)
- Fix cleanup to use correct CDP session (old code used default session, breaking iframe cleanup)
- Export `extract_ax_string` as `pub(super)` for reuse
- Add 4 regression tests for `find_ax_node_by_role`

Fixes #1123

* style: apply cargo fmt

* chore: remove redundant comments

* refactor: replace marker attribute with temporary ref for element resolution

Eliminates 3 CDP round-trips (DOM.resolveNode, Runtime.callFunctionOn,
Runtime.evaluate cleanup) by registering a temporary ref in the ref_map.
execute_subaction resolves the element via backendNodeId directly.
No more DOM pollution with marker attributes.

* fix: ref counter collision, ref_map leak, and stale fallback name

- Increment next_ref_num after inserting temp ref to prevent id collision
- Remove temp ref after execute_subaction to prevent unbounded ref_map growth
- Return actual AX name from find_ax_node_by_role for accurate fallback resolution
- Add RefMap::remove method

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
2026-04-05 09:10:24 -05:00
Chris Tate 1205e2ca9c v0.24.1 (#1142)
* v0.24.1

* fix: e2e test failures on CI

- e2e_relaunch_on_options_change: use headless for all launches;
  the third launch only changes extensions, which is sufficient to
  trigger the relaunch hash mismatch without needing an X display
- e2e_auth_login flake: reduce SPA render delay from 1200ms to 800ms
  to add headroom within the 5s preferred selector window on slower
  CI runners
2026-04-04 12:49:40 -05:00
9f8e518a46 feat: reuse Chrome profile login state via --profile <name> (#1131)
* feat(chrome): add Chrome profile name resolution and copy for --profile flag

When --profile receives a name without path separators (e.g., "Default"),
it now resolves the name against installed Chrome profiles, copies the
profile to a temp directory (excluding large cache dirs), and launches
Chrome with the copied profile to reuse login state.

Key changes:
- Add profile resolution: is_chrome_profile_name, find_chrome_user_data_dir,
  list_chrome_profiles, resolve_chrome_profile (3-tier matching)
- Add copy_chrome_profile with best-effort copy and exclusion list
- Wire preprocessing into launch_chrome before retry loop
- Add use_real_keychain field to LaunchOptions for conditional keychain flags
- Make --password-store=basic and --use-mock-keychain conditional

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(cli): add `profiles` command to list available Chrome profiles

Adds `agent-browser profiles` command that reads Chrome's Local State
file to list available profiles with directory names and display names.
Supports --json output. Added help text in print_command_help and
print_help.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add Chrome profile reuse documentation across all locations

Update all 5 documentation locations per AGENTS.md:
- output.rs: updated --profile help text and examples
- README.md: added Chrome Profile Reuse section, updated options table
- SKILL.md: added profile reuse as Option 2
- docs/src/app/sessions/page.mdx: added Chrome profile reuse section
- chrome.rs: added doc comments to get_chrome_user_data_dirs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix formatting and clippy warning in chrome.rs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: simplify profile resolution and launch integration

- Only clone LaunchOptions when profile name requires resolution
  (avoids unnecessary allocation on every Chrome launch)
- Remove redundant is_file() check before copy of Local State
  (copy() handles missing files naturally)
- Extract format_profile_list() to deduplicate error formatting
- Remove unnecessary section comments in tests

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(tests): use RAII TempDir guard for test cleanup

Replace manual remove_dir_all calls with a TempDir struct that
auto-cleans on drop, preventing temp dir leaks on test panics.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 11:21:11 -05:00
Chris Tateandctate 354dd8b615 fix: pass --ignore-certificate-errors Chrome flag when --ignore-https-errors is set (#1132)
* fix: pass --ignore-certificate-errors Chrome flag when --ignore-https-errors is set

The existing CDP-level Security.setIgnoreCertificateErrors only takes
effect after Chrome opens a connection, but some TLS errors (e.g.
ERR_SSL_PROTOCOL_ERROR) are rejected at the network layer before CDP
can intervene. Adding the Chrome launch flag ensures certificate errors
are bypassed from process start.

Fixes #1124

* test: add unit tests for --ignore-certificate-errors Chrome flag

---------

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 11:15:48 -05:00
Chris Tateandctate 9b0205ef50 fix: prevent orphaned Chrome processes on daemon exit (#1137)
Three changes to ensure headless Chrome process trees are fully cleaned
up when the daemon exits, whether gracefully or abnormally:

1. Spawn Chrome in its own process group (`setpgid(0,0)`) and kill the
   entire group (`kill(-pgid, SIGKILL)`) in `ChromeProcess::kill()`.
   This takes down all helper processes (GPU, renderer, utility,
   crashpad) instead of only the main Chrome PID.

2. On Linux, set `PR_SET_PDEATHSIG(SIGKILL)` on the Chrome process so
   the kernel automatically kills it when the daemon dies for any
   reason, including SIGKILL/OOM. No macOS equivalent exists.

3. Replace `process::exit(0)` in the daemon's close handler with a
   `Notify` signal back to the main loop, so Rust destructors
   (including `ChromeProcess::Drop`) actually run.

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 11:15:26 -05:00
Chris Tateandctate c69f611d78 Fix CDP attach hang on real browser sessions (Chrome 144+) (#1133)
When connecting to a real, already-running browser (Chrome 144+) via CDP,
targets may be paused waiting for the debugger after attach. Without an
explicit Runtime.runIfWaitingForDebugger call, page-level commands hang
indefinitely even though the WebSocket connection is live.

Add Runtime.runIfWaitingForDebugger after Runtime.enable in all target
attachment paths: enable_domains (covers initial attach, tab_new,
tab_switch), enable_domains_direct (provider proxies), and the iframe
auto-attach handler. The call is placed before Network.enable to avoid
the documented deadlock when Network.enable precedes the resume. It is
a no-op for targets that are not paused.

Fixes #1130

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 11:12:35 -05:00
Chris Tateandctate 2911d91ce3 Fix stale daemon after upgrade causing silent CDP failures (#1134)
After upgrading agent-browser, the old daemon process keeps running.
ensure_daemon() only checks socket connectivity, not version, so the
new CLI silently reuses the old daemon — causing broken CDP behavior
with no error or warning.

Add a version sidecar file (.version) written by the daemon on startup.
ensure_daemon() now compares it against the CLI's compiled version and
automatically kills/restarts on mismatch. Missing version files (from
pre-fix or Node.js-era daemons) are treated as mismatches so the first
upgrade to this version also benefits.

Fixes #1127

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 11:07:30 -05:00
Chris Tateandctate 5e33672d08 fix: recover from stale daemon/socket state (#1136)
When a daemon is killed or crashes without cleaning up, stale .sock/.pid
files are left behind. Previously, `close --all` would fail to connect to
these zombie daemons and simply report an error, leaving the stale files
in place and poisoning all future sessions.

Three fixes:

1. `close --all` now force-kills unreachable daemon processes and removes
   all stale files (pid, sock, stream) instead of reporting failure. It
   also cleans up dead-but-lingering PID files during enumeration and
   scans for orphaned .sock files without corresponding .pid files.

2. `ensure_daemon` handles concurrent startup races: when a spawned
   daemon exits with "Address already in use" (another instance won the
   bind race), it checks whether the winner is accepting connections and
   piggybacks on it instead of failing.

3. `cleanup_stale_files` is now public so `close --all` can reuse it.

Fixes #1118

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 10:54:46 -05:00
jin.2andhyunjinee c976212db4 fix: idle timeout not respected due to sleep future reset in select loop (#1110)
* fix: idle timeout not respected on Unix/macOS (#1101)

The idle sleep future was recreated inside the select loop on every
iteration.  Because the drain interval ticks every 500 ms the future
was dropped and replaced before it could reach its deadline, so the
daemon never shut down.

Move the pinned Sleep future outside the loop so it survives drain
ticks and only resets on actual command receipt (reset_rx).  Apply the
same fix to the Windows path where accept events caused an identical
timer reset.

* style: apply cargo fmt

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
2026-04-04 10:51:55 -05:00
05d86fadf5 fix: relaunch browser when launch options change (#996)
* fix: relaunch browser when launch options change (#993)

  When the daemon already held a running browser, handle_launch only
  checked connection type and liveness to decide reuse. Config changes
  like adding extensions to config.json were silently ignored.

  Store a hash of the relaunch-relevant LaunchOptions fields and compare
  on each launch command. If the hash differs the browser is closed and
  relaunched with the new options.

* fmt

* fix

* fix

* fmt

---------

Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 10:41:02 -05:00
Hung-Che Lo 4b5ba9f245 fix(native): auto_launch() honours AGENT_BROWSER_PROVIDER for cloud providers (#1126)
When a non-launch command (e.g. open, snapshot) triggers auto_launch()
before the explicit launch command is processed, auto_launch() now checks
AGENT_BROWSER_PROVIDER and connects via the provider API instead of
always falling back to a local Chrome instance.

Also redirects daemon stderr to /dev/null when not in debug mode to
prevent crashes from broken pipe after the CLI drops the piped stderr
handle. Cloud providers may write to stderr during connection setup.

Fixes #1125
Related: #979
2026-04-04 10:29:04 -05:00
Chris Tateandctate c52d25d576 Fix HAR capture missing API requests under heavy traffic (#1135)
The CDP event broadcast buffer (256 events) was too small for pages with
many concurrent API requests, causing silent event drops. Modern SPAs
routinely fire 100+ API calls during page load, generating 300+ CDP
network events that would overflow the buffer between drain cycles.

Changes:
- Increase CDP broadcast buffer from 256 to 4096 (event channel) and
  512 to 4096 (raw channel)
- Reduce background drain interval from 500ms to 100ms
- Handle Network.loadingFailed events in HAR recording
- Enable Network.enable on cross-origin iframe sessions during HAR
  recording and request tracking
- Allow Network events from iframe sessions through the session filter
- Log a warning when buffer overflow occurs instead of silently dropping

Fixes #1128

Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
2026-04-04 10:25:12 -05:00
272 changed files with 23210 additions and 45300 deletions
+32 -1
View File
@@ -15,6 +15,11 @@ jobs:
- name: Checkout repository
uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: .node-version
- name: Check version sync
run: node scripts/check-version-sync.js
@@ -44,6 +49,29 @@ jobs:
- name: Run Rust tests
run: cargo test --profile ci --manifest-path cli/Cargo.toml
dashboard:
name: Dashboard
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: .node-version
- name: Install pnpm
uses: pnpm/action-setup@v4
- name: Install dependencies
run: pnpm install --filter dashboard
working-directory: packages/dashboard
- name: Build dashboard
run: pnpm build
working-directory: packages/dashboard
rust-cross:
name: Rust (${{ matrix.os }} - ${{ matrix.target }})
if: github.event_name != 'pull_request'
@@ -96,6 +124,9 @@ jobs:
run: |
cargo run --manifest-path cli/Cargo.toml -- install --with-deps
- name: Install ffmpeg
run: sudo apt-get update && sudo apt-get install -y ffmpeg
- name: Run e2e tests
run: cargo test --profile ci --manifest-path cli/Cargo.toml e2e -- --ignored --test-threads=1
@@ -181,7 +212,7 @@ jobs:
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: 22
node-version-file: .node-version
- name: Setup Rust toolchain
uses: dtolnay/rust-toolchain@stable
+46 -36
View File
@@ -9,20 +9,29 @@ on:
concurrency: ${{ github.workflow }}-${{ github.ref }}
permissions:
contents: write
contents: read
jobs:
check-release:
name: Check for new version
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
outputs:
should_release: ${{ steps.check.outputs.should_release }}
needs_github_release: ${{ steps.check.outputs.needs_github_release }}
version: ${{ steps.check.outputs.version }}
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Compare package.json version to npm
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: .node-version
- name: Compare package.json version to npm and check GitHub release
id: check
run: |
LOCAL_VERSION=$(node -p "require('./package.json').version")
@@ -34,16 +43,30 @@ jobs:
if [ "$LOCAL_VERSION" != "$NPM_VERSION" ]; then
echo "Version changed: $NPM_VERSION -> $LOCAL_VERSION"
echo "should_release=true" >> "$GITHUB_OUTPUT"
echo "needs_github_release=true" >> "$GITHUB_OUTPUT"
else
echo "Version unchanged, skipping release"
echo "Version unchanged on npm, skipping build and publish"
echo "should_release=false" >> "$GITHUB_OUTPUT"
# Check if GitHub release exists; it may be missing if a prior run
# published to npm but failed before creating the release.
TAG="v$LOCAL_VERSION"
if gh release view "$TAG" &>/dev/null; then
echo "GitHub release $TAG exists"
echo "needs_github_release=false" >> "$GITHUB_OUTPUT"
else
echo "GitHub release $TAG is missing, will rebuild and create it"
echo "needs_github_release=true" >> "$GITHUB_OUTPUT"
fi
fi
echo "version=$LOCAL_VERSION" >> "$GITHUB_OUTPUT"
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
build-binaries:
name: Build ${{ matrix.name }}
needs: check-release
if: needs.check-release.outputs.should_release == 'true'
if: needs.check-release.outputs.should_release == 'true' || needs.check-release.outputs.needs_github_release == 'true'
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
@@ -91,13 +114,11 @@ jobs:
- name: Setup pnpm
uses: pnpm/action-setup@v4
with:
version: 9
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
node-version-file: .node-version
cache: pnpm
- name: Install npm dependencies
@@ -106,6 +127,9 @@ jobs:
- name: Sync version
run: pnpm run version:sync
- name: Build dashboard
run: pnpm --filter dashboard build
- name: Setup Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -168,20 +192,24 @@ jobs:
publish:
name: Publish to npm
needs: [check-release, build-binaries]
if: needs.check-release.outputs.should_release == 'true'
runs-on: ubuntu-latest
timeout-minutes: 15
environment: Release
permissions:
contents: read
id-token: write
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Setup pnpm
uses: pnpm/action-setup@v4
with:
version: 9
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
node-version-file: .node-version
cache: pnpm
registry-url: 'https://registry.npmjs.org'
@@ -236,14 +264,16 @@ jobs:
echo "All 7 platform binaries present and valid"
- name: Publish to npm
run: pnpm publish --no-git-checks
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_VERCEL_TOKEN_ELEVATED }}
run: npm publish --provenance
github-release:
name: Create GitHub Release
needs: [check-release, publish]
needs: [check-release, build-binaries, publish]
if: always() && needs.build-binaries.result == 'success' && needs.check-release.outputs.needs_github_release == 'true'
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: write
steps:
- name: Checkout repository
uses: actions/checkout@v4
@@ -271,26 +301,6 @@ jobs:
fi
echo "Found $BINARY_COUNT binaries"
- name: Setup pnpm
uses: pnpm/action-setup@v4
with:
version: 9
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
cache: pnpm
- name: Install dependencies
run: pnpm install --frozen-lockfile
- name: Build dashboard
run: pnpm --filter dashboard build
- name: Create dashboard.zip
run: cd packages/dashboard/out && zip -r ../../../dashboard.zip .
- name: Extract changelog entry
run: |
VERSION="${{ needs.check-release.outputs.version }}"
@@ -310,13 +320,13 @@ jobs:
if gh release view "$TAG" &>/dev/null; then
echo "Release $TAG already exists, uploading assets..."
gh release upload "$TAG" bin/agent-browser-* dashboard.zip --clobber
gh release upload "$TAG" bin/agent-browser-* --clobber
else
echo "Creating release $TAG..."
gh release create "$TAG" \
--title "$TAG" \
--notes-file /tmp/release-notes.md \
bin/agent-browser-* dashboard.zip
bin/agent-browser-*
fi
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+3
View File
@@ -61,6 +61,9 @@ docs/package-lock.json
# pnpm
.pnpm-store/
# TypeScript
*.tsbuildinfo
# next
.next/
out/
+8
View File
@@ -0,0 +1,8 @@
if [ "${SKIP_CLAWHUB_SYNC:-0}" = "1" ]; then
echo "Skipping ClawHub sync (SKIP_CLAWHUB_SYNC=1)"
exit 0
fi
pnpm run clawhub:sync || {
echo "ClawHub sync failed. Push continues. Run 'pnpm run clawhub:sync' manually after fixing login/network."
}
+1
View File
@@ -0,0 +1 @@
24
+9 -10
View File
@@ -19,7 +19,7 @@ When adding or changing user-facing features (new flags, commands, behaviors, en
1. `cli/src/output.rs``--help` output (flags list, examples, environment variables)
2. `README.md` — Options table, relevant feature sections, examples
3. `skills/agent-browser/SKILL.md` — so AI agents know about the feature
3. `skill-data/core/SKILL.md` (and its `references/`) — so AI agents know about the feature when they load the core skill. Edit `skill-data/core/SKILL.md` for overview/workflow changes; edit `skill-data/core/references/*.md` for detailed reference content. Do **not** put feature content in `skills/agent-browser/SKILL.md` — that file is an intentionally thin discovery stub for `npx skills add` and exists only to redirect agents to `agent-browser skills get core`.
4. `docs/src/app/` — the Next.js docs site (MDX pages)
5. Inline doc comments in the relevant source files
@@ -41,7 +41,7 @@ To prepare a release:
1. Create a branch (e.g. `prepare-v0.24.0`)
2. Bump `version` in `package.json`
3. Run `pnpm version:sync` to update `cli/Cargo.toml`, `cli/Cargo.lock`, and `packages/dashboard/package.json`
4. Write the changelog entry in `CHANGELOG.md` at the top, under a new `## <version>` heading, wrapped in `<!-- release:start -->` and `<!-- release:end -->` markers
4. Write the changelog entry in `CHANGELOG.md` at the top, under a new `## <version>` heading, wrapped in `<!-- release:start -->` and `<!-- release:end -->` markers. Remove the `<!-- release:start -->` and `<!-- release:end -->` markers from the previous release entry so only the new release has markers.
5. Add a matching entry to `docs/src/app/changelog/page.mdx` at the top (below the `# Changelog` heading)
6. Open a PR and merge to `main`
@@ -51,16 +51,12 @@ When the PR merges, CI compares `package.json` version to what's on npm. If it d
Review the git log since the last release and write the entry in `CHANGELOG.md`. Follow the existing format and voice. Group changes under `### New Features`, `### Bug Fixes`, `### Improvements`, etc. Bold the feature/fix name, then describe it concisely. Reference PR numbers in parentheses.
Wrap the release notes (everything between the `## <version>` heading and the previous version) in markers so CI can extract them for the GitHub release:
Wrap the release notes (everything between the `## <version>` heading and the previous version) in markers so CI can extract them for the GitHub release. Only the current release should have markers; remove the `<!-- release:start -->` and `<!-- release:end -->` markers from any previous release entry:
```markdown
## 0.24.0
## 0.24.1
<!-- release:start -->
### New Features
- **Foo command** - Added `foo` command for bar (#1234)
### Bug Fixes
- Fixed **baz** not working when qux is enabled (#1235)
@@ -68,10 +64,13 @@ Wrap the release notes (everything between the `## <version>` heading and the pr
### Contributors
- @ctate
- @somecontributor
<!-- release:end -->
## 0.23.3
## 0.24.0
### New Features
- **Foo command** - Added `foo` command for bar (#1234)
```
Include a `### Contributors` section listing the GitHub usernames (with `@` prefix) of everyone who contributed to the release. Check the git log between the previous tag and HEAD to find them.
-753
View File
@@ -1,753 +0,0 @@
# agent-browser
## 0.24.0
<!-- release:start -->
### New Features
- **AWS Bedrock AgentCore provider** - Added AWS Bedrock AgentCore as a cloud browser provider. Connect with `--provider agentcore` or `AGENT_BROWSER_PROVIDER=agentcore`. Uses lightweight manual SigV4 signing for authentication with support for the full AWS credential provider chain (environment variables, AWS CLI, SSO, IAM roles). Configure with `AGENTCORE_REGION`, `AGENTCORE_PROFILE_ID`, and `AGENTCORE_BROWSER_ID` environment variables. Returns session ID and Live View URL in the launch response (#397)
### Documentation
- Added AgentCore provider page to docs site, README options table, SKILL.md, and dashboard provider icons (#1120)
### Contributors
- @ctate
- @pahud
<!-- release:end -->
## 0.23.4
### Bug Fixes
- Fixed **daemon hang on Linux** caused by a `waitpid(-1)` race condition in the SIGCHLD handler that stole exit statuses from Rust's `Child` handles, leaving the daemon in a broken state. Replaced the global signal handler with targeted crash detection via the existing drain interval (#1098)
## 0.23.3
### Bug Fixes
- Fixed **drag and drop** not working because `mouseMoved` events during the drag omitted the `buttons` bitmask, causing the browser to see `event.buttons === 0` and never fire `dragstart`/`dragover`/`drop` (#1087)
## 0.23.2
### Patch Changes
- 3c942e2: ### New Features
- **Dashboard session creation** - Sessions can now be created directly from the dashboard UI. A new session dialog provides a unified selector grid for local engines (Chrome, Lightpanda) and cloud providers (Browserbase, Browserless, Browser Use, Kernel) with async creation, loading state, and error display (#1092)
- **Dashboard provider icons** - The session sidebar now shows the provider or engine icon for each session, making it easy to identify which backend a session is using (#1092)
### Bug Fixes
- Fixed **Browser Use** provider using an intermediate API call instead of connecting directly via WSS (`wss://connect.browser-use.com`), which caused connection failures (#1092)
- Fixed **Browserbase** provider not sending an explicit JSON body and `Content-Type` header, causing session creation to fail (#1092)
- Fixed **provider navigation** hanging because `wait_for_lifecycle` waited for page load events that remote providers may not emit. Navigation with `--provider` now automatically sets `waitUntil=none` (#1092)
- Fixed **remote CDP connections** timing out by increasing the CDP connect timeout from 10s to 25s for cloud providers (#1092)
- Fixed **zombie daemon processes** not being cleaned up when a provider connection fails during session creation from the dashboard (#1092)
## 0.23.1
### Patch Changes
- fbcab37: ### New Features
- **Auto-dismissal for alert and beforeunload dialogs** - JavaScript `alert()` and `beforeunload` dialogs are now automatically accepted to prevent the agent from blocking indefinitely. `confirm` and `prompt` dialogs still require explicit `dialog accept/dismiss` commands. Disable with `--no-auto-dialog` flag or `AGENT_BROWSER_NO_AUTO_DIALOG` environment variable (#1075)
- **Puppeteer browser cache fallback** - Chrome discovery now searches `~/.cache/puppeteer/chrome/` (or `PUPPETEER_CACHE_DIR`) for Chrome binaries, so users with an existing Puppeteer installation can use agent-browser without a separate install step (#1088)
- **Console output improvements** - `console.log` of objects now shows the actual object preview (e.g. `{userId: "abc", count: 42}`) instead of `"Object"`. JSON output includes a raw `args` array for programmatic access (#1040)
### Bug Fixes
- Fixed **same-document navigation** (e.g. SPA hash routing) hanging forever because `wait_for_lifecycle` waited for a `Page.loadEventFired` that never fires on same-document navigations (#1059)
- Fixed **save_state** only capturing cookies and localStorage for the current origin, silently dropping cross-domain data (e.g. SSO/CAS auth cookies). Now uses `Network.getAllCookies` and collects localStorage from all visited origins (#1064)
- Fixed **externally opened tabs** not appearing in `tab list` when using `--cdp` mode. Tabs opened by the user or another CDP client are now detected and tracked (#1042)
- Fixed **dashboard server** not picking up installed files without a restart. `dashboard install` now takes effect immediately on a running server (#1066)
- Fixed **Windows Chrome extraction** failing because zip path normalization used forward slashes while the extraction code expected backslashes (#1088)
## 0.23.0
### Minor Changes
- 0f0f300: ### New Features
- **Observability dashboard** - Added a local web UI (`dashboard`) that shows live browser viewports, command activity feeds, console output, network requests, storage, and extensions for all sessions. Manage it with `dashboard start`, `dashboard stop`, and `dashboard install`. The dashboard runs as a standalone background process and all sessions stream to it automatically (#1034)
- **Runtime stream management** - Added `stream enable`, `stream disable`, and `stream status` commands to control WebSocket streaming at runtime. Streaming is now always enabled by default; `AGENT_BROWSER_STREAM_PORT` overrides the port instead of toggling the feature (#951)
- **Close all sessions** - Added `close --all` flag to close every active browser session at once
### Bug Fixes
- Fixed **Lightpanda engine** compatibility (#1050)
- Fixed **Windows daemon TCP bind** failing when Hyper-V reserves the port by falling back to an OS-assigned port and writing it to a `.port` file (#1041)
- Fixed **Windows dashboard relay** using Unix socket instead of TCP (#1038)
- Fixed **radio/checkbox elements** being dropped from compact snapshot tree because the `ref=` check required a leading `[` that those elements lack (#1008)
## 0.22.3
### Patch Changes
- eb64ca4: ### Bug Fixes
- **Re-apply download behavior on recording context** - Fixed an issue where downloads were silently dropped in recording contexts because `Browser.setDownloadBehavior` set at launch only applied to the default context. The download behavior is now re-applied when a new recording context is created (#1019)
- **Reap zombie Chrome process and fast-detect crash for auto-restart** - Added a non-blocking process-exit check before attempting CDP connection checks. This prevents a 3-second CDP timeout when Chrome has already crashed or exited, enabling faster detection and auto-restart of the browser (#1023)
- **Route keyboard `type` through text input** - Fixed keyboard `type` subaction to correctly route through the text input handler, and added support for an `insertText` subaction using `Input.insertText` (#1014)
- **Handle `--clear` flag in `console` command** - Fixed the `console` command to accept and process a `clear` parameter, allowing console event history to be cleared (#1015)
## 0.22.2
### Patch Changes
- a098197: ### New Features
- **Dialog status command** - Added `dialog status` command to check whether a JavaScript dialog is currently open (#999)
- **Dialog warning field** - Command responses now include a `warning` field when a JavaScript dialog is pending, indicating the dialog type and message (#999)
### Improvements
- **Standard proxy environment variables** - The proxy setting now automatically falls back to standard environment variables (`HTTP_PROXY`, `HTTPS_PROXY`, `ALL_PROXY`, and their lowercase variants), with `NO_PROXY`/`no_proxy` respected for bypass rules (#1000)
- **Font packages for `--with-deps`** - Installing with `--with-deps` now includes CJK and emoji font packages on Linux (Debian, RPM, and yum-based distros) to prevent missing glyphs when rendering international content (#1002)
### Bug Fixes
- Fixed `state show` always failing with "Missing 'path' parameter" due to a mismatched JSON field name (`filename``path`) (#994)
- Fixed `console` command returning only `Done` due to a JSON field name mismatch in the response (#986)
- Fixed browser-domain CDP events being dropped during downloads due to a `sessionId` mismatch (#998)
- Fixed proxy authentication by handling credentials via the CDP `Fetch.authRequired` event rather than passing them inline (#1000)
## 0.22.1
### Patch Changes
- 3a3317b: ### Bug Fixes
- Fixed **modifier key chords** (e.g. `Control+a`, `Shift+Enter`, `Control+Shift+a`) not being handled correctly when using `press`. Modifier keys (`Alt`, `Control`/`Ctrl`, `Meta`/`Cmd`, `Shift`) are now parsed and forwarded as CDP modifier bitmasks rather than treated as part of the key name (#980)
- Fixed **query parameters being dropped** from `--cdp` HTTP URLs (e.g. `http://host:9222?mode=Hello`). Query strings are now preserved and forwarded to the remote CDP endpoint (#982)
## 0.22.0
### Minor Changes
- be30bc9: ### New Features
- **Cross-origin iframe support** - Added support for snapshots and interactions within cross-origin iframes via `Target.setAutoAttach` (#949)
- **Network request detail and filtering** - Added `network request <requestId>` command to view full request/response detail, and new filtering options for `network requests` including `--type` (e.g. `xhr,fetch`), `--method` (e.g. `POST`), and `--status` (e.g. `2xx`, `400-499`) (#935)
### Improvements
- **Snapshot usability** - Reduced AI cognitive load by filtering semantic noise from snapshot output; cursor-interactive elements are now included by default, making the `-C` flag unnecessary (#968)
- **Upgrade command** - Improved robustness of installation method detection in the upgrade command (#960)
- **Target tracking** - Enhanced target tracking and page information handling for more reliable browser session management (#969)
### Bug Fixes
- Fixed **viewport dimensions** being reported incorrectly in streaming status messages and screencast (#952)
- Fixed **`find` command** flags such as `--exact` and `--name` leaking into fill values when used with fill actions (#955)
- Fixed **state commands** incorrectly starting the daemon when no `session_name` is provided (#677, #964)
- Fixed **auto-connect** triggering when the daemon is already running, preventing duplicate connections (#971)
- Fixed **Enter key press** not working by adding a text field to `keyDown` events (#972)
- Fixed **download command** to properly handle absolute paths and correctly click target elements (#970)
### Breaking Changes
- The `-C` / `--cursor` flag for `snapshot` is deprecated; cursor-interactive elements are now included by default and the flag has no additional effect (#968)
### Documentation
- Updated `README.md` with new `network requests` filtering options and `network request <requestId>` command usage
- Removed references to the deprecated `-C` / `--cursor` snapshot flag from docs and command reference
## 0.21.4
### Patch Changes
- aed466b: ### Bug Fixes
- **Auth login readiness** - `agent-browser auth login` now navigates with `load`, waits for usable login form selectors, and uses staged username detection (targeted email/username selectors first, then broad text-input fallback). This reduces SPA timing failures, avoids false matches on unrelated text fields, and prevents `networkidle` hangs on pages with continuous background requests.
## 0.21.3
### Patch Changes
- 6daad22: ### Bug Fixes
- **WebSocket keepalive for remote browsers** - Added WebSocket Ping frames and TCP `SO_KEEPALIVE` to prevent CDP connections from being silently dropped by intermediate proxies (reverse proxies, load balancers, service meshes) during idle periods (#936)
- **XPath selector support** - Fixed element resolution to correctly handle the `xpath=` selector prefix (#908)
### Performance
- **Fast-path for identical snapshots** - Short-circuits the Myers diff algorithm when comparing a snapshot to itself, avoiding unnecessary computation in retry and loop workloads where repeated identical snapshots are common (#922)
### Documentation
- Migrated page metadata from MDX files to `layout.tsx` (#904)
- Added search functionality and color improvements to docs (#927)
- Fixed desktop browser list in the iOS comparison table (#926)
- Created a new `providers/` section with dedicated provider pages (#928)
## 0.21.2
### Patch Changes
- 757626f: ### Bug Fixes
- **Deduplicate text content in snapshots** - Fixed an issue where duplicate text content appeared in page snapshots (#909)
- **Native mouse drag state** - Fixed incorrect raw native mouse drag state not being properly tracked across `down`, `move`, and `up` events (#872)
- **Chrome headless launch failures** - Fixed browser launch failures caused by the `--enable-unsafe-swiftshader` flag in Chrome headless mode (#915)
- **Origin-scoped `--headers` persistence** - Restored correct persistence of origin-scoped headers set via `--headers` across navigation commands (#894)
- **Relative URLs in WebSocket domain filter** - Fixed handling of relative URLs in the WebSocket domain filter script (#624)
## 0.21.1
### Patch Changes
- 1e7619d: ### New Features
- **HAR 1.2 network capture** - Added commands to capture and export network traffic in HAR 1.2 format, including accurate request/response timing, headers, body sizes, and resource types sourced from Chrome DevTools Protocol events (#864)
- **Built-in `upgrade` command** - Added `agent-browser upgrade` to self-update the CLI; automatically detects your installation method (npm, Homebrew, or Cargo) and runs the appropriate update command (#898)
### Documentation
- Added `upgrade` command to the README command reference and installation guide
- Added a dedicated **Updating** section to the README with usage instructions for `agent-browser upgrade`
## 0.21.0
### Minor Changes
- c6de80b: ### New Features
- **`batch` command** -- Execute multiple commands from stdin in a single invocation. Accepts a JSON array of string arrays and returns results sequentially. Supports `--bail` to stop on first error and `--json` for structured output (#865)
- **iframe support** -- CLI interactions and snapshots now traverse into iframe content, enabling automation of cross-frame pages (#869)
- **`network har start/stop` command** -- Capture and export network traffic in HAR 1.2 format (#874)
- **WebSocket fallback for CDP discovery** -- When HTTP-based CDP endpoint discovery fails, the CLI now falls back to a WebSocket connection automatically (#873)
### Improvements
- **`--full`/`-f` refactored to command-level flag** -- Moved from a global flag to a per-command flag for clearer scoping (#877)
- **Enhanced Chrome launch** -- Added `--user-data-dir` support and configurable launch timeout for more reliable browser startup (#852)
### Bug Fixes
- Fixed `/json/list` fallback when `/json/version` endpoint is unavailable, improving compatibility with non-standard CDP implementations (#861)
- Fixed daemon liveness detection for PID namespace isolation (e.g. `unshare`). Uses socket connectivity as the sole liveness check instead of `kill(pid, 0)`, which fails when the caller cannot see the daemon's PID (#879)
- Fixed Ubuntu dependency install accidentally removing system packages (#884)
## 0.20.14
### Patch Changes
- c0d4cf6: ### New Features
- **Idle timeout for daemon auto-shutdown** - Added `--idle-timeout` CLI flag (and `AGENT_BROWSER_IDLE_TIMEOUT_MS` environment variable) to automatically shut down the daemon after a period of inactivity. Accepts human-friendly formats such as `10s`, `3m`, `1h`, or raw milliseconds (#856)
- **Cursor-interactive elements in snapshot tree** - Cursor-interactive elements are now embedded directly into the snapshot tree for richer context (#855)
### Bug Fixes
- Fixed **remote host support** in CDP discovery, enabling connection to browsers running on non-local hosts (#854)
- Fixed **CDP flag propagation** to the daemon process, ensuring reliable CDP reconnection across sessions (#857)
- Fixed **Windows auto-connect profiling** to correctly handle browser connection on Windows (#835, #840)
- Fixed **Windows transient error detection** by recognising Windows-specific socket error codes (`os error 10061` connection refused, `os error 10054` connection reset) during daemon reconnection attempts
## 0.20.13
### Patch Changes
- eda956b: ### Bug Fixes
- **Network idle detection for cached pages** - Fixed an issue where `poll_network_idle` could return immediately when no network events were observed (e.g. pages served from cache). The idle timer is now only satisfied after a consistent **500 ms idle period** has elapsed, preventing false-positive idle detection. The core polling logic has also been extracted into a standalone `poll_network_idle` function to improve testability (#847)
## 0.20.12
### Patch Changes
- 5fa2396: ### Bug Fixes
- Fixed **`snapshot -C`** and **`screenshot --annotate`** hanging when connected over WSS (WebSocket Secure) due to sequential CDP round-trips per interactive element (#842)
### Performance
- **`snapshot -C` (cursor-interactive mode)** now batches CDP calls instead of issuing N×2 sequential round-trips per cursor-interactive element, preventing timeouts on high-latency WSS connections (#842)
- **`screenshot --annotate`** now batches element queries, reducing completion time from potentially 2040s (e.g. 50+ buttons over WSS) to within expected bounds (#842)
## 0.20.11
### Patch Changes
- 4b5fc78: ### Bug Fixes
- **Material Design checkbox/radio parity** - Restored Playwright-parity behavior for `check`/`uncheck` actions on Material Design controls. These components hide the native `<input>` off-screen and use overlay elements that intercept coordinate-based clicks; the actions now detect this pattern and fall back to a JS `.click()` to correctly toggle state. Also improves `ischecked` to handle nested hidden inputs and ARIA-only checkboxes (#837)
- **Punctuation handling in `type` command** - Fixed incorrect virtual key (VK) codes being used for punctuation characters (e.g. `.`, `@`) in the `type` action, which previously caused those characters to be dropped or mistyped (#836)
## 0.20.10
### Patch Changes
- a3d9662: ### Bug Fixes
- **Restored WebSocket streaming** - Fixed broken WebSocket streaming in the native daemon by keeping the **StreamServer** instance alive so the broadcast channel remains open, and ensuring CDP session IDs and connection status are correctly propagated to stream clients (#826)
- **Filtered internal Chrome targets** - Fixed auto-connect discovery incorrectly attempting to attach to Chrome-internal pages (e.g. `chrome://`, `chrome-extension://`, `devtools://` URLs), which could cause unexpected connection failures (#827)
## 0.20.9
### Patch Changes
- 51d9ab4: ### Bug Fixes
- **Appium v3 iOS capabilities** - Added `appium:` vendor prefix to iOS capabilities (e.g., `appium:automationName`, `appium:deviceName`, `appium:platformVersion`) to comply with the Appium v3 WebDriver protocol requirements (#810)
- **Snapshot `--selector` scoping** - Fixed `snapshot --selector` so that the output is properly scoped to the matched element's subtree rather than returning the full accessibility tree. The selector now resolves the target DOM node's backend IDs and filters the accessibility tree to only include nodes within that subtree (#825)
## 0.20.8
### Patch Changes
- daf7263: ### Bug Fixes
- Fixed **video duration** being reported incorrectly when using real-time ffmpeg encoding for screen recording (#812)
- Removed obsolete **`BrowserManager` TypeScript API** references that no longer reflect the current CLI-based usage model (#821)
### Documentation
- Updated README to replace outdated **`BrowserManager` programmatic API** examples with the current CLI-based approach using `execSync` and `agent-browser` commands (#821)
- Removed the **Programmatic API** section covering `BrowserManager` screencast and input injection methods, which are no longer part of the public API (#821)
## 0.20.7
### Patch Changes
- 25a1526: ### New Features
- **Brave Browser support** - Added auto-discovery of Brave Browser for CDP connections on macOS, Linux, and Windows. The agent will now automatically detect and connect to Brave alongside Chrome, Chromium, and Canary installations (#817)
### Improvements
- **Postinstall message** - The post-install message now detects existing Chrome installations on the system. If a compatible browser is found, it confirms the path and notes it will be used automatically instead of prompting an install. If no browser is detected, the warning is clearer and mentions that installation can be skipped when using `--cdp`, `--provider`, `--engine`, or `--executable-path` (#815)
## 0.20.6
### Patch Changes
- fa91c22: ### Bug Fixes
- **Stale accessibility tree reference fallback** - Fixed an issue where interacting with an element whose **`backend_node_id`** had become stale (e.g. after the DOM was replaced) would fail with a `Could not compute box model` CDP error. Element resolution now re-queries the accessibility tree using role/name lookup to obtain a fresh node ID before retrying the operation (#806)
## 0.20.5
### Patch Changes
- fc091d2: ### Bug Fixes
- **Daemon panic on broken stderr pipe** - Replaced all `eprintln!` calls with `writeln!(std::io::stderr(), ...)` wrapped in `let _ =` to silently discard write errors, preventing the daemon from panicking when the parent process drops the stderr pipe during Chrome launch (#802)
## 0.20.4
### Patch Changes
- e2ebde2: ### Bug Fixes
- **Broadcast channel lag handling** - Fixed an issue where **broadcast channel lag** errors were incorrectly treated as stream closure, causing premature termination of event listeners in reload, response body, download, and navigation wait operations. Lagged messages are now skipped and the loop continues instead of breaking (#797)
### Improvements
- Removed unused **pnpm setup** steps from the `global-install` CI job, simplifying the workflow configuration (#798)
## 0.20.3
### Patch Changes
- e365909: ### Bug Fixes
- **Chrome launch retry** - Chrome will now retry launching up to 3 times with a 500ms delay between attempts, improving resilience against transient startup failures (#791)
- **Remote CDP snapshot hang** - Resolved an issue where snapshots would hang indefinitely over remote CDP (WSS) connections by removing WebSocket message and frame size limits to accommodate large responses (e.g. `Accessibility.getFullAXTree`), accepting binary frames from remote proxies such as Browserless, and immediately clearing pending commands when the connection closes rather than waiting for the 30-second timeout (#792)
## 0.20.2
### Patch Changes
- 944fa01: ### New Features
- **Linux musl (Alpine) builds** - Added pre-built binaries for **linux-musl** targeting both **x64** and **arm64** architectures, enabling native support for Alpine Linux and other musl-based distributions without requiring glibc (#784)
### Improvements
- **Consecutive `--auto-connect` commands** - Added support for issuing multiple consecutive `--auto-connect` commands without requiring a full browser relaunch; external connections are now correctly identified and reused (#786)
- **External browser disconnect behavior** - When using `--auto-connect` or `--cdp`, closing the agent session now disconnects cleanly without shutting down the user's browser process
### Bug Fixes
- **Restored `refs` dict in `--json` snapshot output** - The `refs` map containing role and name metadata for referenced elements is now correctly included in JSON snapshot responses (#787)
- Fixed e2e test assertions for `diff_snapshot` and `domain_filter` to correctly reflect expected behavior (#783)
- Fixed Chrome temp-dir cleanup test failing on Windows (#766)
## 0.20.1
### Patch Changes
- bd05917: ### Bug Fixes
- Fixed **AX tree deserialization** to accept integer `nodeId` and `childIds` values for compatibility with Lightpanda, which sends numeric IDs where Chrome sends strings (#775)
- Fixed **misleading SIGPIPE comment** to accurately describe the default Rust SIGPIPE behavior and why it is reset to `SIG_DFL` (#776)
- Fixed **WebM recording output** to use the VP9 codec (`libvpx-vp9`) instead of H.264, producing valid WebM files; also adds a padding filter to ensure even frame dimensions (#779)
## 0.20.0
### Minor Changes
- 235fa88: ### Full Native Rust
- **100% native Rust** -- Removed the entire Node.js/Playwright daemon. The Rust native daemon is now the only implementation. No Node.js runtime or Playwright dependency required. (#754)
- **99x smaller install** -- Install size reduced from 710 MB to 7 MB by eliminating the Node.js dependency tree.
- **18x less memory** -- Daemon memory usage reduced from 143 MB to 8 MB.
- **1.6x faster cold start** -- Cold start time reduced from 1002ms to 617ms.
- **Benchmarks** -- Added benchmark suite comparing native vs Node.js daemon performance.
- **Chromium installer hardened** -- Fixed zip path traversal vulnerability in Chrome for Testing installer.
### Bug Fixes
- Fixed `--headed false` flag not being respected in CLI (#757)
- Fixed "not found" error pattern in `to_ai_friendly_error` incorrectly catching non-element errors (#759)
- Fixed storage local key lookup parsing and text output (#761)
- Fixed Lightpanda engine launch with release binaries (#760)
- Hardened Lightpanda startup timeouts (#762)
## 0.19.0
### Minor Changes
- 56bb92b: ### New Features
- **Browserless.io provider** -- Added browserless.io as a browser provider, supported in both Node.js and native daemon paths. Connect to remote Browserless instances with `--provider browserless` or `AGENT_BROWSER_PROVIDER=browserless`. Configurable via `BROWSERLESS_API_KEY`, `BROWSERLESS_API_URL`, and `BROWSERLESS_BROWSER_TYPE` environment variables. (#502, #746)
- **`clipboard` command** -- Read from and write to the browser clipboard. Supports `read`, `write <text>`, `copy` (simulates Ctrl+C), and `paste` (simulates Ctrl+V) operations. (#749)
- **Screenshot output configuration** -- New global flags `--screenshot-dir`, `--screenshot-quality`, `--screenshot-format` and corresponding `AGENT_BROWSER_SCREENSHOT_DIR`, `AGENT_BROWSER_SCREENSHOT_QUALITY`, `AGENT_BROWSER_SCREENSHOT_FORMAT` environment variables for persistent screenshot settings. (#749)
### Bug Fixes
- Fixed `wait --text` not working in native daemon path (#749)
- Fixed `BrowserManager.navigate()` and package entry point (#748)
- Fixed extensions not being loaded from `config.json` (#750)
- Fixed scroll on page load (#747)
- Fixed HTML retrieval by using `browser.getLocator()` for selector operations (#745)
## 0.18.0
### Minor Changes
- 942b8cd: ### New Features
- **`inspect` command** - Opens Chrome DevTools for the active page by launching a local proxy server that forwards the DevTools frontend to the browser's CDP WebSocket. Commands continue to work while DevTools is open. Implemented in both Node.js and native paths. (#736)
- **`get cdp-url` subcommand** - Retrieve the Chrome DevTools Protocol WebSocket URL for the active page, useful for external debugging tools. (#736)
- **Native screenshot annotate** - The `--annotate` flag for screenshots now works in the native Rust daemon, bringing parity with the Node.js path. (#706)
### Improvements
- **KERNEL_API_KEY now optional** - External credential injection no longer requires `KERNEL_API_KEY` to be set, making it easier to use Kernel with pre-configured environments. (#687)
- **Browserbase simplified** - Removed the `BROWSERBASE_PROJECT_ID` requirement, reducing setup friction for Browserbase users. (#625)
### Bug Fixes
- Fixed Browserbase API using incorrect endpoint to release sessions (#707)
- Fixed CDP connect paths using hardcoded 10s timeout instead of `getDefaultTimeout()` (#704)
- Fixed lone Unicode surrogates causing errors by sanitizing with `toWellFormed()` (#720)
- Fixed CDP connection failure on IPv6-first systems (#717)
- Fixed recordings not inheriting the current viewport settings (#718)
## 0.17.1
### Patch Changes
- 94cd888: Added support for device scale factor (retina display) in the viewport command via an optional scale parameter. Also added webview target type support for better Electron application compatibility, and the pages list now includes target type information.
## 0.17.0
### Minor Changes
- 94521e7: ### New Features
- **Lightpanda browser engine support** - Added `--engine <name>` flag to select the browser engine (`chrome` by default, or `lightpanda`), implying `--native` mode. Configurable via `AGENT_BROWSER_ENGINE` environment variable (#646)
- **Dialog dismiss command** - Added support for `dismiss` subcommand in dialog command parsing (#605)
### Improvements
- **Daemon startup error reporting** - Daemon startup errors are now surfaced directly instead of showing an opaque timeout message (#614)
- **CDP port discovery** - Replaced broken hand-rolled HTTP client with `reqwest` for more reliable CDP port discovery (#619)
- **Chrome extensions** - Extensions now load correctly by forcing headed mode when extensions are present (#652)
- **Google Translate bar suppression** - Suppressed the Google Translate bar in native headless mode to avoid interference (#649)
- **Auth cookie persistence** - Auth cookies are now persisted on browser close in native mode (#650)
### Bug Fixes
- Fixed native auth login failing due to incompatible encryption format (#648)
### Documentation
- Improved snapshot usage guidance and added reproducibility check (#630)
- Added `--engine` flag to the README options table
### Performance
- Added benchmarks to the CLI codebase (#637)
## 0.16.3
### Patch Changes
- 7d2c895: Fixed an issue where the --native flag was being passed to child processes even when not explicitly specified on the command line. The flag is now only forwarded when the user explicitly provides it, consistent with how other CLI flags like --allow-file-access and --download-path are handled.
## 0.16.2
### Patch Changes
- 01ac557: Added AGENT_BROWSER_HEADED environment variable support for running the browser in headed mode, and improved temporary profile cleanup when launching Chrome directly. Also includes documentation clarification that browser extensions work in both headed and headless modes.
## 0.16.1
### Patch Changes
- c4180c8: Improved Chrome launch reliability by automatically detecting containerized environments (Docker, Podman, Kubernetes) and enabling --no-sandbox when needed. Added support for discovering Playwright-installed Chromium browsers and enhanced error messages with helpful diagnostics when Chrome fails to launch.
## 0.16.0
### Minor Changes
- 05018b3: Added experimental native Rust daemon (`--native` flag, `AGENT_BROWSER_NATIVE=1` env, or `"native": true` in config). The native daemon communicates with Chrome directly via CDP, eliminating Node.js and Playwright dependencies. Supports 150+ commands with full parity to the default Node.js daemon. Includes WebDriver backend for Safari/iOS, CDP protocol codegen, request tracking, frame context management, and comprehensive e2e and parity tests.
## 0.15.3
### Patch Changes
- 62241b5: Fixed Windows compatibility issues including proper handling of extended-length path prefixes from canonicalize(), prevention of MSYS/Git Bash path translation that could mangle arguments, and improved daemon startup reliability. Also added ARM64 Windows support in postinstall shims and expanded CI testing with a full daemon lifecycle test on Windows.
## 0.15.2
### Patch Changes
- 6aea316: Documentation site improvements and internal tooling updates including enhanced code blocks, mobile navigation, and docs chat components. CLI connection and output handling refinements. Skill creator reference documentation and scripts have been reorganized.
## 0.15.1
### Patch Changes
- 7bd8ce9: Added support for chrome:// and chrome-extension:// URLs in navigation and recording commands. These special browser URLs are now preserved as-is instead of having https:// incorrectly prepended.
## 0.15.0
### Minor Changes
- 2e38882: - Added security hardening: authentication vault, content boundary markers, domain allowlist, action policy, action confirmation, and output length limits.
- Added `--download-path` flag (and `AGENT_BROWSER_DOWNLOAD_PATH` env / `downloadPath` config key) to set a default download directory.
- Added `--selector` flag to `scroll` command for scrolling within specific container elements.
## 0.14.0
### Minor Changes
- b7665e5: - Added `keyboard` command for raw keyboard input -- type with real keystrokes, insert text, and press shortcuts at the currently focused element without needing a selector.
- Added `--color-scheme` flag and `AGENT_BROWSER_COLOR_SCHEME` env var for persistent dark/light mode preference across browser sessions.
- Fixed IPC EAGAIN errors (os error 35/11) by adding backpressure-aware socket writes, command serialization, and lowering the default Playwright timeout to 25s (configurable via `AGENT_BROWSER_DEFAULT_TIMEOUT`).
- Fixed remote debugging (CDP) reconnection.
- Fixed state load failing when no browser is running.
- Fixed `--annotate` flag warning appearing when not explicitly passed via CLI.
## 0.13.0
### Minor Changes
- ebd8717: Added new diff commands for comparing snapshots, screenshots, and URLs between page states. You can now run visual pixel diffs against baseline images, compare accessibility tree snapshots with customizable depth and selectors, and diff two URLs side-by-side with optional screenshot comparison.
## 0.12.0
### Minor Changes
- 69ffad0: Add annotated screenshots with the new --annotate flag, which overlays numbered labels on interactive elements and prints a legend mapping each label to its element ref. This enables multimodal AI models to reason about visual layout while using the same @eN refs for subsequent interactions. The flag can also be set via the AGENT_BROWSER_ANNOTATE environment variable.
## 0.11.1
### Patch Changes
- c6fc7df: Added documentation for command chaining with && across README, CLI help output, docs, and skill files, explaining how to efficiently chain multiple agent-browser commands in a single shell invocation since the browser persists via a background daemon.
## 0.11.0
### Minor Changes
- 5dc40b4: Added configuration file support with automatic loading from user and project directories, new profiler commands for Chrome DevTools profiling, computed styles getter, browser extension loading, storage state management, and iOS device emulation. Expanded click command with new-tab option, improved find command with additional actions and filtering options, and enhanced CDP connection to accept WebSocket URLs. Documentation has been significantly expanded with new sections for configuration, profiling, and proxy support.
## 0.10.0
### Minor Changes
- 1112a16: Added session persistence with automatic save/restore of cookies and localStorage across browser restarts using --session-name flag, with optional AES-256-GCM encryption for saved state data. New state management commands allow listing, showing, renaming, clearing, and cleaning up old session files. Also added --new-tab option for click commands to open links in new tabs.
## 0.9.4
### Patch Changes
- 323b6cd: Fix all Clippy lint warnings in the Rust CLI: remove redundant import, use `.first()` instead of `.get(0)`, use `.copied()` instead of `.map(|s| *s)`, use `.contains()` instead of `.iter().any()`, use `then_some` instead of lazy `then`, and simplify redundant match guards.
## 0.9.3
### Patch Changes
- d03e238: Added support for custom executable path in CLI browser launch options. Documentation site received UI improvements including a new chat component with sheet-based interface and updated dependencies.
## 0.9.2
### Patch Changes
- 76d23db: Documentation site migrated to MDX for improved content authoring, added AI-powered docs chat feature, and updated README with Homebrew installation instructions for macOS users.
## 0.9.1
### Patch Changes
- ae34945: Added --allow-file-access flag to enable opening and interacting with local file:// URLs (PDFs, HTML files) by passing Chromium flags that allow JavaScript access to local files. Added -C/--cursor flag for snapshots to include cursor-interactive elements like divs with onclick handlers or cursor:pointer styles, which is useful for modern web apps using custom clickable elements.
## 0.9.0
### Minor Changes
- 9d021bd: Add iOS Simulator and real device support for mobile Safari testing via Appium. New CLI commands include `device list` to show available simulators, `tap` and `swipe` for touch interactions, and the `--device` flag to specify which iOS device to use. Configure with `-p ios` provider flag or `AGENT_BROWSER_PROVIDER=ios` environment variable.
## 0.8.10
### Patch Changes
- 17dba8f: Add --stdin flag for eval command to read JavaScript from stdin, enabling heredoc usage for multiline scripts
- daeede4: Add --stdin flag for the eval command to read JavaScript from stdin, enabling heredoc usage for multiline scripts. Also fix binary permission issues on macOS/Linux when postinstall scripts don't run (e.g., with bun).
## 0.8.9
### Patch Changes
- 0dc36f2: Add --stdin flag for eval command to read JavaScript from stdin, enabling heredoc usage for multiline scripts
## 0.8.8
### Patch Changes
- 2771588: Added base64 encoding support for the eval command with -b/--base64 flag to avoid shell escaping issues when executing JavaScript. Updated documentation with AI agent setup instructions and reorganized the docs structure by consolidating agent mode content into the installation page.
## 0.8.7
### Patch Changes
- d24f753: Fixed browser launch options not being passed correctly when using persistent profiles, ensuring args, userAgent, proxy, and ignoreHTTPSErrors settings now work properly. Added pre-flight checks for socket path length limits and directory write permissions to provide clearer error messages when daemon startup fails. Improved error handling to properly exit with failure status when browser launch fails.
## 0.8.6
### Patch Changes
- d75350a: Improved daemon connection reliability by adding automatic retry logic for transient errors like connection resets, broken pipes, and temporary resource unavailability. The CLI now cleans up stale socket and PID files before starting a new daemon, and includes better detection of daemon responsiveness to handle race conditions during shutdown.
## 0.8.5
### Patch Changes
- cb2f8c3: Fixed version synchronization to automatically update Cargo.lock alongside Cargo.toml during releases, and made the CLI binary executable. This ensures the Rust CLI version stays in sync with the npm package version.
## 0.8.4
### Patch Changes
- 759302e: Fixed "Daemon not found" error when running through AI agents (e.g., Claude Code) by resolving symlinks in the executable path. Previously, npm global bin symlinks weren't being resolved correctly, causing intermittent daemon discovery failures.
## 0.8.3
### Patch Changes
- 4116a8a: Replaced shell-based CLI wrappers with a cross-platform Node.js wrapper to enable npx support on Windows. Added postinstall logic to patch npm's bin entry on global installs, allowing the native binary to be invoked directly with zero overhead. Added CI tests to verify global installation works correctly across all platforms.
## 0.8.2
### Patch Changes
- 7e6336f: Fixed the Windows CMD wrapper to use the native binary directly instead of routing through Node.js, improving startup performance and reliability. Added retry logic to the CI install command to handle transient failures during browser installation.
## 0.8.1
### Patch Changes
- 8eec634: Improved release workflow to validate binary file sizes and ensure binaries are executable after npm install. Updated documentation site with a new mobile navigation system and added v0.8.0 changelog entries. Reformatted CHANGELOG.md for better readability.
## v0.8.0
### New Features
- **Kernel cloud browser provider** - Connect to Kernel (https://kernel.sh) for remote browser infrastructure via `-p kernel` flag or `AGENT_BROWSER_PROVIDER=kernel`. Supports stealth mode, persistent profiles, and automatic profile find-or-create.
- **Ignore HTTPS certificate errors** - New `--ignore-https-errors` flag for working with self-signed certificates and development environments
- **Enhanced cookie management** - Extended `cookies set` command with `--url`, `--domain`, `--path`, `--httpOnly`, `--secure`, `--sameSite`, and `--expires` flags for setting cookies before page load
### Bug Fixes
- Fixed tab list command not recognizing new pages opened via clicks or `target="_blank"` links (#275)
- Fixed `check` command hanging indefinitely (#272)
- Fixed `set device` not applying deviceScaleFactor - HiDPI screenshots now work correctly (#270)
- Fixed state load and profile persistence not working in v0.7.6 (#268)
- Screenshots now save to temp directory when no path is provided (#247)
### Security
- Daemon and stream server now reject cross-origin connections (#274)
## 0.7.6
### Patch Changes
- a4d0c26: Allow null values for the screenshot selector field. Previously, passing a null selector would fail validation, but now it is properly handled as an optional value.
## 0.7.5
### Patch Changes
- 8c2a6ec: Fix GitHub release workflow to handle existing releases. If a release already exists, binaries are uploaded to it instead of failing.
## 0.7.4
### Patch Changes
- 957b5e5: Fix binary permissions on install. npm doesn't preserve execute bits, so postinstall now ensures the native binary is executable.
## 0.7.3
### Patch Changes
- 161d8f5: Fix native binary distribution in npm package. Native binaries for all platforms (Linux x64/arm64, macOS x64/arm64, Windows x64) are now correctly included when publishing.
## 0.7.2
### Patch Changes
- 6afede2: Fix native binary distribution in npm package
Native binaries for all platforms (Linux x64/arm64, macOS x64/arm64, Windows x64) are now included in the npm package. Previously, the release workflow published to npm before building binaries, causing "No binary found" errors on installation.
## 0.7.1
### Patch Changes
- Fix native binary distribution in npm package. Native binaries for all platforms (Linux x64/arm64, macOS x64/arm64, Windows x64) are now included in the npm package. Previously, the release workflow published to npm before building binaries, causing "No binary found" errors on installation.
## 0.7.0
### Minor Changes
- 316e649: ## New Features
- **Cloud browser providers** - Connect to Browserbase or Browser Use for remote browser infrastructure via `-p` flag or `AGENT_BROWSER_PROVIDER` env var
- **Persistent browser profiles** - Store cookies, localStorage, and login sessions across browser restarts with `--profile`
- **Remote CDP WebSocket URLs** - Connect to remote browser services via WebSocket URL (e.g., `--cdp "wss://..."`)
- **Download commands** - New `download` command and `wait --download` for file downloads with ref support
- **Browser launch configuration** - New `--args`, `--user-agent`, and `--proxy-bypass` flags for fine-grained browser control
- **Enhanced skills** - Hierarchical structure with references and templates for Claude Code
## Bug Fixes
- Screenshot command now supports refs and has improved error messages
- WebSocket URLs work in `connect` command
- Fixed socket file location (uses `~/.agent-browser` instead of TMPDIR)
- Windows binary path fix (.exe extension)
- State load and path-based actions now show correct output messages
## Documentation
- Added Claude Code marketplace plugin installation instructions
- Updated skill documentation with references and templates
- Improved error documentation
+62 -1317
View File
File diff suppressed because it is too large Load Diff
-4
View File
@@ -1,4 +0,0 @@
# Vercel Sandbox credentials
SANDBOX_VERCEL_TOKEN=
SANDBOX_VERCEL_TEAM_ID=
SANDBOX_VERCEL_PROJECT_ID=
-2
View File
@@ -1,2 +0,0 @@
node_modules/
results.json
-76
View File
@@ -1,76 +0,0 @@
# agent-browser Daemon Benchmarks
Compares command latency and system metrics between the **Node.js daemon** (published npm version) and the **Rust native daemon** (built from source), running inside a [Vercel Sandbox](https://vercel.com/docs/sandbox) microVM.
## What it measures
**Command latency** -- per-scenario timing with warmup, multiple iterations, and stddev:
- `navigate` -- page load round-trip
- `snapshot` -- accessibility tree generation
- `screenshot` -- viewport capture
- `evaluate` -- JavaScript execution
- `click` -- element interaction
- `fill` -- form input
- `agent-loop` -- snapshot/click/snapshot cycle (typical AI agent pattern)
- `full-workflow` -- realistic 7-command sequence
**System metrics** -- collected while the daemon is running:
- Cold start time (daemon spawn + browser launch)
- Binary size and total distribution size (including browser download)
- Daemon RSS and peak RSS (separated from browser process memory)
- Browser RSS (Chrome processes, same for both daemons)
- Daemon CPU time
- Process counts
## Prerequisites
- Node.js 18+
- pnpm
- Vercel Sandbox credentials (token, team ID, project ID)
## Setup
```bash
cd benchmarks
pnpm install
cp .env.example .env
```
Fill in your Vercel Sandbox credentials in `.env`:
```
SANDBOX_VERCEL_TOKEN=your_token
SANDBOX_VERCEL_TEAM_ID=your_team_id
SANDBOX_VERCEL_PROJECT_ID=your_project_id
```
## Usage
```bash
pnpm bench # 10 iterations, 1 warmup, 8 vCPUs
pnpm bench -- --iterations 20 # more iterations for tighter stats
pnpm bench -- --warmup 2 # extra warmup iterations
pnpm bench -- --json # write results.json
pnpm bench -- --branch main # build native from a different branch
pnpm bench -- --vcpus 16 # more vCPUs (faster Rust build)
```
## How it works
1. Creates a Vercel Sandbox (Amazon Linux, configurable vCPUs)
2. Installs Chromium system dependencies
3. **Phase 1 -- Node.js daemon**: installs `agent-browser` from npm (last version with the Node daemon), runs all scenarios, collects metrics
4. **Phase 2 -- Rust native daemon**: installs Rust toolchain, clones the repo, runs `cargo build --release`, replaces the binary, runs the same scenarios, collects metrics
5. Prints comparison tables and optionally writes `results.json`
## Interpreting results
**Command latency** is dominated by Chrome (CDP round-trips), not the daemon. Both daemons are thin relays between the CLI and Chrome, so per-command speedups are typically small. The stddev column helps distinguish real differences from noise.
**Where the native daemon wins** is in cold start (no Node.js runtime to boot), daemon memory (single Rust binary vs V8 heap), and distribution size (no Playwright dependency).
The **daemon RSS** metric isolates the daemon process memory from Chrome. This is the apples-to-apples comparison -- both daemons talk to the same Chrome, but Node.js adds ~140 MB of V8 overhead while the Rust daemon uses ~7 MB.
**Distribution size** includes the daemon plus its browser download. The Node version includes the npm package + Playwright's bundled Chromium. The Rust version is just the binary + Chrome for Testing.
-900
View File
@@ -1,900 +0,0 @@
/**
* Node.js Daemon vs Rust Native Daemon benchmark.
*
* Compares the last published npm version (Node.js daemon) against the
* Rust-only build from a given branch, running real agent-browser commands
* inside a Vercel Sandbox.
*
* Captures:
* - Command latency (per-scenario, with warmup + measured iterations + stddev)
* - Cold start time (first launch to daemon ready)
* - Daemon memory (RSS, peak RSS) separated from browser memory
* - Daemon CPU time
* - Process tree (daemon + browser children)
* - Binary and distribution size on disk
*
* Usage:
* pnpm bench # default: 10 iterations, 1 warmup
* pnpm bench -- --iterations 20 # override iterations
* pnpm bench -- --warmup 2 # override warmup count
* pnpm bench -- --json # write results.json
* pnpm bench -- --branch my-branch # override native branch (default: ctate/native-2)
* pnpm bench -- --vcpus 8 # sandbox vCPUs (default: 8, higher = faster Rust build)
*/
import { Sandbox } from "@vercel/sandbox";
import { readFileSync, writeFileSync } from "fs";
import { scenarios, type Scenario } from "./scenarios.js";
// ---------------------------------------------------------------------------
// Env
// ---------------------------------------------------------------------------
function loadEnv() {
try {
const content = readFileSync(".env", "utf-8");
for (const line of content.split("\n")) {
const trimmed = line.trim();
if (!trimmed || trimmed.startsWith("#")) continue;
const eq = trimmed.indexOf("=");
if (eq === -1) continue;
const key = trimmed.slice(0, eq);
let val = trimmed.slice(eq + 1);
if (
(val.startsWith('"') && val.endsWith('"')) ||
(val.startsWith("'") && val.endsWith("'"))
) {
val = val.slice(1, -1);
}
process.env[key] = val;
}
} catch {}
}
loadEnv();
const credentials = {
token: process.env.SANDBOX_VERCEL_TOKEN!,
teamId: process.env.SANDBOX_VERCEL_TEAM_ID!,
projectId: process.env.SANDBOX_VERCEL_PROJECT_ID!,
};
if (!credentials.token || !credentials.teamId || !credentials.projectId) {
console.error(
"Missing credentials. Set SANDBOX_VERCEL_TOKEN, SANDBOX_VERCEL_TEAM_ID, SANDBOX_VERCEL_PROJECT_ID in .env",
);
process.exit(1);
}
// ---------------------------------------------------------------------------
// CLI args
// ---------------------------------------------------------------------------
function parseArgs() {
const args = process.argv.slice(2);
let iterations = 10;
let warmup = 1;
let json = false;
let branch = "ctate/native-2";
let vcpus = 8;
for (let i = 0; i < args.length; i++) {
if (args[i] === "--iterations" && args[i + 1]) {
iterations = parseInt(args[++i], 10);
} else if (args[i] === "--warmup" && args[i + 1]) {
warmup = parseInt(args[++i], 10);
} else if (args[i] === "--json") {
json = true;
} else if (args[i] === "--branch" && args[i + 1]) {
branch = args[++i];
} else if (args[i] === "--vcpus" && args[i + 1]) {
vcpus = parseInt(args[++i], 10);
}
}
return { iterations, warmup, json, branch, vcpus };
}
const config = parseArgs();
// ---------------------------------------------------------------------------
// Constants
// ---------------------------------------------------------------------------
const TIMEOUT_MS = 30 * 60 * 1000;
const REPO_URL = "https://github.com/vercel-labs/agent-browser.git";
const CHROMIUM_SYSTEM_DEPS = [
"nss",
"nspr",
"libxkbcommon",
"atk",
"at-spi2-atk",
"at-spi2-core",
"libXcomposite",
"libXdamage",
"libXrandr",
"libXfixes",
"libXcursor",
"libXi",
"libXtst",
"libXScrnSaver",
"libXext",
"mesa-libgbm",
"libdrm",
"mesa-libGL",
"mesa-libEGL",
"cups-libs",
"alsa-lib",
"pango",
"cairo",
"gtk3",
"dbus-libs",
];
// ---------------------------------------------------------------------------
// Sandbox helpers
// ---------------------------------------------------------------------------
type SandboxInstance = InstanceType<typeof Sandbox>;
async function run(
sandbox: SandboxInstance,
cmd: string,
args: string[],
): Promise<string> {
const result = await sandbox.runCommand(cmd, args);
const stdout = await result.stdout();
const stderr = await result.stderr();
if (result.exitCode !== 0) {
throw new Error(
`Command failed (exit ${result.exitCode}): ${cmd} ${args.join(" ")}\n${stderr || stdout}`,
);
}
return stdout;
}
async function shell(sandbox: SandboxInstance, script: string): Promise<string> {
return run(sandbox, "sh", ["-c", script]);
}
async function shellSafe(sandbox: SandboxInstance, script: string): Promise<string> {
const result = await sandbox.runCommand("sh", ["-c", script]);
return (await result.stdout()).trim();
}
// ---------------------------------------------------------------------------
// Stats
// ---------------------------------------------------------------------------
interface Stats {
avgMs: number;
stddevMs: number;
minMs: number;
maxMs: number;
p50Ms: number;
samples: number[];
}
function computeStats(samples: number[]): Stats {
const sorted = [...samples].sort((a, b) => a - b);
const sum = sorted.reduce((a, b) => a + b, 0);
const avg = sum / sorted.length;
const variance =
sorted.reduce((acc, v) => acc + (v - avg) ** 2, 0) / sorted.length;
return {
avgMs: Math.round(avg),
stddevMs: Math.round(Math.sqrt(variance)),
minMs: sorted[0],
maxMs: sorted[sorted.length - 1],
p50Ms: sorted[Math.floor(sorted.length / 2)],
samples: sorted,
};
}
// ---------------------------------------------------------------------------
// Metrics collection
// ---------------------------------------------------------------------------
interface ProcessMetrics {
pid: number;
rssKb: number;
vszKb: number;
cpuPercent: number;
memPercent: number;
cpuTimeSec: number;
command: string;
}
interface DaemonMetrics {
coldStartMs: number;
binarySizeBytes: number;
distributionSizeBytes: number;
daemonProcesses: ProcessMetrics[];
browserProcesses: ProcessMetrics[];
daemonRssKb: number;
browserRssKb: number;
daemonPeakRssKb: number;
daemonCpuTimeSec: number;
totalCpuTimeSec: number;
}
async function findDaemonPids(
sandbox: SandboxInstance,
_session: string,
): Promise<number[]> {
// The daemon process name is "agent-browser" but session/daemon flags are
// env vars, not command-line args, so we can't grep them from `ps`.
// Instead, find all agent-browser processes that look like long-running daemons
// (not short-lived CLI invocations -- those exit immediately).
const raw = await shellSafe(
sandbox,
`pgrep -x agent-browser 2>/dev/null || true`,
);
if (!raw) {
// Fallback: broader match on process name
const fallback = await shellSafe(
sandbox,
`pgrep -f 'agent-browser' 2>/dev/null | head -5 || true`,
);
if (!fallback) return [];
return fallback.split("\n").map(Number).filter(Boolean);
}
return raw.split("\n").map(Number).filter(Boolean);
}
async function collectProcessMetrics(
sandbox: SandboxInstance,
pid: number,
): Promise<ProcessMetrics | null> {
const raw = await shellSafe(
sandbox,
`ps -p ${pid} -o pid=,rss=,vsz=,%cpu=,%mem=,cputime=,comm= 2>/dev/null || true`,
);
if (!raw) return null;
const parts = raw.trim().split(/\s+/);
if (parts.length < 7) return null;
// Parse cputime "HH:MM:SS" or "MM:SS" to seconds
const timeParts = parts[5].split(":").map(Number);
let cpuTimeSec = 0;
if (timeParts.length === 3) {
cpuTimeSec = timeParts[0] * 3600 + timeParts[1] * 60 + timeParts[2];
} else if (timeParts.length === 2) {
cpuTimeSec = timeParts[0] * 60 + timeParts[1];
}
return {
pid: Number(parts[0]),
rssKb: Number(parts[1]),
vszKb: Number(parts[2]),
cpuPercent: Number(parts[3]),
memPercent: Number(parts[4]),
cpuTimeSec,
command: parts.slice(6).join(" "),
};
}
async function getPeakRssKb(
sandbox: SandboxInstance,
pid: number,
): Promise<number> {
const raw = await shellSafe(
sandbox,
`cat /proc/${pid}/status 2>/dev/null | grep VmHWM | awk '{print $2}' || echo 0`,
);
return Number(raw) || 0;
}
async function getChildPids(
sandbox: SandboxInstance,
pid: number,
): Promise<number[]> {
const raw = await shellSafe(
sandbox,
`pgrep -P ${pid} 2>/dev/null || true`,
);
if (!raw) return [];
return raw.split("\n").map(Number).filter(Boolean);
}
async function getAllDescendantPids(
sandbox: SandboxInstance,
pid: number,
): Promise<number[]> {
const all: number[] = [];
const queue = [pid];
while (queue.length > 0) {
const current = queue.shift()!;
all.push(current);
const children = await getChildPids(sandbox, current);
queue.push(...children);
}
return all;
}
async function collectDaemonMetrics(
sandbox: SandboxInstance,
session: string,
coldStartMs: number,
binarySizeBytes: number,
distributionSizeBytes: number,
): Promise<DaemonMetrics> {
// Find daemon PIDs -- the agent-browser process itself
const daemonPids = await findDaemonPids(sandbox, session);
// Also find the full process tree (daemon + Chrome children)
let allPids: number[] = [];
for (const pid of daemonPids) {
const descendants = await getAllDescendantPids(sandbox, pid);
allPids.push(...descendants);
}
allPids = [...new Set(allPids)];
// If no daemon PIDs found via pgrep, fall back to grabbing all
// agent-browser and chrome processes for metrics
if (allPids.length === 0) {
const fallback = await shellSafe(
sandbox,
`ps -eo pid,comm | grep -E 'agent-browser|chrome' | grep -v grep | awk '{print $1}' || true`,
);
if (fallback) {
allPids = fallback.split("\n").map(Number).filter(Boolean);
}
}
const daemonProcs: ProcessMetrics[] = [];
const browserProcs: ProcessMetrics[] = [];
let daemonPeakRssKb = 0;
for (const pid of allPids) {
const metrics = await collectProcessMetrics(sandbox, pid);
if (!metrics) continue;
const isBrowser = /chrome|chromium/i.test(metrics.command);
if (isBrowser) {
browserProcs.push(metrics);
} else {
daemonProcs.push(metrics);
const peak = await getPeakRssKb(sandbox, pid);
daemonPeakRssKb = Math.max(daemonPeakRssKb, peak);
}
}
const daemonRssKb = daemonProcs.reduce((sum, p) => sum + p.rssKb, 0);
const browserRssKb = browserProcs.reduce((sum, p) => sum + p.rssKb, 0);
const daemonCpuTimeSec = daemonProcs.reduce((sum, p) => sum + p.cpuTimeSec, 0);
const allProcs = [...daemonProcs, ...browserProcs];
const totalCpuTimeSec = allProcs.reduce((sum, p) => sum + p.cpuTimeSec, 0);
return {
coldStartMs,
binarySizeBytes,
distributionSizeBytes,
daemonProcesses: daemonProcs,
browserProcesses: browserProcs,
daemonRssKb,
browserRssKb,
daemonPeakRssKb,
daemonCpuTimeSec,
totalCpuTimeSec,
};
}
async function getBinarySize(
sandbox: SandboxInstance,
): Promise<number> {
// Follow symlinks to get the real binary/script size
const raw = await shellSafe(
sandbox,
`stat -L -c %s "$(readlink -f "$(which agent-browser)")" 2>/dev/null || echo 0`,
);
return Number(raw) || 0;
}
async function getDistributionSize(
sandbox: SandboxInstance,
mode: DaemonMode,
): Promise<number> {
if (mode === "node") {
// Total size of the npm package + Playwright browser
const npmPkg = await shellSafe(
sandbox,
`du -sb "$(npm root -g)/agent-browser" 2>/dev/null | awk '{print $1}' || echo 0`,
);
const pwBrowser = await shellSafe(
sandbox,
`du -sb "$HOME/.cache/ms-playwright" 2>/dev/null | awk '{print $1}' || echo 0`,
);
return (Number(npmPkg) || 0) + (Number(pwBrowser) || 0);
} else {
// Rust binary + Chrome for Testing (checks multiple possible cache paths)
const binary = await shellSafe(
sandbox,
`stat -L -c %s "$(readlink -f "$(which agent-browser)")" 2>/dev/null || echo 0`,
);
const chrome = await shellSafe(
sandbox,
[
`size=0`,
`for d in "$HOME/.cache/agent-browser" "$HOME/.cache/ms-playwright" "$HOME/.agent-browser/chrome"; do`,
` if [ -d "$d" ]; then size=$(du -sb "$d" 2>/dev/null | awk '{print $1}'); break; fi`,
`done`,
`echo $size`,
].join("; "),
);
return (Number(binary) || 0) + (Number(chrome) || 0);
}
}
function formatBytes(bytes: number): string {
if (bytes >= 1024 * 1024) return `${(bytes / 1024 / 1024).toFixed(1)} MB`;
if (bytes >= 1024) return `${(bytes / 1024).toFixed(1)} KB`;
return `${bytes} B`;
}
function formatKb(kb: number): string {
if (kb >= 1024) return `${(kb / 1024).toFixed(1)} MB`;
return `${kb} KB`;
}
// ---------------------------------------------------------------------------
// Scenario runner
// ---------------------------------------------------------------------------
type DaemonMode = "node" | "native";
function daemonEnv(mode: DaemonMode): Record<string, string> {
return { AGENT_BROWSER_SESSION: `bench-${mode}` };
}
async function agentBrowser(
sandbox: SandboxInstance,
args: string[],
mode: DaemonMode,
): Promise<void> {
const result = await sandbox.runCommand({
cmd: "agent-browser",
args,
env: daemonEnv(mode),
});
if (result.exitCode !== 0) {
const stderr = await result.stderr();
const stdout = await result.stdout();
throw new Error(
`agent-browser ${args.join(" ")} failed (exit ${result.exitCode}): ${stderr || stdout}`,
);
}
}
async function timedAgentBrowser(
sandbox: SandboxInstance,
args: string[],
mode: DaemonMode,
): Promise<number> {
const start = Date.now();
const result = await sandbox.runCommand({
cmd: "agent-browser",
args,
env: daemonEnv(mode),
});
const elapsed = Date.now() - start;
if (result.exitCode !== 0) {
const stderr = await result.stderr();
const stdout = await result.stdout();
throw new Error(
`agent-browser ${args.join(" ")} failed (exit ${result.exitCode}): ${stderr || stdout}`,
);
}
return elapsed;
}
interface ScenarioResult {
name: string;
description: string;
stats: Stats;
error?: string;
}
async function runScenario(
sandbox: SandboxInstance,
scenario: Scenario,
mode: DaemonMode,
iterations: number,
warmup: number,
): Promise<ScenarioResult> {
try {
if (scenario.setup) {
for (const cmd of scenario.setup) {
await agentBrowser(sandbox, cmd, mode);
}
}
for (let w = 0; w < warmup; w++) {
for (const cmd of scenario.commands) {
await agentBrowser(sandbox, cmd, mode);
}
}
const samples: number[] = [];
for (let i = 0; i < iterations; i++) {
let totalMs = 0;
for (const cmd of scenario.commands) {
totalMs += await timedAgentBrowser(sandbox, cmd, mode);
}
samples.push(totalMs);
}
if (scenario.teardown) {
for (const cmd of scenario.teardown) {
await agentBrowser(sandbox, cmd, mode);
}
}
return {
name: scenario.name,
description: scenario.description,
stats: computeStats(samples),
};
} catch (err: unknown) {
const message = err instanceof Error ? err.message : String(err);
return {
name: scenario.name,
description: scenario.description,
stats: { avgMs: -1, stddevMs: -1, minMs: -1, maxMs: -1, p50Ms: -1, samples: [] },
error: message,
};
}
}
// ---------------------------------------------------------------------------
// Benchmark phases
// ---------------------------------------------------------------------------
interface DaemonResults {
mode: DaemonMode;
label: string;
scenarios: ScenarioResult[];
metrics: DaemonMetrics;
}
async function benchmarkDaemon(
sandbox: SandboxInstance,
mode: DaemonMode,
label: string,
): Promise<DaemonResults> {
console.log(`\n--- ${label} ---`);
// Measure sizes before launch
const binarySizeBytes = await getBinarySize(sandbox);
const distributionSizeBytes = await getDistributionSize(sandbox, mode);
// Cold start: time the first launch (daemon spawn + browser launch)
const coldStartBegin = Date.now();
await agentBrowser(sandbox, ["open", "about:blank"], mode);
const coldStartMs = Date.now() - coldStartBegin;
console.log(` Cold start: ${coldStartMs}ms`);
console.log(` Binary size: ${formatBytes(binarySizeBytes)}`);
console.log(` Distribution size: ${formatBytes(distributionSizeBytes)}`);
// Run all scenarios
const results: ScenarioResult[] = [];
for (const scenario of scenarios) {
process.stdout.write(` ${scenario.name} `);
const result = await runScenario(
sandbox,
scenario,
mode,
config.iterations,
config.warmup,
);
if (result.error) {
console.log(`FAILED: ${result.error.slice(0, 120)}`);
} else {
const dots = ".".repeat(Math.max(1, 30 - scenario.name.length));
const s = result.stats;
console.log(
`${dots} ${s.avgMs}ms avg +/-${s.stddevMs}ms (p50: ${s.p50Ms}ms, min: ${s.minMs}ms, max: ${s.maxMs}ms)`,
);
}
results.push(result);
}
// Collect system metrics after scenarios (daemon is still running)
const session = `bench-${mode}`;
const metrics = await collectDaemonMetrics(
sandbox,
session,
coldStartMs,
binarySizeBytes,
distributionSizeBytes,
);
// Also grab a full process snapshot for context
const psOutput = await shellSafe(
sandbox,
`ps aux --sort=-rss | head -20`,
);
console.log(`\n Process snapshot (top by RSS):`);
for (const line of psOutput.split("\n").slice(0, 10)) {
console.log(` ${line}`);
}
console.log(`\n Daemon processes (${metrics.daemonProcesses.length}):`);
console.log(` RSS: ${formatKb(metrics.daemonRssKb)} (peak: ${formatKb(metrics.daemonPeakRssKb)})`);
console.log(` CPU time: ${metrics.daemonCpuTimeSec.toFixed(1)}s`);
for (const p of metrics.daemonProcesses) {
console.log(` PID ${p.pid}: ${p.command} (RSS: ${formatKb(p.rssKb)}, CPU: ${p.cpuPercent}%)`);
}
console.log(` Browser processes (${metrics.browserProcesses.length}):`);
console.log(` RSS: ${formatKb(metrics.browserRssKb)}`);
for (const p of metrics.browserProcesses) {
console.log(` PID ${p.pid}: ${p.command} (RSS: ${formatKb(p.rssKb)}, CPU: ${p.cpuPercent}%)`);
}
await agentBrowser(sandbox, ["close"], mode);
console.log(` Browser closed.`);
return { mode, label, scenarios: results, metrics };
}
// ---------------------------------------------------------------------------
// Install helpers
// ---------------------------------------------------------------------------
async function installChromiumDeps(sandbox: SandboxInstance) {
console.log("Installing Chromium system dependencies...");
await shell(
sandbox,
`sudo dnf clean all 2>&1 && sudo dnf install -y --skip-broken ${CHROMIUM_SYSTEM_DEPS.join(" ")} 2>&1 && sudo ldconfig 2>&1`,
);
}
async function installNodeDaemon(sandbox: SandboxInstance) {
console.log("Installing agent-browser from npm (Node.js daemon)...");
await run(sandbox, "npm", ["install", "-g", "agent-browser"]);
await run(sandbox, "npx", ["agent-browser", "install"]);
const version = await shell(sandbox, "agent-browser --version 2>&1 || true");
console.log(` version: ${version.trim()}`);
}
async function installNativeDaemon(sandbox: SandboxInstance, branch: string) {
console.log(`\nBuilding native daemon from ${branch}...`);
console.log(" Installing build tools and Rust toolchain...");
const rustStart = Date.now();
await shell(
sandbox,
"sudo dnf install -y gcc gcc-c++ make perl-core openssl-devel 2>&1",
);
await shell(
sandbox,
"curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y 2>&1",
);
console.log(` Rust + build tools installed (${Math.round((Date.now() - rustStart) / 1000)}s)`);
console.log(` Cloning repo (branch: ${branch})...`);
const cloneStart = Date.now();
await shell(
sandbox,
`git clone --depth 1 --branch ${branch} ${REPO_URL} /tmp/agent-browser 2>&1`,
);
console.log(` Cloned (${Math.round((Date.now() - cloneStart) / 1000)}s)`);
console.log(" Building release binary (cargo build --release)...");
const buildStart = Date.now();
await shell(
sandbox,
"source $HOME/.cargo/env && cd /tmp/agent-browser/cli && cargo build --release 2>&1",
);
console.log(` Built (${Math.round((Date.now() - buildStart) / 1000)}s)`);
const npmBinPath = (await shell(sandbox, "which agent-browser")).trim();
console.log(` Replacing ${npmBinPath} with native build...`);
await shell(
sandbox,
`sudo cp /tmp/agent-browser/cli/target/release/agent-browser ${npmBinPath}`,
);
const version = await shell(sandbox, "agent-browser --version 2>&1 || true");
console.log(` version: ${version.trim()}`);
}
// ---------------------------------------------------------------------------
// Output
// ---------------------------------------------------------------------------
function printResults(node: DaemonResults, native: DaemonResults) {
console.log("\n\n========== COMMAND LATENCY ==========\n");
const header =
"Scenario".padEnd(20) + "| Node avg +/-sd | Rust avg +/-sd | Speedup";
const sep = "-".repeat(20) + "|-----------------|-----------------|--------";
console.log(header);
console.log(sep);
for (let i = 0; i < node.scenarios.length; i++) {
const n = node.scenarios[i];
const r = native.scenarios[i];
const name = n.name.padEnd(20);
if (n.error || r.error) {
const nodeVal = n.error ? "FAILED".padEnd(15) : `${n.stats.avgMs}ms`.padEnd(15);
const rustVal = r.error ? "FAILED".padEnd(15) : `${r.stats.avgMs}ms`.padEnd(15);
console.log(`${name}| ${nodeVal} | ${rustVal} | --`);
continue;
}
const nodeVal = `${n.stats.avgMs} +/-${n.stats.stddevMs}ms`.padEnd(15);
const rustVal = `${r.stats.avgMs} +/-${r.stats.stddevMs}ms`.padEnd(15);
const speedup =
r.stats.avgMs > 0
? (n.stats.avgMs / r.stats.avgMs).toFixed(2) + "x"
: "--";
console.log(`${name}| ${nodeVal} | ${rustVal} | ${speedup.padStart(6)}`);
}
console.log("\n\n========== SYSTEM METRICS ==========\n");
const nm = node.metrics;
const rm = native.metrics;
function ratio(a: number, b: number): string {
if (b <= 0) return "--";
return (a / b).toFixed(2) + "x";
}
const metricRows: [string, string, string, string][] = [
[
"Cold start",
`${nm.coldStartMs}ms`,
`${rm.coldStartMs}ms`,
ratio(nm.coldStartMs, rm.coldStartMs),
],
[
"Binary size",
formatBytes(nm.binarySizeBytes),
formatBytes(rm.binarySizeBytes),
ratio(nm.binarySizeBytes, rm.binarySizeBytes),
],
[
"Distribution size",
formatBytes(nm.distributionSizeBytes),
formatBytes(rm.distributionSizeBytes),
ratio(nm.distributionSizeBytes, rm.distributionSizeBytes),
],
[
"Daemon RSS",
formatKb(nm.daemonRssKb),
formatKb(rm.daemonRssKb),
ratio(nm.daemonRssKb, rm.daemonRssKb),
],
[
"Daemon peak RSS",
formatKb(nm.daemonPeakRssKb),
formatKb(rm.daemonPeakRssKb),
ratio(nm.daemonPeakRssKb, rm.daemonPeakRssKb),
],
[
"Browser RSS",
formatKb(nm.browserRssKb),
formatKb(rm.browserRssKb),
ratio(nm.browserRssKb, rm.browserRssKb),
],
[
"Daemon CPU time",
`${nm.daemonCpuTimeSec.toFixed(1)}s`,
`${rm.daemonCpuTimeSec.toFixed(1)}s`,
ratio(nm.daemonCpuTimeSec, rm.daemonCpuTimeSec),
],
[
"Daemon processes",
String(nm.daemonProcesses.length),
String(rm.daemonProcesses.length),
"--",
],
[
"Browser processes",
String(nm.browserProcesses.length),
String(rm.browserProcesses.length),
"--",
],
];
const mHeader =
"Metric".padEnd(20) + "| Node".padEnd(14) + "| Rust".padEnd(14) + "| Ratio";
const mSep = "-".repeat(20) + "|" + "-".repeat(13) + "|" + "-".repeat(13) + "|--------";
console.log(mHeader);
console.log(mSep);
for (const [metric, nodeVal, rustVal, ratio] of metricRows) {
console.log(
`${metric.padEnd(20)}| ${nodeVal.padEnd(12)}| ${rustVal.padEnd(12)}| ${ratio}`,
);
}
}
// ---------------------------------------------------------------------------
// Main
// ---------------------------------------------------------------------------
async function main() {
console.log("agent-browser Daemon Benchmark (Node.js vs Rust Native)");
console.log(`Branch: ${config.branch}`);
console.log(`Iterations: ${config.iterations} (+ ${config.warmup} warmup)`);
console.log(`vCPUs: ${config.vcpus}\n`);
console.log("Creating sandbox...");
const sandbox = await Sandbox.create({
...credentials,
timeout: TIMEOUT_MS,
runtime: "node22",
networkPolicy: "allow-all" as const,
resources: { vcpus: config.vcpus },
});
console.log(`Sandbox: ${sandbox.sandboxId}`);
try {
await installChromiumDeps(sandbox);
// Phase 1: Node.js daemon (from published npm package)
await installNodeDaemon(sandbox);
const nodeResults = await benchmarkDaemon(
sandbox,
"node",
"Node.js Daemon (npm)",
);
// Phase 2: Rust native daemon (built from branch)
await installNativeDaemon(sandbox, config.branch);
const nativeResults = await benchmarkDaemon(
sandbox,
"native",
`Rust Native Daemon (${config.branch})`,
);
printResults(nodeResults, nativeResults);
if (config.json) {
const output = {
timestamp: new Date().toISOString(),
branch: config.branch,
vcpus: config.vcpus,
iterations: config.iterations,
warmup: config.warmup,
node: {
scenarios: nodeResults.scenarios.map((s) => ({
name: s.name,
description: s.description,
...s.stats,
error: s.error,
})),
metrics: nodeResults.metrics,
},
native: {
scenarios: nativeResults.scenarios.map((s) => ({
name: s.name,
description: s.description,
...s.stats,
error: s.error,
})),
metrics: nativeResults.metrics,
},
};
writeFileSync("results.json", JSON.stringify(output, null, 2));
console.log("\nResults written to results.json");
}
} catch (err: unknown) {
const message = err instanceof Error ? err.message : String(err);
console.error(`\nFatal error: ${message}`);
process.exit(1);
} finally {
try {
await sandbox.stop();
console.log("\nSandbox stopped.");
} catch {
console.warn("Warning: failed to stop sandbox.");
}
}
}
main();
-13
View File
@@ -1,13 +0,0 @@
{
"name": "agent-browser-benchmarks",
"version": "1.0.0",
"private": true,
"type": "module",
"scripts": {
"bench": "tsx bench.ts"
},
"dependencies": {
"@vercel/sandbox": "^1.8.0",
"tsx": "^4.19.0"
}
}
-472
View File
@@ -1,472 +0,0 @@
lockfileVersion: '9.0'
settings:
autoInstallPeers: true
excludeLinksFromLockfile: false
importers:
.:
dependencies:
'@vercel/sandbox':
specifier: ^1.8.0
version: 1.8.1
tsx:
specifier: ^4.19.0
version: 4.21.0
packages:
'@esbuild/aix-ppc64@0.27.4':
resolution: {integrity: sha512-cQPwL2mp2nSmHHJlCyoXgHGhbEPMrEEU5xhkcy3Hs/O7nGZqEpZ2sUtLaL9MORLtDfRvVl2/3PAuEkYZH0Ty8Q==}
engines: {node: '>=18'}
cpu: [ppc64]
os: [aix]
'@esbuild/android-arm64@0.27.4':
resolution: {integrity: sha512-gdLscB7v75wRfu7QSm/zg6Rx29VLdy9eTr2t44sfTW7CxwAtQghZ4ZnqHk3/ogz7xao0QAgrkradbBzcqFPasw==}
engines: {node: '>=18'}
cpu: [arm64]
os: [android]
'@esbuild/android-arm@0.27.4':
resolution: {integrity: sha512-X9bUgvxiC8CHAGKYufLIHGXPJWnr0OCdR0anD2e21vdvgCI8lIfqFbnoeOz7lBjdrAGUhqLZLcQo6MLhTO2DKQ==}
engines: {node: '>=18'}
cpu: [arm]
os: [android]
'@esbuild/android-x64@0.27.4':
resolution: {integrity: sha512-PzPFnBNVF292sfpfhiyiXCGSn9HZg5BcAz+ivBuSsl6Rk4ga1oEXAamhOXRFyMcjwr2DVtm40G65N3GLeH1Lvw==}
engines: {node: '>=18'}
cpu: [x64]
os: [android]
'@esbuild/darwin-arm64@0.27.4':
resolution: {integrity: sha512-b7xaGIwdJlht8ZFCvMkpDN6uiSmnxxK56N2GDTMYPr2/gzvfdQN8rTfBsvVKmIVY/X7EM+/hJKEIbbHs9oA4tQ==}
engines: {node: '>=18'}
cpu: [arm64]
os: [darwin]
'@esbuild/darwin-x64@0.27.4':
resolution: {integrity: sha512-sR+OiKLwd15nmCdqpXMnuJ9W2kpy0KigzqScqHI3Hqwr7IXxBp3Yva+yJwoqh7rE8V77tdoheRYataNKL4QrPw==}
engines: {node: '>=18'}
cpu: [x64]
os: [darwin]
'@esbuild/freebsd-arm64@0.27.4':
resolution: {integrity: sha512-jnfpKe+p79tCnm4GVav68A7tUFeKQwQyLgESwEAUzyxk/TJr4QdGog9sqWNcUbr/bZt/O/HXouspuQDd9JxFSw==}
engines: {node: '>=18'}
cpu: [arm64]
os: [freebsd]
'@esbuild/freebsd-x64@0.27.4':
resolution: {integrity: sha512-2kb4ceA/CpfUrIcTUl1wrP/9ad9Atrp5J94Lq69w7UwOMolPIGrfLSvAKJp0RTvkPPyn6CIWrNy13kyLikZRZQ==}
engines: {node: '>=18'}
cpu: [x64]
os: [freebsd]
'@esbuild/linux-arm64@0.27.4':
resolution: {integrity: sha512-7nQOttdzVGth1iz57kxg9uCz57dxQLHWxopL6mYuYthohPKEK0vU0C3O21CcBK6KDlkYVcnDXY099HcCDXd9dA==}
engines: {node: '>=18'}
cpu: [arm64]
os: [linux]
'@esbuild/linux-arm@0.27.4':
resolution: {integrity: sha512-aBYgcIxX/wd5n2ys0yESGeYMGF+pv6g0DhZr3G1ZG4jMfruU9Tl1i2Z+Wnj9/KjGz1lTLCcorqE2viePZqj4Eg==}
engines: {node: '>=18'}
cpu: [arm]
os: [linux]
'@esbuild/linux-ia32@0.27.4':
resolution: {integrity: sha512-oPtixtAIzgvzYcKBQM/qZ3R+9TEUd1aNJQu0HhGyqtx6oS7qTpvjheIWBbes4+qu1bNlo2V4cbkISr8q6gRBFA==}
engines: {node: '>=18'}
cpu: [ia32]
os: [linux]
'@esbuild/linux-loong64@0.27.4':
resolution: {integrity: sha512-8mL/vh8qeCoRcFH2nM8wm5uJP+ZcVYGGayMavi8GmRJjuI3g1v6Z7Ni0JJKAJW+m0EtUuARb6Lmp4hMjzCBWzA==}
engines: {node: '>=18'}
cpu: [loong64]
os: [linux]
'@esbuild/linux-mips64el@0.27.4':
resolution: {integrity: sha512-1RdrWFFiiLIW7LQq9Q2NES+HiD4NyT8Itj9AUeCl0IVCA459WnPhREKgwrpaIfTOe+/2rdntisegiPWn/r/aAw==}
engines: {node: '>=18'}
cpu: [mips64el]
os: [linux]
'@esbuild/linux-ppc64@0.27.4':
resolution: {integrity: sha512-tLCwNG47l3sd9lpfyx9LAGEGItCUeRCWeAx6x2Jmbav65nAwoPXfewtAdtbtit/pJFLUWOhpv0FpS6GQAmPrHA==}
engines: {node: '>=18'}
cpu: [ppc64]
os: [linux]
'@esbuild/linux-riscv64@0.27.4':
resolution: {integrity: sha512-BnASypppbUWyqjd1KIpU4AUBiIhVr6YlHx/cnPgqEkNoVOhHg+YiSVxM1RLfiy4t9cAulbRGTNCKOcqHrEQLIw==}
engines: {node: '>=18'}
cpu: [riscv64]
os: [linux]
'@esbuild/linux-s390x@0.27.4':
resolution: {integrity: sha512-+eUqgb/Z7vxVLezG8bVB9SfBie89gMueS+I0xYh2tJdw3vqA/0ImZJ2ROeWwVJN59ihBeZ7Tu92dF/5dy5FttA==}
engines: {node: '>=18'}
cpu: [s390x]
os: [linux]
'@esbuild/linux-x64@0.27.4':
resolution: {integrity: sha512-S5qOXrKV8BQEzJPVxAwnryi2+Iq5pB40gTEIT69BQONqR7JH1EPIcQ/Uiv9mCnn05jff9umq/5nqzxlqTOg9NA==}
engines: {node: '>=18'}
cpu: [x64]
os: [linux]
'@esbuild/netbsd-arm64@0.27.4':
resolution: {integrity: sha512-xHT8X4sb0GS8qTqiwzHqpY00C95DPAq7nAwX35Ie/s+LO9830hrMd3oX0ZMKLvy7vsonee73x0lmcdOVXFzd6Q==}
engines: {node: '>=18'}
cpu: [arm64]
os: [netbsd]
'@esbuild/netbsd-x64@0.27.4':
resolution: {integrity: sha512-RugOvOdXfdyi5Tyv40kgQnI0byv66BFgAqjdgtAKqHoZTbTF2QqfQrFwa7cHEORJf6X2ht+l9ABLMP0dnKYsgg==}
engines: {node: '>=18'}
cpu: [x64]
os: [netbsd]
'@esbuild/openbsd-arm64@0.27.4':
resolution: {integrity: sha512-2MyL3IAaTX+1/qP0O1SwskwcwCoOI4kV2IBX1xYnDDqthmq5ArrW94qSIKCAuRraMgPOmG0RDTA74mzYNQA9ow==}
engines: {node: '>=18'}
cpu: [arm64]
os: [openbsd]
'@esbuild/openbsd-x64@0.27.4':
resolution: {integrity: sha512-u8fg/jQ5aQDfsnIV6+KwLOf1CmJnfu1ShpwqdwC0uA7ZPwFws55Ngc12vBdeUdnuWoQYx/SOQLGDcdlfXhYmXQ==}
engines: {node: '>=18'}
cpu: [x64]
os: [openbsd]
'@esbuild/openharmony-arm64@0.27.4':
resolution: {integrity: sha512-JkTZrl6VbyO8lDQO3yv26nNr2RM2yZzNrNHEsj9bm6dOwwu9OYN28CjzZkH57bh4w0I2F7IodpQvUAEd1mbWXg==}
engines: {node: '>=18'}
cpu: [arm64]
os: [openharmony]
'@esbuild/sunos-x64@0.27.4':
resolution: {integrity: sha512-/gOzgaewZJfeJTlsWhvUEmUG4tWEY2Spp5M20INYRg2ZKl9QPO3QEEgPeRtLjEWSW8FilRNacPOg8R1uaYkA6g==}
engines: {node: '>=18'}
cpu: [x64]
os: [sunos]
'@esbuild/win32-arm64@0.27.4':
resolution: {integrity: sha512-Z9SExBg2y32smoDQdf1HRwHRt6vAHLXcxD2uGgO/v2jK7Y718Ix4ndsbNMU/+1Qiem9OiOdaqitioZwxivhXYg==}
engines: {node: '>=18'}
cpu: [arm64]
os: [win32]
'@esbuild/win32-ia32@0.27.4':
resolution: {integrity: sha512-DAyGLS0Jz5G5iixEbMHi5KdiApqHBWMGzTtMiJ72ZOLhbu/bzxgAe8Ue8CTS3n3HbIUHQz/L51yMdGMeoxXNJw==}
engines: {node: '>=18'}
cpu: [ia32]
os: [win32]
'@esbuild/win32-x64@0.27.4':
resolution: {integrity: sha512-+knoa0BDoeXgkNvvV1vvbZX4+hizelrkwmGJBdT17t8FNPwG2lKemmuMZlmaNQ3ws3DKKCxpb4zRZEIp3UxFCg==}
engines: {node: '>=18'}
cpu: [x64]
os: [win32]
'@vercel/oidc@3.2.0':
resolution: {integrity: sha512-UycprH3T6n3jH0k44NHMa7pnFHGu/N05MjojYr+Mc6I7obkoLIJujSWwin1pCvdy/eOxrI/l3uDLQsmcrOb4ug==}
engines: {node: '>= 20'}
'@vercel/sandbox@1.8.1':
resolution: {integrity: sha512-txohjI20aMxZiAzBL/KJi5EqTYsesBdOyIOtpTIyebPLTqYtDYfNhQ4OeYiUcPMUo0XBt8gSet/rIdLQEjj3/A==}
async-retry@1.3.3:
resolution: {integrity: sha512-wfr/jstw9xNi/0teMHrRW7dsz3Lt5ARhYNZ2ewpadnhaIp5mbALhOAP+EAdsC7t4Z6wqsDVv9+W6gm1Dk9mEyw==}
b4a@1.8.0:
resolution: {integrity: sha512-qRuSmNSkGQaHwNbM7J78Wwy+ghLEYF1zNrSeMxj4Kgw6y33O3mXcQ6Ie9fRvfU/YnxWkOchPXbaLb73TkIsfdg==}
peerDependencies:
react-native-b4a: '*'
peerDependenciesMeta:
react-native-b4a:
optional: true
bare-events@2.8.2:
resolution: {integrity: sha512-riJjyv1/mHLIPX4RwiK+oW9/4c3TEUeORHKefKAKnZ5kyslbN+HXowtbaVEqt4IMUB7OXlfixcs6gsFeo/jhiQ==}
peerDependencies:
bare-abort-controller: '*'
peerDependenciesMeta:
bare-abort-controller:
optional: true
esbuild@0.27.4:
resolution: {integrity: sha512-Rq4vbHnYkK5fws5NF7MYTU68FPRE1ajX7heQ/8QXXWqNgqqJ/GkmmyxIzUnf2Sr/bakf8l54716CcMGHYhMrrQ==}
engines: {node: '>=18'}
hasBin: true
events-universal@1.0.1:
resolution: {integrity: sha512-LUd5euvbMLpwOF8m6ivPCbhQeSiYVNb8Vs0fQ8QjXo0JTkEHpz8pxdQf0gStltaPpw0Cca8b39KxvK9cfKRiAw==}
fast-fifo@1.3.2:
resolution: {integrity: sha512-/d9sfos4yxzpwkDkuN7k2SqFKtYNmCTzgfEpz82x34IM9/zc8KGxQoXg1liNC/izpRM/MBdt44Nmx41ZWqk+FQ==}
fsevents@2.3.3:
resolution: {integrity: sha512-5xoDfX+fL7faATnagmWPpbFtwh/R77WmMMqqHGS65C3vvB0YHrgF+B1YmZ3441tMj5n63k0212XNoJwzlhffQw==}
engines: {node: ^8.16.0 || ^10.6.0 || >=11.0.0}
os: [darwin]
get-tsconfig@4.13.6:
resolution: {integrity: sha512-shZT/QMiSHc/YBLxxOkMtgSid5HFoauqCE3/exfsEcwg1WkeqjG+V40yBbBrsD+jW2HDXcs28xOfcbm2jI8Ddw==}
jsonlines@0.1.1:
resolution: {integrity: sha512-ekDrAGso79Cvf+dtm+mL8OBI2bmAOt3gssYs833De/C9NmIpWDWyUO4zPgB5x2/OhY366dkhgfPMYfwZF7yOZA==}
ms@2.1.3:
resolution: {integrity: sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==}
os-paths@4.4.0:
resolution: {integrity: sha512-wrAwOeXp1RRMFfQY8Sy7VaGVmPocaLwSFOYCGKSyo8qmJ+/yaafCl5BCA1IQZWqFSRBrKDYFeR9d/VyQzfH/jg==}
engines: {node: '>= 6.0'}
picocolors@1.1.1:
resolution: {integrity: sha512-xceH2snhtb5M9liqDsmEw56le376mTZkEX/jEb/RxNFyegNul7eNslCXP9FDj/Lcu0X8KEyMceP2ntpaHrDEVA==}
resolve-pkg-maps@1.0.0:
resolution: {integrity: sha512-seS2Tj26TBVOC2NIc2rOe2y2ZO7efxITtLZcGSOnHHNOQ7CkiUBfw0Iw2ck6xkIhPwLhKNLS8BO+hEpngQlqzw==}
retry@0.13.1:
resolution: {integrity: sha512-XQBQ3I8W1Cge0Seh+6gjj03LbmRFWuoszgK9ooCpwYIrhhoO80pfq4cUkU5DkknwfOfFteRwlZ56PYOGYyFWdg==}
engines: {node: '>= 4'}
streamx@2.23.0:
resolution: {integrity: sha512-kn+e44esVfn2Fa/O0CPFcex27fjIL6MkVae0Mm6q+E6f0hWv578YCERbv+4m02cjxvDsPKLnmxral/rR6lBMAg==}
tar-stream@3.1.7:
resolution: {integrity: sha512-qJj60CXt7IU1Ffyc3NJMjh6EkuCFej46zUqJ4J7pqYlThyd9bO0XBTmcOIhSzZJVWfsLks0+nle/j538YAW9RQ==}
text-decoder@1.2.7:
resolution: {integrity: sha512-vlLytXkeP4xvEq2otHeJfSQIRyWxo/oZGEbXrtEEF9Hnmrdly59sUbzZ/QgyWuLYHctCHxFF4tRQZNQ9k60ExQ==}
tsx@4.21.0:
resolution: {integrity: sha512-5C1sg4USs1lfG0GFb2RLXsdpXqBSEhAaA/0kPL01wxzpMqLILNxIxIOKiILz+cdg/pLnOUxFYOR5yhHU666wbw==}
engines: {node: '>=18.0.0'}
hasBin: true
undici@7.24.1:
resolution: {integrity: sha512-5xoBibbmnjlcR3jdqtY2Lnx7WbrD/tHlT01TmvqZUFVc9Q1w4+j5hbnapTqbcXITMH1ovjq/W7BkqBilHiVAaA==}
engines: {node: '>=20.18.1'}
xdg-app-paths@5.1.0:
resolution: {integrity: sha512-RAQ3WkPf4KTU1A8RtFx3gWywzVKe00tfOPFfl2NDGqbIFENQO4kqAJp7mhQjNj/33W5x5hiWWUdyfPq/5SU3QA==}
engines: {node: '>=6'}
xdg-portable@7.3.0:
resolution: {integrity: sha512-sqMMuL1rc0FmMBOzCpd0yuy9trqF2yTTVe+E9ogwCSWQCdDEtQUwrZPT6AxqtsFGRNxycgncbP/xmOOSPw5ZUw==}
engines: {node: '>= 6.0'}
zod@3.24.4:
resolution: {integrity: sha512-OdqJE9UDRPwWsrHjLN2F8bPxvwJBK22EHLWtanu0LSYr5YqzsaaW3RMgmjwr8Rypg5k+meEJdSPXJZXE/yqOMg==}
snapshots:
'@esbuild/aix-ppc64@0.27.4':
optional: true
'@esbuild/android-arm64@0.27.4':
optional: true
'@esbuild/android-arm@0.27.4':
optional: true
'@esbuild/android-x64@0.27.4':
optional: true
'@esbuild/darwin-arm64@0.27.4':
optional: true
'@esbuild/darwin-x64@0.27.4':
optional: true
'@esbuild/freebsd-arm64@0.27.4':
optional: true
'@esbuild/freebsd-x64@0.27.4':
optional: true
'@esbuild/linux-arm64@0.27.4':
optional: true
'@esbuild/linux-arm@0.27.4':
optional: true
'@esbuild/linux-ia32@0.27.4':
optional: true
'@esbuild/linux-loong64@0.27.4':
optional: true
'@esbuild/linux-mips64el@0.27.4':
optional: true
'@esbuild/linux-ppc64@0.27.4':
optional: true
'@esbuild/linux-riscv64@0.27.4':
optional: true
'@esbuild/linux-s390x@0.27.4':
optional: true
'@esbuild/linux-x64@0.27.4':
optional: true
'@esbuild/netbsd-arm64@0.27.4':
optional: true
'@esbuild/netbsd-x64@0.27.4':
optional: true
'@esbuild/openbsd-arm64@0.27.4':
optional: true
'@esbuild/openbsd-x64@0.27.4':
optional: true
'@esbuild/openharmony-arm64@0.27.4':
optional: true
'@esbuild/sunos-x64@0.27.4':
optional: true
'@esbuild/win32-arm64@0.27.4':
optional: true
'@esbuild/win32-ia32@0.27.4':
optional: true
'@esbuild/win32-x64@0.27.4':
optional: true
'@vercel/oidc@3.2.0': {}
'@vercel/sandbox@1.8.1':
dependencies:
'@vercel/oidc': 3.2.0
async-retry: 1.3.3
jsonlines: 0.1.1
ms: 2.1.3
picocolors: 1.1.1
tar-stream: 3.1.7
undici: 7.24.1
xdg-app-paths: 5.1.0
zod: 3.24.4
transitivePeerDependencies:
- bare-abort-controller
- react-native-b4a
async-retry@1.3.3:
dependencies:
retry: 0.13.1
b4a@1.8.0: {}
bare-events@2.8.2: {}
esbuild@0.27.4:
optionalDependencies:
'@esbuild/aix-ppc64': 0.27.4
'@esbuild/android-arm': 0.27.4
'@esbuild/android-arm64': 0.27.4
'@esbuild/android-x64': 0.27.4
'@esbuild/darwin-arm64': 0.27.4
'@esbuild/darwin-x64': 0.27.4
'@esbuild/freebsd-arm64': 0.27.4
'@esbuild/freebsd-x64': 0.27.4
'@esbuild/linux-arm': 0.27.4
'@esbuild/linux-arm64': 0.27.4
'@esbuild/linux-ia32': 0.27.4
'@esbuild/linux-loong64': 0.27.4
'@esbuild/linux-mips64el': 0.27.4
'@esbuild/linux-ppc64': 0.27.4
'@esbuild/linux-riscv64': 0.27.4
'@esbuild/linux-s390x': 0.27.4
'@esbuild/linux-x64': 0.27.4
'@esbuild/netbsd-arm64': 0.27.4
'@esbuild/netbsd-x64': 0.27.4
'@esbuild/openbsd-arm64': 0.27.4
'@esbuild/openbsd-x64': 0.27.4
'@esbuild/openharmony-arm64': 0.27.4
'@esbuild/sunos-x64': 0.27.4
'@esbuild/win32-arm64': 0.27.4
'@esbuild/win32-ia32': 0.27.4
'@esbuild/win32-x64': 0.27.4
events-universal@1.0.1:
dependencies:
bare-events: 2.8.2
transitivePeerDependencies:
- bare-abort-controller
fast-fifo@1.3.2: {}
fsevents@2.3.3:
optional: true
get-tsconfig@4.13.6:
dependencies:
resolve-pkg-maps: 1.0.0
jsonlines@0.1.1: {}
ms@2.1.3: {}
os-paths@4.4.0: {}
picocolors@1.1.1: {}
resolve-pkg-maps@1.0.0: {}
retry@0.13.1: {}
streamx@2.23.0:
dependencies:
events-universal: 1.0.1
fast-fifo: 1.3.2
text-decoder: 1.2.7
transitivePeerDependencies:
- bare-abort-controller
- react-native-b4a
tar-stream@3.1.7:
dependencies:
b4a: 1.8.0
fast-fifo: 1.3.2
streamx: 2.23.0
transitivePeerDependencies:
- bare-abort-controller
- react-native-b4a
text-decoder@1.2.7:
dependencies:
b4a: 1.8.0
transitivePeerDependencies:
- react-native-b4a
tsx@4.21.0:
dependencies:
esbuild: 0.27.4
get-tsconfig: 4.13.6
optionalDependencies:
fsevents: 2.3.3
undici@7.24.1: {}
xdg-app-paths@5.1.0:
dependencies:
xdg-portable: 7.3.0
xdg-portable@7.3.0:
dependencies:
os-paths: 4.4.0
zod@3.24.4: {}
-105
View File
@@ -1,105 +0,0 @@
/**
* Benchmark scenarios for comparing Node.js daemon vs Rust native daemon.
*
* Each scenario defines CLI commands run via `sandbox.runCommand("agent-browser", args)`.
* Setup/teardown commands run once and are not timed.
* The `commands` array is timed over N iterations.
*/
export interface Scenario {
name: string;
description: string;
setup?: string[][];
commands: string[][];
teardown?: string[][];
}
const FORM_HTML = [
"<html><head><title>Bench</title></head><body>",
"<h1>Benchmark Page</h1>",
"<input id='name' type='text' placeholder='Name'>",
"<input id='email' type='email' placeholder='Email'>",
"<select id='color'><option value='red'>Red</option><option value='blue'>Blue</option></select>",
"<input id='agree' type='checkbox'>",
"<textarea id='bio' placeholder='Bio'></textarea>",
"<button id='submit'>Submit</button>",
"<p id='status'>Ready</p>",
"<a id='link' href='javascript:void(0)' onclick=\"document.getElementById('status').textContent='Clicked'\">Click me</a>",
"<ul>",
...Array.from({ length: 20 }, (_, i) => `<li class='item'>Item ${i + 1}</li>`),
"</ul>",
"</body></html>",
].join("");
const INJECT_FORM_SCRIPT = `document.open(); document.write(${JSON.stringify(FORM_HTML)}); document.close(); 'ok'`;
const SETUP_PAGE: string[][] = [
["open", "about:blank"],
["eval", INJECT_FORM_SCRIPT],
];
export const scenarios: Scenario[] = [
{
name: "navigate",
description: "Page navigation (about:blank round-trip)",
commands: [["open", "about:blank"]],
},
{
name: "snapshot",
description: "DOM snapshot (accessibility tree)",
setup: SETUP_PAGE,
commands: [["snapshot"]],
},
{
name: "screenshot",
description: "Screenshot capture",
setup: SETUP_PAGE,
commands: [["screenshot"]],
},
{
name: "evaluate",
description: "JavaScript evaluation",
setup: SETUP_PAGE,
commands: [
[
"eval",
"document.title + ' ' + document.querySelectorAll('li').length",
],
],
},
{
name: "click",
description: "Element click interaction",
setup: SETUP_PAGE,
commands: [["click", "#link"]],
},
{
name: "fill",
description: "Form field fill",
setup: SETUP_PAGE,
commands: [["fill", "#name", "Benchmark User"]],
},
{
name: "agent-loop",
description: "AI agent loop: snapshot -> click -> snapshot (typical agent cycle)",
setup: SETUP_PAGE,
commands: [["snapshot"], ["click", "#link"], ["snapshot"]],
},
{
name: "full-workflow",
description:
"Realistic workflow: navigate, inject form, snapshot, click, fill, evaluate, screenshot",
commands: [
["open", "about:blank"],
["eval", INJECT_FORM_SCRIPT],
["snapshot"],
["click", "#link"],
["fill", "#name", "Agent User"],
[
"eval",
"document.getElementById('name').value",
],
["screenshot"],
],
},
];
-13
View File
@@ -1,13 +0,0 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "ESNext",
"moduleResolution": "bundler",
"esModuleInterop": true,
"strict": true,
"skipLibCheck": true,
"outDir": "dist",
"declaration": true
},
"include": ["*.ts"]
}
BIN
View File
Binary file not shown.
+1
View File
@@ -0,0 +1 @@
/Users/leo/github.com/agent-browser/cli/target/release/agent-browser: /Users/leo/github.com/agent-browser/cli/build.rs /Users/leo/github.com/agent-browser/cli/cdp-protocol/browser_protocol.json /Users/leo/github.com/agent-browser/cli/cdp-protocol/js_protocol.json /Users/leo/github.com/agent-browser/cli/src/color.rs /Users/leo/github.com/agent-browser/cli/src/commands.rs /Users/leo/github.com/agent-browser/cli/src/connection.rs /Users/leo/github.com/agent-browser/cli/src/flags.rs /Users/leo/github.com/agent-browser/cli/src/install.rs /Users/leo/github.com/agent-browser/cli/src/main.rs /Users/leo/github.com/agent-browser/cli/src/output.rs /Users/leo/github.com/agent-browser/cli/src/validation.rs
+148 -2
View File
@@ -44,8 +44,8 @@ dependencies = [
]
[[package]]
name = "agent-browser"
version = "0.24.0"
name = "agent-browser-stealth"
version = "0.27.0-fork.9"
dependencies = [
"aes-gcm",
"async-trait",
@@ -58,12 +58,15 @@ dependencies = [
"hmac",
"image",
"libc",
"regex-lite",
"reqwest",
"rust-embed",
"serde",
"serde_json",
"sha2",
"similar",
"socket2",
"tempfile",
"time",
"tokio",
"tokio-tungstenite",
@@ -529,6 +532,12 @@ dependencies = [
"zune-inflate",
]
[[package]]
name = "fastrand"
version = "2.4.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9f1f227452a390804cdb637b74a86990f2a7d7ba4b7d5693aac9b4dd6defd8d6"
[[package]]
name = "fax"
version = "0.2.6"
@@ -605,6 +614,12 @@ version = "0.3.32"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7e3450815272ef58cec6d564423f6e755e25379b217b0bc688e295ba24df6b1d"
[[package]]
name = "futures-io"
version = "0.3.32"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cecba35d7ad927e23624b22ad55235f2239cfa44fd10428eecbeba6d6a717718"
[[package]]
name = "futures-macro"
version = "0.3.32"
@@ -635,9 +650,11 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "389ca41296e6190b48053de0321d02a77f32f8a5d2461dd38762c0593805c6d6"
dependencies = [
"futures-core",
"futures-io",
"futures-macro",
"futures-sink",
"futures-task",
"memchr",
"pin-project-lite",
"slab",
]
@@ -1152,6 +1169,12 @@ dependencies = [
"libc",
]
[[package]]
name = "linux-raw-sys"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "df1d3c3b53da64cf5760482273a98e575c651a67eec7f77df96b5b642de8f039"
[[package]]
name = "litemap"
version = "0.8.1"
@@ -1669,6 +1692,12 @@ dependencies = [
"thiserror 1.0.69",
]
[[package]]
name = "regex-lite"
version = "0.1.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cab834c73d247e67f4fae452806d17d3c7501756d98c8808d7c9c7aa7d18f973"
[[package]]
name = "reqwest"
version = "0.12.28"
@@ -1678,6 +1707,7 @@ dependencies = [
"base64",
"bytes",
"futures-core",
"futures-util",
"http",
"http-body",
"http-body-util",
@@ -1697,12 +1727,14 @@ dependencies = [
"sync_wrapper",
"tokio",
"tokio-rustls",
"tokio-util",
"tower",
"tower-http",
"tower-service",
"url",
"wasm-bindgen",
"wasm-bindgen-futures",
"wasm-streams",
"web-sys",
"webpki-roots 1.0.5",
]
@@ -1727,12 +1759,59 @@ dependencies = [
"windows-sys 0.52.0",
]
[[package]]
name = "rust-embed"
version = "8.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "04113cb9355a377d83f06ef1f0a45b8ab8cd7d8b1288160717d66df5c7988d27"
dependencies = [
"rust-embed-impl",
"rust-embed-utils",
"walkdir",
]
[[package]]
name = "rust-embed-impl"
version = "8.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "da0902e4c7c8e997159ab384e6d0fc91c221375f6894346ae107f47dd0f3ccaa"
dependencies = [
"proc-macro2",
"quote",
"rust-embed-utils",
"syn",
"walkdir",
]
[[package]]
name = "rust-embed-utils"
version = "8.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5bcdef0be6fe7f6fa333b1073c949729274b05f123a0ad7efcb8efd878e5c3b1"
dependencies = [
"sha2",
"walkdir",
]
[[package]]
name = "rustc-hash"
version = "2.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "357703d41365b4b27c590e3ed91eabb1b663f07c4c084095e60cbed4362dff0d"
[[package]]
name = "rustix"
version = "1.1.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "146c9e247ccc180c1f61615433868c99f3de3ae256a30a43b49f67c2d9171f34"
dependencies = [
"bitflags",
"errno",
"libc",
"linux-raw-sys",
"windows-sys 0.61.2",
]
[[package]]
name = "rustls"
version = "0.23.37"
@@ -1780,6 +1859,15 @@ version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9774ba4a74de5f7b1c1451ed6cd5285a32eddb5cccb8cc655a4e50009e06477f"
[[package]]
name = "same-file"
version = "1.0.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "93fc1dc3aaa9bfed95e02e6eadabb4baf7e3078b0bd1b4d7b6b0b68378900502"
dependencies = [
"winapi-util",
]
[[package]]
name = "semver"
version = "1.0.27"
@@ -1965,6 +2053,19 @@ dependencies = [
"syn",
]
[[package]]
name = "tempfile"
version = "3.25.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0136791f7c95b1f6dd99f9cc786b91bb81c3800b639b3478e561ddb7be95e5f1"
dependencies = [
"fastrand",
"getrandom 0.4.1",
"once_cell",
"rustix",
"windows-sys 0.61.2",
]
[[package]]
name = "thiserror"
version = "1.0.69"
@@ -2128,6 +2229,19 @@ dependencies = [
"webpki-roots 0.26.11",
]
[[package]]
name = "tokio-util"
version = "0.7.18"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9ae9cec805b01e8fc3fd2fe289f89149a9b66dd16786abd8b19cfa7b48cb0098"
dependencies = [
"bytes",
"futures-core",
"futures-sink",
"pin-project-lite",
"tokio",
]
[[package]]
name = "tower"
version = "0.5.3"
@@ -2316,6 +2430,16 @@ version = "0.9.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a"
[[package]]
name = "walkdir"
version = "2.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "29790946404f91d9c5d06f9874efddea1dc06c5efe94541a7d6863108e3a5e4b"
dependencies = [
"same-file",
"winapi-util",
]
[[package]]
name = "want"
version = "0.3.1"
@@ -2430,6 +2554,19 @@ dependencies = [
"wasmparser",
]
[[package]]
name = "wasm-streams"
version = "0.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "15053d8d85c7eccdbefef60f06769760a563c7f0a9d6902a13d35c7800b0ad65"
dependencies = [
"futures-util",
"js-sys",
"wasm-bindgen",
"wasm-bindgen-futures",
"web-sys",
]
[[package]]
name = "wasmparser"
version = "0.244.0"
@@ -2486,6 +2623,15 @@ version = "0.1.12"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a28ac98ddc8b9274cb41bb4d9d4d5c425b6020c50c46f25559911905610b4a88"
[[package]]
name = "winapi-util"
version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c2a7b1c03c876122aa43f3020e6c3c3ee5c05081c9a00739faf7503aeba10d22"
dependencies = [
"windows-sys 0.61.2",
]
[[package]]
name = "windows-core"
version = "0.62.2"
+14 -5
View File
@@ -1,18 +1,23 @@
[package]
name = "agent-browser"
version = "0.24.0"
name = "agent-browser-stealth"
version = "0.27.0-fork.9"
edition = "2021"
description = "Fast browser automation CLI for AI agents"
license = "Apache-2.0"
repository = "https://github.com/vercel-labs/agent-browser"
homepage = "https://agent-browser.dev"
repository = "https://github.com/leeguooooo/agent-browser-stealth"
homepage = "https://github.com/leeguooooo/agent-browser-stealth"
readme = "../README.md"
keywords = ["browser", "automation", "ai", "cdp", "chrome"]
categories = ["command-line-utilities", "web-programming"]
[[bin]]
name = "agent-browser"
path = "src/main.rs"
[dependencies]
serde = { version = "1.0", features = ["derive"] }
serde_json = "1.0"
regex-lite = "0.1"
dirs = "5.0"
base64 = "0.22"
getrandom = "0.2"
@@ -22,7 +27,7 @@ futures-util = "0.3"
url = "2"
uuid = { version = "1", features = ["v4"] }
image = "0.25"
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls-webpki-roots"] }
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls-webpki-roots", "stream"] }
sha2 = "0.10"
aes-gcm = "0.10"
async-trait = "0.1"
@@ -34,6 +39,7 @@ hmac = "0.12"
hex = "0.4"
chrono = "0.4"
urlencoding = "2"
rust-embed = "8"
[target.'cfg(unix)'.dependencies]
libc = "0.2"
@@ -41,6 +47,9 @@ libc = "0.2"
[target.'cfg(windows)'.dependencies]
windows-sys = { version = "0.52", features = ["Win32_System_Threading", "Win32_Foundation"] }
[dev-dependencies]
tempfile = "3"
[build-dependencies]
serde = { version = "1.0", features = ["derive"] }
serde_json = "1.0"
+17
View File
@@ -3,7 +3,24 @@ use std::env;
use std::fs;
use std::path::Path;
/// Ensure `packages/dashboard/out/` exists so `rust-embed` doesn't fail during
/// Rust-only dev builds where the dashboard hasn't been built. The placeholder
/// `index.html` is only written when the directory is completely absent.
fn ensure_dashboard_dir() {
let dashboard_out = Path::new("../packages/dashboard/out");
println!("cargo:rerun-if-changed=../packages/dashboard/out");
if !dashboard_out.join("index.html").exists() {
let _ = fs::create_dir_all(dashboard_out);
let _ = fs::write(
dashboard_out.join("index.html"),
"<!DOCTYPE html><html><body><p>Dashboard not built. Run: cd packages/dashboard &amp;&amp; pnpm build</p></body></html>\n",
);
}
}
fn main() {
ensure_dashboard_dir();
let protocol_dir = Path::new("cdp-protocol");
let out_dir = env::var("OUT_DIR").unwrap();
let out_path = Path::new(&out_dir).join("cdp_generated.rs");
+503
View File
@@ -0,0 +1,503 @@
use std::io::Write as _;
use std::process::exit;
use serde_json::{json, Value};
use crate::color;
use crate::flags::Flags;
use crate::native::stream::chat;
const DEFAULT_MODEL: &str = "anthropic/claude-sonnet-4.6";
#[derive(Clone, Copy, PartialEq)]
enum Verbosity {
Quiet,
Normal,
Verbose,
}
pub fn run_chat(flags: &Flags, message: Option<String>) {
if !chat::is_chat_enabled() {
if flags.json {
println!(
"{}",
json!({"success": false, "error": "AI_GATEWAY_API_KEY not set. Set the AI_GATEWAY_API_KEY environment variable to enable chat."})
);
} else {
eprintln!(
"{} AI_GATEWAY_API_KEY not set. Set the AI_GATEWAY_API_KEY environment variable to enable chat.",
color::error_indicator()
);
}
exit(1);
}
let verbosity = if flags.quiet {
Verbosity::Quiet
} else if flags.verbose {
Verbosity::Verbose
} else {
Verbosity::Normal
};
let model = flags
.model
.clone()
.unwrap_or_else(|| DEFAULT_MODEL.to_string());
let rt = tokio::runtime::Runtime::new().expect("Failed to create tokio runtime");
let is_tty = std::io::IsTerminal::is_terminal(&std::io::stdin());
match message {
Some(msg) => {
rt.block_on(run_single_turn(
&flags.session,
&model,
&msg,
verbosity,
flags.json,
));
}
None if !is_tty => {
let mut input = String::new();
if let Err(e) = std::io::stdin().read_line(&mut input) {
if flags.json {
println!(
"{}",
json!({"success": false, "error": format!("Failed to read stdin: {}", e)})
);
} else {
eprintln!("{} Failed to read stdin: {}", color::error_indicator(), e);
}
exit(1);
}
let input = input.trim();
if input.is_empty() {
if flags.json {
println!(
"{}",
json!({"success": false, "error": "No input provided"})
);
} else {
eprintln!("{} No input provided", color::error_indicator());
}
exit(1);
}
rt.block_on(run_single_turn(
&flags.session,
&model,
input,
verbosity,
flags.json,
));
}
None => {
rt.block_on(run_interactive(
&flags.session,
&model,
verbosity,
flags.json,
));
}
}
}
async fn run_single_turn(
session: &str,
model: &str,
message: &str,
verbosity: Verbosity,
json_mode: bool,
) {
let mut openai_messages: Vec<Value> =
vec![json!({"role": "system", "content": chat::get_system_prompt()})];
openai_messages.push(json!({"role": "user", "content": message}));
let result = run_chat_turn(session, model, &mut openai_messages, verbosity, json_mode).await;
if !result {
exit(1);
}
}
async fn run_interactive(session: &str, model: &str, verbosity: Verbosity, json_mode: bool) {
let mut openai_messages: Vec<Value> =
vec![json!({"role": "system", "content": chat::get_system_prompt()})];
let gateway_url = std::env::var("AI_GATEWAY_URL")
.unwrap_or_else(|_| chat::DEFAULT_AI_GATEWAY_URL.to_string())
.trim_end_matches('/')
.to_string();
let api_key = std::env::var("AI_GATEWAY_API_KEY").unwrap_or_default();
let url = format!("{}/v1/chat/completions", gateway_url);
let client = chat::http_client();
loop {
if !json_mode {
eprint!("{} ", color::cyan(">"));
let _ = std::io::stderr().flush();
}
let mut input = String::new();
match std::io::stdin().read_line(&mut input) {
Ok(0) => break,
Err(_) => break,
Ok(_) => {}
}
let input = input.trim();
if input.is_empty() {
continue;
}
if matches!(input, "quit" | "exit" | "q") {
break;
}
openai_messages.push(json!({"role": "user", "content": input}));
// Compaction check
let total_chars = chat::estimate_chars(&openai_messages);
if total_chars > chat::COMPACT_THRESHOLD_CHARS
&& openai_messages.len() > chat::KEEP_RECENT_MESSAGES + 2
{
let split = chat::find_safe_split(&openai_messages, chat::KEEP_RECENT_MESSAGES);
let to_summarize = &openai_messages[1..split];
if let Some(summary) =
chat::summarize_for_compaction(client, &url, &api_key, model, to_summarize).await
{
let summary_msg = json!({
"role": "system",
"content": format!("[Conversation summary]\n{}", summary)
});
let recent = openai_messages[split..].to_vec();
openai_messages = vec![openai_messages[0].clone(), summary_msg];
openai_messages.extend(recent);
}
}
let success =
run_chat_turn(session, model, &mut openai_messages, verbosity, json_mode).await;
if !success && !json_mode {
// Continue the loop on error; don't exit interactive mode
}
if !json_mode {
eprintln!();
}
}
}
/// Runs one chat turn: sends messages to the gateway, streams text/tool calls,
/// executes tools in a loop until the model is done. Appends assistant and tool
/// messages to `openai_messages`. Returns true on success.
async fn run_chat_turn(
session: &str,
model: &str,
openai_messages: &mut Vec<Value>,
verbosity: Verbosity,
json_mode: bool,
) -> bool {
let gateway_url = std::env::var("AI_GATEWAY_URL")
.unwrap_or_else(|_| chat::DEFAULT_AI_GATEWAY_URL.to_string())
.trim_end_matches('/')
.to_string();
let api_key = match std::env::var("AI_GATEWAY_API_KEY") {
Ok(k) => k,
Err(_) => {
if json_mode {
println!(
"{}",
json!({"success": false, "error": "AI_GATEWAY_API_KEY not set"})
);
} else {
eprintln!("{} AI_GATEWAY_API_KEY not set", color::error_indicator());
}
return false;
}
};
let tools: Value = serde_json::from_str(chat::CHAT_TOOLS).unwrap();
let url = format!("{}/v1/chat/completions", gateway_url);
let client = chat::http_client();
let total_deadline = tokio::time::Instant::now() + std::time::Duration::from_secs(300);
let tool_timeout = std::time::Duration::from_secs(60);
let mut all_text = String::new();
let mut all_tool_calls: Vec<Value> = Vec::new();
let mut had_text = false;
for _step in 0..50 {
if tokio::time::Instant::now() >= total_deadline {
if json_mode {
println!(
"{}",
json!({"success": false, "error": "Chat session timed out (5 minute limit)."})
);
} else {
eprintln!(
"\n{} Chat session timed out (5 minute limit).",
color::error_indicator()
);
}
return false;
}
let gateway_body = json!({
"model": model,
"messages": openai_messages,
"tools": tools,
"stream": true,
});
let gw_response = match client
.post(&url)
.header("Authorization", format!("Bearer {}", api_key))
.header("Content-Type", "application/json")
.body(gateway_body.to_string())
.send()
.await
{
Ok(r) => r,
Err(e) => {
if json_mode {
println!(
"{}",
json!({"success": false, "error": format!("Gateway request failed: {}", e)})
);
} else {
eprintln!(
"\n{} Gateway request failed: {}",
color::error_indicator(),
e
);
}
return false;
}
};
if !gw_response.status().is_success() {
let body_text = gw_response.text().await.unwrap_or_default();
if json_mode {
println!("{}", json!({"success": false, "error": body_text}));
} else {
eprintln!("\n{} {}", color::error_indicator(), body_text);
}
return false;
}
let (text_chunks, tool_calls) =
parse_gateway_stream(gw_response, verbosity, json_mode).await;
if !text_chunks.is_empty() {
let text = text_chunks.join("");
all_text.push_str(&text);
if !json_mode {
if !had_text && verbosity != Verbosity::Quiet {
// Add blank line before text if we showed tool calls
if !all_tool_calls.is_empty() {
println!();
}
}
had_text = true;
}
let mut content = json!(text);
if let Some(last) = openai_messages.last() {
if last.get("role").and_then(|r| r.as_str()) == Some("assistant")
&& last.get("tool_calls").is_some()
{
content = json!(text);
}
}
openai_messages.push(json!({"role": "assistant", "content": content}));
}
if tool_calls.is_empty() {
break;
}
let tc_values: Vec<Value> = tool_calls
.iter()
.map(|(id, name, args)| {
json!({"id": id, "type": "function", "function": {"name": name, "arguments": args}})
})
.collect();
if text_chunks.is_empty() {
openai_messages.push(json!({"role": "assistant", "tool_calls": tc_values}));
} else {
// If we had both text and tool calls in the same response, merge them
if let Some(last) = openai_messages.last_mut() {
if last.get("role").and_then(|r| r.as_str()) == Some("assistant")
&& last.get("tool_calls").is_none()
{
last["tool_calls"] = json!(tc_values);
} else {
openai_messages.push(json!({"role": "assistant", "tool_calls": tc_values}));
}
}
}
for (tc_id, _tc_name, tc_args) in &tool_calls {
let input: Value = serde_json::from_str(tc_args).unwrap_or(json!({}));
let command = input.get("command").and_then(|c| c.as_str()).unwrap_or("");
if !json_mode && verbosity != Verbosity::Quiet {
eprintln!("{}", color::dim(&format!("> {}", command)));
}
let result =
match tokio::time::timeout(tool_timeout, chat::execute_chat_tool(session, command))
.await
{
Ok(r) => r,
Err(_) => "Tool execution timed out after 60 seconds.".to_string(),
};
if !json_mode && verbosity == Verbosity::Verbose {
for line in result.lines() {
eprintln!(" {}", color::dim(line));
}
}
all_tool_calls.push(json!({
"command": command,
"output": result
}));
openai_messages.push(json!({
"role": "tool",
"tool_call_id": tc_id,
"content": result
}));
}
}
if json_mode {
println!(
"{}",
json!({
"success": true,
"text": all_text,
"tool_calls": all_tool_calls
})
);
} else if !had_text && !json_mode {
// Model returned only tool calls with no final text; print newline for clean output
println!();
}
true
}
/// Parses the SSE stream from the AI gateway, printing text deltas to stdout in
/// real-time. Returns (collected_text_chunks, tool_calls).
async fn parse_gateway_stream(
gw_response: reqwest::Response,
verbosity: Verbosity,
json_mode: bool,
) -> (Vec<String>, Vec<(String, String, String)>) {
use futures_util::StreamExt as _;
let mut text_chunks: Vec<String> = Vec::new();
let mut tool_call_args: std::collections::HashMap<usize, (String, String, String)> =
std::collections::HashMap::new();
let mut byte_stream = gw_response.bytes_stream();
let mut buffer = String::new();
while let Some(chunk_result) = byte_stream.next().await {
let chunk = match chunk_result {
Ok(c) => c,
Err(_) => break,
};
buffer.push_str(&String::from_utf8_lossy(&chunk));
while let Some(newline_pos) = buffer.find('\n') {
let line = buffer[..newline_pos].trim_end_matches('\r').to_string();
buffer = buffer[newline_pos + 1..].to_string();
if line.is_empty() {
continue;
}
let Some(data) = line.strip_prefix("data: ") else {
continue;
};
if data == "[DONE]" {
let tool_calls = collect_tool_calls(&mut tool_call_args);
if !json_mode && !text_chunks.is_empty() {
// End the streamed text line
let _ = std::io::stdout().flush();
}
return (text_chunks, tool_calls);
}
let Ok(sse_json) = serde_json::from_str::<Value>(data) else {
continue;
};
let delta = sse_json
.get("choices")
.and_then(|c| c.get(0))
.and_then(|c| c.get("delta"));
let Some(delta) = delta else { continue };
if let Some(text) = delta.get("content").and_then(|c| c.as_str()) {
if !text.is_empty() {
text_chunks.push(text.to_string());
if !json_mode && verbosity != Verbosity::Quiet {
print!("{}", text);
let _ = std::io::stdout().flush();
}
}
}
if let Some(tcs) = delta.get("tool_calls").and_then(|t| t.as_array()) {
for tc in tcs {
let idx = tc.get("index").and_then(|i| i.as_u64()).unwrap_or(0) as usize;
if let std::collections::hash_map::Entry::Vacant(e) = tool_call_args.entry(idx)
{
let id = tc
.get("id")
.and_then(|i| i.as_str())
.unwrap_or("")
.to_string();
let name = tc
.get("function")
.and_then(|f| f.get("name"))
.and_then(|n| n.as_str())
.unwrap_or("")
.to_string();
e.insert((id, name, String::new()));
}
if let Some(arg_delta) = tc
.get("function")
.and_then(|f| f.get("arguments"))
.and_then(|a| a.as_str())
{
let entry = tool_call_args.get_mut(&idx).unwrap();
entry.2.push_str(arg_delta);
}
}
}
}
}
if !json_mode && !text_chunks.is_empty() {
let _ = std::io::stdout().flush();
}
let tool_calls = collect_tool_calls(&mut tool_call_args);
(text_chunks, tool_calls)
}
fn collect_tool_calls(
map: &mut std::collections::HashMap<usize, (String, String, String)>,
) -> Vec<(String, String, String)> {
let mut indices: Vec<usize> = map.keys().copied().collect();
indices.sort();
indices
.into_iter()
.filter_map(|idx| map.remove(&idx))
.collect()
}
+20 -5
View File
@@ -1,15 +1,30 @@
//! Color output utilities respecting NO_COLOR environment variable.
//! Color output utilities.
//!
//! When the NO_COLOR environment variable is present (regardless of value),
//! all color formatting is disabled per https://no-color.org/
//! Colors are off by default (agent-friendly). Enable with
//! `AGENT_BROWSER_COLOR=1`. Setting `NO_COLOR` to any value disables
//! colors per <https://no-color.org/>.
use std::env;
use std::sync::OnceLock;
/// Returns true if color output is enabled (NO_COLOR is NOT set)
fn env_is_truthy(name: &str) -> Option<bool> {
env::var(name)
.ok()
.map(|val| !matches!(val.to_lowercase().as_str(), "0" | "false" | "no"))
}
/// Returns true if color output is enabled.
///
/// Priority: `NO_COLOR` (presence disables, per spec) >
/// `AGENT_BROWSER_COLOR` (truthy enables) > default (off).
pub fn is_enabled() -> bool {
static COLORS_ENABLED: OnceLock<bool> = OnceLock::new();
*COLORS_ENABLED.get_or_init(|| env::var("NO_COLOR").is_err())
*COLORS_ENABLED.get_or_init(|| {
if env::var_os("NO_COLOR").is_some() {
return false;
}
env_is_truthy("AGENT_BROWSER_COLOR").unwrap_or(false)
})
}
/// Format text in red (errors)
+1022 -36
View File
File diff suppressed because it is too large Load Diff
+410 -4
View File
@@ -12,6 +12,11 @@ use std::time::Duration;
#[cfg(unix)]
use std::os::unix::net::UnixStream;
#[cfg(windows)]
use windows_sys::Win32::Foundation::CloseHandle;
#[cfg(windows)]
use windows_sys::Win32::System::Threading::{OpenProcess, PROCESS_QUERY_LIMITED_INFORMATION};
#[derive(Serialize)]
#[allow(dead_code)]
pub struct Request {
@@ -118,12 +123,31 @@ fn get_pid_path(session: &str) -> PathBuf {
get_socket_dir().join(format!("{}.pid", session))
}
fn get_version_path(session: &str) -> PathBuf {
get_socket_dir().join(format!("{}.version", session))
}
/// Path to the sidecar file that records the URL the previous daemon was on,
/// used to restore navigation after a version-mismatch restart. Only written
/// when the version-mismatch branch fires; cleared after the new daemon
/// reads it. Manual `close` does not write this file, so a clean shutdown
/// won't trigger surprise navigation.
pub fn get_restore_url_path(session: &str) -> PathBuf {
get_socket_dir().join(format!("{}.restore-url", session))
}
/// Clean up stale socket and PID files for a session
fn cleanup_stale_files(session: &str) {
pub fn cleanup_stale_files(session: &str) {
let pid_path = get_pid_path(session);
let _ = fs::remove_file(&pid_path);
let version_path = get_version_path(session);
let _ = fs::remove_file(&version_path);
let stream_path = get_socket_dir().join(format!("{}.stream", session));
let _ = fs::remove_file(&stream_path);
// Note: the .restore-url sidecar is intentionally NOT removed here —
// it lives across the brief window between killing the old daemon
// and the new daemon reading it back. The new daemon deletes it after
// restoring (see actions::auto_launch).
#[cfg(unix)]
{
@@ -138,6 +162,186 @@ fn cleanup_stale_files(session: &str) {
}
}
/// Returns whether a process with the given PID is currently alive.
///
/// On unix, EPERM (process exists but we can't signal it) counts as alive
/// so we don't mis-clean a live daemon owned by a different uid. Only ESRCH
/// ("no such process") is treated as dead.
pub fn is_pid_alive(pid: u32) -> bool {
#[cfg(unix)]
unsafe {
if libc::kill(pid as i32, 0) == 0 {
return true;
}
std::io::Error::last_os_error().raw_os_error() != Some(libc::ESRCH)
}
#[cfg(windows)]
unsafe {
let handle = OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION, 0, pid);
if handle != 0 {
CloseHandle(handle);
true
} else {
false
}
}
}
/// A currently-running daemon session discovered by [`walk_daemons`].
#[derive(Debug, Clone)]
pub struct ActiveSession {
pub name: String,
pub pid: u32,
/// Contents of the session's `.version` file if present and non-empty.
pub version: Option<String>,
}
/// Why a session's sidecar files were cleaned up during a walk.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum CleanReason {
/// The `.pid` file referenced a process that no longer exists.
ProcessGone,
/// The `.pid` file could not be parsed as a PID.
UnreadablePidFile,
/// A `.sock` file had no corresponding `.pid` file (unix only).
OrphanedSocket,
/// The `dashboard.pid` referenced a process that no longer exists.
DashboardGone,
}
/// A session whose sidecar files were removed as a side effect of a walk.
#[derive(Debug, Clone)]
pub struct CleanedSession {
pub name: String,
pub reason: CleanReason,
}
/// Information about the standalone dashboard process, if any.
#[derive(Debug, Clone, Copy)]
pub struct DashboardInfo {
pub pid: u32,
pub alive: bool,
}
/// Snapshot of daemon state under [`get_socket_dir()`] after a walk. Stale
/// sidecar files are cleaned up as a side effect and recorded in `cleaned`.
#[derive(Debug, Default)]
pub struct DaemonInventory {
pub sessions: Vec<ActiveSession>,
pub cleaned: Vec<CleanedSession>,
pub dashboard: Option<DashboardInfo>,
}
/// Read the session's `.version` sidecar if present and non-empty.
pub fn read_session_version(session: &str) -> Option<String> {
let path = get_socket_dir().join(format!("{}.version", session));
fs::read_to_string(&path)
.ok()
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty())
}
/// Walk the socket directory and classify each `.pid` / `.sock` entry.
///
/// - Live daemons go into `sessions` with their `.version` file contents.
/// - Stale entries (process gone, unreadable pid, orphaned `.sock`) are
/// cleaned via [`cleanup_stale_files`] and recorded in `cleaned`.
/// - `dashboard.pid` lands in `dashboard` with liveness info; if the
/// process is gone, the pid file is removed and a `DashboardGone` entry
/// is added to `cleaned`.
///
/// If the socket directory doesn't exist, returns an empty inventory with
/// no side effects.
pub fn walk_daemons() -> DaemonInventory {
let socket_dir = get_socket_dir();
let mut inventory = DaemonInventory::default();
let entries = match fs::read_dir(&socket_dir) {
Ok(e) => e,
Err(_) => return inventory,
};
for entry in entries.flatten() {
let name = entry.file_name().to_string_lossy().to_string();
if name == "dashboard.pid" {
if let Ok(s) = fs::read_to_string(entry.path()) {
if let Ok(pid) = s.trim().parse::<u32>() {
let alive = is_pid_alive(pid);
inventory.dashboard = Some(DashboardInfo { pid, alive });
if !alive {
let _ = fs::remove_file(entry.path());
inventory.cleaned.push(CleanedSession {
name: "dashboard".to_string(),
reason: CleanReason::DashboardGone,
});
}
}
}
continue;
}
let session_name = match name.strip_suffix(".pid") {
Some(s) if !s.is_empty() => s.to_string(),
_ => continue,
};
let pid = match fs::read_to_string(entry.path())
.ok()
.and_then(|s| s.trim().parse::<u32>().ok())
{
Some(p) => p,
None => {
cleanup_stale_files(&session_name);
inventory.cleaned.push(CleanedSession {
name: session_name,
reason: CleanReason::UnreadablePidFile,
});
continue;
}
};
if !is_pid_alive(pid) {
cleanup_stale_files(&session_name);
inventory.cleaned.push(CleanedSession {
name: session_name,
reason: CleanReason::ProcessGone,
});
continue;
}
let version = read_session_version(&session_name);
inventory.sessions.push(ActiveSession {
name: session_name,
pid,
version,
});
}
// Orphaned .sock files without a corresponding .pid (unix only).
#[cfg(unix)]
if let Ok(entries) = fs::read_dir(&socket_dir) {
for entry in entries.flatten() {
let name = entry.file_name().to_string_lossy().to_string();
if let Some(session_name) = name.strip_suffix(".sock") {
if session_name.is_empty() {
continue;
}
let pid_path = socket_dir.join(format!("{}.pid", session_name));
if !pid_path.exists() {
cleanup_stale_files(session_name);
inventory.cleaned.push(CleanedSession {
name: session_name.to_string(),
reason: CleanReason::OrphanedSocket,
});
}
}
}
}
inventory
}
#[cfg(windows)]
fn get_port_path(session: &str) -> PathBuf {
get_socket_dir().join(format!("{}.port", session))
@@ -198,6 +402,8 @@ pub struct DaemonOptions<'a> {
pub debug: bool,
pub executable_path: Option<&'a str>,
pub extensions: &'a [String],
pub init_scripts: &'a [String],
pub enable: &'a [String],
pub args: Option<&'a str>,
pub user_agent: Option<&'a str>,
pub proxy: Option<&'a str>,
@@ -206,6 +412,7 @@ pub struct DaemonOptions<'a> {
pub proxy_password: Option<&'a str>,
pub ignore_https_errors: bool,
pub allow_file_access: bool,
pub hide_scrollbars: bool,
pub profile: Option<&'a str>,
pub state: Option<&'a str>,
pub provider: Option<&'a str>,
@@ -217,7 +424,9 @@ pub struct DaemonOptions<'a> {
pub confirm_actions: Option<&'a str>,
pub engine: Option<&'a str>,
pub auto_connect: bool,
pub force_launch: bool,
pub idle_timeout: Option<&'a str>,
pub default_timeout: Option<u64>,
pub cdp: Option<&'a str>,
pub no_auto_dialog: bool,
}
@@ -238,6 +447,12 @@ fn apply_daemon_env(cmd: &mut Command, session: &str, opts: &DaemonOptions) {
if !opts.extensions.is_empty() {
cmd.env("AGENT_BROWSER_EXTENSIONS", opts.extensions.join(","));
}
if !opts.init_scripts.is_empty() {
cmd.env("AGENT_BROWSER_INIT_SCRIPTS", opts.init_scripts.join(","));
}
if !opts.enable.is_empty() {
cmd.env("AGENT_BROWSER_ENABLE", opts.enable.join(","));
}
if let Some(a) = opts.args {
cmd.env("AGENT_BROWSER_ARGS", a);
}
@@ -262,6 +477,10 @@ fn apply_daemon_env(cmd: &mut Command, session: &str, opts: &DaemonOptions) {
if opts.allow_file_access {
cmd.env("AGENT_BROWSER_ALLOW_FILE_ACCESS", "1");
}
cmd.env(
"AGENT_BROWSER_HIDE_SCROLLBARS",
if opts.hide_scrollbars { "1" } else { "0" },
);
if let Some(prof) = opts.profile {
cmd.env("AGENT_BROWSER_PROFILE", prof);
}
@@ -295,9 +514,15 @@ fn apply_daemon_env(cmd: &mut Command, session: &str, opts: &DaemonOptions) {
if opts.auto_connect {
cmd.env("AGENT_BROWSER_AUTO_CONNECT", "1");
}
if opts.force_launch {
cmd.env("AGENT_BROWSER_FORCE_LAUNCH", "1");
}
if let Some(idle) = opts.idle_timeout {
cmd.env("AGENT_BROWSER_IDLE_TIMEOUT_MS", idle);
}
if let Some(timeout) = opts.default_timeout {
cmd.env("AGENT_BROWSER_DEFAULT_TIMEOUT", timeout.to_string());
}
if let Some(cdp) = opts.cdp {
cmd.env("AGENT_BROWSER_CDP", cdp);
}
@@ -306,6 +531,86 @@ fn apply_daemon_env(cmd: &mut Command, session: &str, opts: &DaemonOptions) {
}
}
/// Check if the running daemon's version matches this CLI binary.
/// Returns false when the version file is missing — an unversioned daemon
/// is most likely a stale leftover from before version tracking was added
/// (or from the Node.js era), and silently reusing it is the exact bug
/// this check exists to prevent. The one-time cost of an unnecessary
/// restart on the first upgrade is preferable to silent failures.
fn daemon_version_matches(session: &str) -> bool {
let version_path = get_version_path(session);
match fs::read_to_string(&version_path) {
Ok(v) => v.trim() == env!("CARGO_PKG_VERSION"),
Err(_) => false,
}
}
/// One-shot socket query for the running daemon's current URL.
/// Returns None on any kind of failure — caller must treat as best-effort.
fn query_current_url(session: &str) -> Option<String> {
let cmd = serde_json::json!({
"id": format!("restore-url-probe-{}", std::process::id()),
"action": "url",
});
let resp = send_command_once(&cmd, session).ok()?;
if !resp.success {
return None;
}
resp.data
.as_ref()
.and_then(|d| d.get("url"))
.and_then(|v| v.as_str())
.map(|s| s.to_string())
}
/// Kill a running daemon by reading its PID file and sending a kill signal.
fn kill_stale_daemon(session: &str) {
// Remove the socket first so no new connections reach the old daemon
#[cfg(unix)]
{
let socket_path = get_socket_path(session);
let _ = fs::remove_file(&socket_path);
}
let pid_path = get_pid_path(session);
if let Ok(pid_str) = fs::read_to_string(&pid_path) {
if let Ok(pid) = pid_str.trim().parse::<u32>() {
#[cfg(unix)]
{
unsafe {
libc::kill(pid as i32, libc::SIGTERM);
}
// Wait up to 1s for graceful shutdown, then force-kill
for _ in 0..10 {
thread::sleep(Duration::from_millis(100));
if unsafe { libc::kill(pid as i32, 0) } != 0 {
break;
}
}
// Force-kill if still alive
if unsafe { libc::kill(pid as i32, 0) } == 0 {
unsafe {
libc::kill(pid as i32, libc::SIGKILL);
}
thread::sleep(Duration::from_millis(100));
}
}
#[cfg(windows)]
{
let _ = Command::new("taskkill")
.args(["/PID", &pid.to_string(), "/F"])
.stdout(Stdio::null())
.stderr(Stdio::null())
.status();
thread::sleep(Duration::from_millis(500));
}
}
}
// Clean up leftover files regardless
cleanup_stale_files(session);
}
pub fn ensure_daemon(session: &str, opts: &DaemonOptions) -> Result<DaemonResult, String> {
// Socket connectivity is the sole liveness check — no PID check — so
// callers in a different PID namespace (e.g. unshare) can still reuse
@@ -316,9 +621,30 @@ pub fn ensure_daemon(session: &str, opts: &DaemonOptions) -> Result<DaemonResult
// (daemon has a 100ms shutdown delay, so we wait longer)
thread::sleep(Duration::from_millis(150));
if daemon_ready(session) {
return Ok(DaemonResult {
already_running: true,
});
// Check version: if the running daemon is from a different CLI
// version (e.g. after an upgrade), kill it and start a fresh one.
if !daemon_version_matches(session) {
eprintln!(
"{} Daemon version mismatch detected, restarting...",
crate::color::warning_indicator()
);
// Best-effort: ask the old daemon for its current URL so the
// new daemon can restore navigation after auto-connect. If the
// query fails (already shutting down, no browser, etc.) we
// silently skip — the user just sees about:blank as before.
if let Some(url) = query_current_url(session) {
if !url.is_empty() && url != "about:blank" {
let path = get_restore_url_path(session);
let _ = fs::write(&path, &url);
}
}
kill_stale_daemon(session);
// Fall through to spawn a new daemon below
} else {
return Ok(DaemonResult {
already_running: true,
});
}
}
}
@@ -429,6 +755,21 @@ pub fn ensure_daemon(session: &str, opts: &DaemonOptions) -> Result<DaemonResult
let _ = stderr.read_to_string(&mut stderr_output);
}
let stderr_trimmed = stderr_output.trim();
// If the daemon failed because another instance won the bind
// race ("Address already in use"), check whether that winner is
// now accepting connections and piggyback on it.
if stderr_trimmed.contains("Address already in use")
|| stderr_trimmed.contains("Failed to bind")
{
thread::sleep(Duration::from_millis(200));
if daemon_ready(session) {
return Ok(DaemonResult {
already_running: true,
});
}
}
if !stderr_trimmed.is_empty() {
let msg = if stderr_trimmed.len() > 500 {
let mut end = 500;
@@ -739,4 +1080,69 @@ mod tests {
assert_eq!(get_port_for_session("work"), 51184);
assert_eq!(get_port_for_session(""), 49152);
}
// === Daemon Version Mismatch Detection Tests ===
#[test]
fn test_daemon_version_matches_same_version() {
let dir = std::env::temp_dir().join("ab-test-version-match");
let _ = fs::create_dir_all(&dir);
let _guard = EnvGuard::new(&["AGENT_BROWSER_SOCKET_DIR", "XDG_RUNTIME_DIR"]);
_guard.set("AGENT_BROWSER_SOCKET_DIR", dir.to_str().unwrap());
let version_path = dir.join("test-session.version");
let _ = fs::write(&version_path, env!("CARGO_PKG_VERSION"));
assert!(daemon_version_matches("test-session"));
let _ = fs::remove_file(&version_path);
let _ = fs::remove_dir(&dir);
}
#[test]
fn test_daemon_version_matches_different_version() {
let dir = std::env::temp_dir().join("ab-test-version-mismatch");
let _ = fs::create_dir_all(&dir);
let _guard = EnvGuard::new(&["AGENT_BROWSER_SOCKET_DIR", "XDG_RUNTIME_DIR"]);
_guard.set("AGENT_BROWSER_SOCKET_DIR", dir.to_str().unwrap());
let version_path = dir.join("test-session.version");
let _ = fs::write(&version_path, "0.0.0-old");
assert!(!daemon_version_matches("test-session"));
let _ = fs::remove_file(&version_path);
let _ = fs::remove_dir(&dir);
}
#[test]
fn test_daemon_version_matches_no_file() {
let dir = std::env::temp_dir().join("ab-test-version-nofile");
let _ = fs::create_dir_all(&dir);
let _guard = EnvGuard::new(&["AGENT_BROWSER_SOCKET_DIR", "XDG_RUNTIME_DIR"]);
_guard.set("AGENT_BROWSER_SOCKET_DIR", dir.to_str().unwrap());
// No version file: treated as mismatch so stale pre-version-tracking
// daemons (including Node.js era) are always restarted.
assert!(!daemon_version_matches("test-session"));
let _ = fs::remove_dir(&dir);
}
#[test]
fn test_cleanup_stale_files_removes_version() {
let dir = std::env::temp_dir().join("ab-test-cleanup-version");
let _ = fs::create_dir_all(&dir);
let _guard = EnvGuard::new(&["AGENT_BROWSER_SOCKET_DIR", "XDG_RUNTIME_DIR"]);
_guard.set("AGENT_BROWSER_SOCKET_DIR", dir.to_str().unwrap());
let version_path = dir.join("test-session.version");
let _ = fs::write(&version_path, "0.1.0");
assert!(version_path.exists());
cleanup_stale_files("test-session");
assert!(!version_path.exists());
let _ = fs::remove_dir(&dir);
}
}
+156
View File
@@ -0,0 +1,156 @@
//! Check the Chrome install: binary path, version, cache dirs, user-data
//! dir, and the optional lightpanda engine.
use std::env;
use std::path::{Path, PathBuf};
use super::helpers::which_exists;
use super::{Check, Status};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Chrome";
let chrome = crate::native::cdp::chrome::find_chrome();
match chrome {
Some(path) => {
let label = path.display().to_string();
match query_chrome_version(&path) {
Some(version) => checks.push(Check::new(
"chrome.installed",
category,
Status::Pass,
format!("{} at {}", version, label),
)),
None => checks.push(Check::new(
"chrome.installed",
category,
Status::Pass,
format!("Chrome at {} (version unknown)", label),
)),
}
}
None => checks.push(
Check::new(
"chrome.installed",
category,
Status::Fail,
"No Chrome binary found",
)
.with_fix("agent-browser install"),
),
}
let cache_dir = crate::install::get_browsers_dir();
if cache_dir.exists() {
checks.push(Check::new(
"chrome.cache_dir",
category,
Status::Info,
format!("Cache dir {}", cache_dir.display()),
));
}
if let Some(puppeteer_dir) = puppeteer_cache_dir() {
if puppeteer_dir.exists() {
checks.push(Check::new(
"chrome.puppeteer_cache",
category,
Status::Info,
format!(
"Puppeteer cache also present: {} (will be used as a fallback)",
puppeteer_dir.display()
),
));
}
}
if let Some(user_data_dir) = crate::native::cdp::chrome::find_chrome_user_data_dir() {
let profiles = crate::native::cdp::chrome::list_chrome_profiles(&user_data_dir);
let count = profiles.len();
let dir_label = user_data_dir.display().to_string();
if count == 0 {
checks.push(Check::new(
"chrome.user_data_dir",
category,
Status::Info,
format!(
"Chrome user data dir found ({}), no profiles parsed",
dir_label
),
));
} else {
checks.push(Check::new(
"chrome.user_data_dir",
category,
Status::Info,
format!("{} Chrome profile(s) at {}", count, dir_label),
));
}
}
if let Ok(engine) = env::var("AGENT_BROWSER_ENGINE") {
if engine == "lightpanda" {
// Best-effort PATH lookup; absence is FAIL only when the user
// explicitly opted into the lightpanda engine.
if which_exists("lightpanda") {
checks.push(Check::new(
"chrome.engine_lightpanda",
category,
Status::Pass,
"Lightpanda binary on PATH",
));
} else {
checks.push(
Check::new(
"chrome.engine_lightpanda",
category,
Status::Fail,
"AGENT_BROWSER_ENGINE=lightpanda but no lightpanda binary on PATH",
)
.with_fix("install lightpanda or unset AGENT_BROWSER_ENGINE"),
);
}
}
}
}
fn query_chrome_version(path: &Path) -> Option<String> {
let output = std::process::Command::new(path)
.arg("--version")
.output()
.ok()?;
if !output.status.success() {
return None;
}
let s = String::from_utf8_lossy(&output.stdout).trim().to_string();
if s.is_empty() {
None
} else {
Some(s)
}
}
pub(super) fn puppeteer_cache_dir() -> Option<PathBuf> {
if let Ok(p) = env::var("PUPPETEER_CACHE_DIR") {
return Some(PathBuf::from(p));
}
dirs::home_dir().map(|h| h.join(".cache").join("puppeteer"))
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_puppeteer_cache_dir_returns_sensible_default() {
// When PUPPETEER_CACHE_DIR is unset, we fall back to
// ~/.cache/puppeteer. Mutating env vars here would race with other
// tests, so just verify the fallback path is shaped correctly.
if env::var("PUPPETEER_CACHE_DIR").is_err() {
let dir = puppeteer_cache_dir().expect("home dir should resolve in tests");
let s = dir.to_string_lossy();
assert!(s.contains(".cache"));
assert!(s.ends_with("puppeteer"));
}
}
}
+90
View File
@@ -0,0 +1,90 @@
//! Check user config files: `~/.agent-browser/config.json`,
//! `./agent-browser.json`, and any file referenced by
//! `AGENT_BROWSER_CONFIG`.
use std::env;
use std::path::PathBuf;
use super::helpers::parse_json_file;
use super::{Check, Status};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Config";
let user_path = dirs::home_dir().map(|d| d.join(".agent-browser").join("config.json"));
if let Some(p) = user_path {
if p.exists() {
match parse_json_file(&p) {
Ok(_) => checks.push(Check::new(
"config.user",
category,
Status::Pass,
format!("{} (valid JSON)", p.display()),
)),
Err(e) => checks.push(
Check::new(
"config.user",
category,
Status::Fail,
format!("{}: {}", p.display(), e),
)
.with_fix(format!("edit {}", p.display())),
),
}
}
}
let project_path = PathBuf::from("agent-browser.json");
if project_path.exists() {
match parse_json_file(&project_path) {
Ok(_) => checks.push(Check::new(
"config.project",
category,
Status::Pass,
format!("{} (valid JSON)", project_path.display()),
)),
Err(e) => checks.push(
Check::new(
"config.project",
category,
Status::Fail,
format!("{}: {}", project_path.display(), e),
)
.with_fix(format!("edit {}", project_path.display())),
),
}
}
if let Ok(custom) = env::var("AGENT_BROWSER_CONFIG") {
let p = PathBuf::from(&custom);
if !p.exists() {
checks.push(
Check::new(
"config.custom",
category,
Status::Fail,
format!("AGENT_BROWSER_CONFIG points to missing file: {}", custom),
)
.with_fix("update or unset AGENT_BROWSER_CONFIG"),
);
} else {
match parse_json_file(&p) {
Ok(_) => checks.push(Check::new(
"config.custom",
category,
Status::Pass,
format!("AGENT_BROWSER_CONFIG: {} (valid JSON)", custom),
)),
Err(e) => checks.push(
Check::new(
"config.custom",
category,
Status::Fail,
format!("AGENT_BROWSER_CONFIG: {}: {}", custom, e),
)
.with_fix(format!("edit {}", custom)),
),
}
}
}
}
+70
View File
@@ -0,0 +1,70 @@
//! Check running daemons: inventory of sessions, version match with the
//! CLI, and stale sidecar files cleaned up as a side effect of the walk.
use super::{Check, Status};
use crate::connection::{walk_daemons, CleanReason};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Daemons";
let cli_version = env!("CARGO_PKG_VERSION");
let inventory = walk_daemons();
for cleaned in &inventory.cleaned {
let reason = match cleaned.reason {
CleanReason::ProcessGone | CleanReason::DashboardGone => "process gone",
CleanReason::UnreadablePidFile => "unreadable pid file",
CleanReason::OrphanedSocket => "orphaned socket",
};
checks.push(Check::new(
format!("daemon.cleaned.{}", cleaned.name),
category,
Status::Warn,
format!("Cleaned stale files: {} ({})", cleaned.name, reason),
));
}
if inventory.sessions.is_empty() {
checks.push(Check::new(
"daemon.active",
category,
Status::Pass,
"No active daemons",
));
} else {
for session in &inventory.sessions {
let version_match = session.version.as_deref() == Some(cli_version);
let status = if version_match {
Status::Pass
} else {
Status::Warn
};
let suffix = if version_match {
String::new()
} else {
format!(" (version mismatch with CLI {})", cli_version)
};
let mut check = Check::new(
format!("daemon.session.{}", session.name),
category,
status,
format!("Session {} (pid {}){}", session.name, session.pid, suffix),
);
if !version_match {
check = check.with_fix(format!("agent-browser --session {} close", session.name));
}
checks.push(check);
}
}
if let Some(dashboard) = inventory.dashboard {
if dashboard.alive {
checks.push(Check::new(
"daemon.dashboard",
category,
Status::Pass,
format!("Dashboard server running (pid {})", dashboard.pid),
));
}
}
}
+140
View File
@@ -0,0 +1,140 @@
//! Check the local environment: CLI version, platform, state/socket dirs,
//! and free disk space.
use std::path::Path;
use super::helpers::{disk_free_bytes, human_size, is_writable_dir};
use super::{Check, Status};
use crate::connection::get_socket_dir;
use crate::native::state::get_state_dir;
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Environment";
let version = env!("CARGO_PKG_VERSION");
let platform = format!("{} {}", std::env::consts::OS, std::env::consts::ARCH);
checks.push(Check::new(
"env.version",
category,
Status::Pass,
format!("CLI version {} ({})", version, platform),
));
match dirs::home_dir() {
Some(home) => checks.push(Check::new(
"env.home",
category,
Status::Pass,
format!("Home directory {}", home.display()),
)),
None => checks.push(Check::new(
"env.home",
category,
Status::Fail,
"Could not determine home directory",
)),
}
let state_dir = get_state_dir();
let socket_dir = get_socket_dir();
// Under the default setup, state and socket dirs are the same
// (~/.agent-browser). Collapse to a single line when they match;
// split when XDG_RUNTIME_DIR or AGENT_BROWSER_SOCKET_DIR diverts
// sockets elsewhere.
if state_dir == socket_dir {
push_dir_check(
checks,
"env.state_dir",
category,
"State and socket directory",
&state_dir,
);
} else {
push_dir_check(
checks,
"env.state_dir",
category,
"State directory",
&state_dir,
);
push_dir_check(
checks,
"env.socket_dir",
category,
"Socket directory",
&socket_dir,
);
}
match disk_free_bytes(&state_dir) {
Some(bytes) => {
let mb = bytes / (1024 * 1024);
let human = human_size(bytes);
if mb < 500 {
checks.push(
Check::new(
"env.disk_free",
category,
Status::Warn,
format!("Low disk space at state dir: {} free", human),
)
.with_fix("free up disk space; Chrome installs require ~500 MB"),
);
} else {
checks.push(Check::new(
"env.disk_free",
category,
Status::Pass,
format!("{} free at state dir", human),
));
}
}
None => checks.push(Check::new(
"env.disk_free",
category,
Status::Info,
"Disk free check unavailable on this platform",
)),
}
}
fn push_dir_check(
checks: &mut Vec<Check>,
id: &'static str,
category: &'static str,
label: &str,
dir: &Path,
) {
if dir.exists() {
if is_writable_dir(dir) {
checks.push(Check::new(
id,
category,
Status::Pass,
format!("{} {}", label, dir.display()),
));
} else {
checks.push(
Check::new(
id,
category,
Status::Fail,
format!("{} not writable: {}", label, dir.display()),
)
.with_fix(format!("chmod u+rwx {}", dir.display())),
);
}
} else {
checks.push(Check::new(
id,
category,
Status::Info,
format!(
"{} does not exist yet (will be created on first use): {}",
label,
dir.display()
),
));
}
}
+247
View File
@@ -0,0 +1,247 @@
//! Destructive repair actions behind `--fix`: reinstall Chrome, close
//! version-mismatched daemons, purge expired state files, and generate a
//! missing encryption key.
use std::env;
use std::fs;
use std::path::Path;
use std::time::{Duration, SystemTime};
#[cfg(unix)]
use std::os::unix::fs::PermissionsExt;
use serde_json::json;
use super::helpers::new_id;
use super::{Check, Status};
use crate::connection::{cleanup_stale_files, send_command, walk_daemons};
use crate::native::state::{get_sessions_dir, get_state_dir};
pub(super) fn run(checks: &mut [Check], fixed: &mut Vec<String>) {
// `close_all_sessions` is expensive and closes every session at once, so
// only fire it on the first daemon.session.* Warn we encounter. Subsequent
// daemon.session.* Warn checks piggy-back on the same result.
let mut daemons_closed: Option<usize> = None;
for c in checks.iter_mut() {
match c.id.as_str() {
"chrome.installed" if c.status == Status::Fail => {
let installed = attempt_chrome_install();
if installed {
fixed.push("Reinstalled Chrome".to_string());
c.status = Status::Pass;
c.message = format!("{} (fixed by --fix)", c.message);
c.fix = None;
}
}
id if id.starts_with("daemon.session.") && c.status == Status::Warn => {
let killed = *daemons_closed.get_or_insert_with(|| {
let n = close_all_sessions();
if n > 0 {
fixed.push(format!("Closed {} version-mismatched daemon(s)", n));
}
n
});
if killed > 0 {
c.status = Status::Pass;
c.message = format!("{} (fixed by --fix)", c.message);
c.fix = None;
}
}
"security.state_count" if c.status == Status::Warn => {
let removed = purge_old_state();
if removed > 0 {
fixed.push(format!("Deleted {} expired state file(s)", removed));
c.status = Status::Pass;
c.message = format!("{} (fixed by --fix)", c.message);
c.fix = None;
}
}
"security.encryption_key" if c.status == Status::Info => {
let generated = create_encryption_key();
if generated {
fixed.push("Generated encryption key".to_string());
c.status = Status::Pass;
c.message = format!("{} (fixed by --fix)", c.message);
c.fix = None;
}
}
_ => {}
}
}
}
fn attempt_chrome_install() -> bool {
// run_install() uses process::exit on failure, so we shell out to ourselves
// to avoid taking down the doctor process if the install fails.
let exe = match std::env::current_exe() {
Ok(p) => p,
Err(_) => return false,
};
std::process::Command::new(exe)
.arg("install")
.status()
.map(|s| s.success())
.unwrap_or(false)
}
fn close_all_sessions() -> usize {
let mut killed = 0;
for session in &walk_daemons().sessions {
let cmd = json!({ "id": new_id(), "action": "close" });
if send_command(cmd, &session.name).is_ok() {
killed += 1;
}
cleanup_stale_files(&session.name);
}
killed
}
fn purge_old_state() -> usize {
let dir = get_sessions_dir();
let expire_days = env::var("AGENT_BROWSER_STATE_EXPIRE_DAYS")
.ok()
.and_then(|s| s.parse::<u64>().ok())
.unwrap_or(30);
let cutoff = SystemTime::now()
.checked_sub(Duration::from_secs(expire_days * 86_400))
.unwrap_or(SystemTime::UNIX_EPOCH);
let mut removed = 0;
if let Ok(entries) = fs::read_dir(&dir) {
for entry in entries.flatten() {
if entry.file_type().map(|t| t.is_file()).unwrap_or(false) {
if let Ok(meta) = entry.metadata() {
if let Ok(modified) = meta.modified() {
if modified < cutoff && fs::remove_file(entry.path()).is_ok() {
removed += 1;
}
}
}
}
}
}
removed
}
fn create_encryption_key() -> bool {
create_encryption_key_at(&get_state_dir())
}
fn create_encryption_key_at(dir: &Path) -> bool {
if fs::create_dir_all(dir).is_err() {
return false;
}
#[cfg(unix)]
{
let _ = fs::set_permissions(dir, fs::Permissions::from_mode(0o700));
}
let path = dir.join(".encryption-key");
if path.exists() {
return false;
}
let mut buf = [0u8; 32];
if getrandom::getrandom(&mut buf).is_err() {
return false;
}
let hex: String = buf.iter().map(|b| format!("{:02x}", b)).collect();
if fs::write(&path, format!("{}\n", hex)).is_err() {
return false;
}
#[cfg(unix)]
{
let _ = fs::set_permissions(&path, fs::Permissions::from_mode(0o600));
}
true
}
#[cfg(test)]
mod tests {
use super::*;
use tempfile::TempDir;
#[test]
fn test_create_encryption_key_at_writes_64_char_hex_key() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().join("state");
assert!(create_encryption_key_at(&dir));
let key = dir.join(".encryption-key");
assert!(key.exists(), "key file should be created");
let contents = fs::read_to_string(&key).unwrap();
let trimmed = contents.trim();
assert_eq!(trimmed.len(), 64, "key should be 64 hex chars");
assert!(
trimmed.chars().all(|c| c.is_ascii_hexdigit()),
"key should be all hex digits"
);
}
#[test]
fn test_create_encryption_key_at_is_idempotent() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().join("state");
assert!(create_encryption_key_at(&dir));
let original = fs::read_to_string(dir.join(".encryption-key")).unwrap();
// Second call returns false and must not overwrite the existing key.
assert!(!create_encryption_key_at(&dir));
let after = fs::read_to_string(dir.join(".encryption-key")).unwrap();
assert_eq!(original, after);
}
#[cfg(unix)]
#[test]
fn test_create_encryption_key_at_sets_0600_perms() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().join("state");
assert!(create_encryption_key_at(&dir));
let key = dir.join(".encryption-key");
let mode = fs::metadata(&key).unwrap().permissions().mode() & 0o777;
assert_eq!(mode, 0o600, "key file should be 0600, got {:o}", mode);
}
#[cfg(unix)]
#[test]
fn test_run_fixes_generates_missing_encryption_key() {
// Reaches the Info-status arm in run_fixes that was previously
// unreachable due to an early-continue guard. Overrides HOME so
// get_state_dir() resolves under a temp dir.
let guard = crate::test_utils::EnvGuard::new(&["HOME"]);
let tmp = TempDir::new().unwrap();
guard.set("HOME", tmp.path().to_str().unwrap());
let mut checks = vec![Check::new(
"security.encryption_key",
"Security",
Status::Info,
"No encryption key set",
)
.with_fix("export AGENT_BROWSER_ENCRYPTION_KEY=...")];
let mut fixed = Vec::new();
run(&mut checks, &mut fixed);
assert_eq!(
checks[0].status,
Status::Pass,
"Info check should transition to Pass after --fix"
);
assert!(
checks[0].fix.is_none(),
"fix hint should be cleared after repair"
);
assert!(
fixed.iter().any(|s| s.contains("encryption key")),
"fixed summary should mention the key generation"
);
assert!(
tmp.path().join(".agent-browser/.encryption-key").exists(),
"key file should exist at ~/.agent-browser/.encryption-key"
);
}
}
+185
View File
@@ -0,0 +1,185 @@
//! Stateless helpers shared across doctor submodules.
use std::fs;
use std::path::Path;
use std::sync::atomic::{AtomicU64, Ordering};
use std::time::SystemTime;
use serde_json::Value;
pub(super) fn is_writable_dir(path: &Path) -> bool {
fs::metadata(path)
.map(|m| !m.permissions().readonly())
.unwrap_or(false)
}
pub(super) fn human_size(bytes: u64) -> String {
const UNITS: &[&str] = &["B", "KB", "MB", "GB", "TB"];
let mut value = bytes as f64;
let mut unit = 0;
while value >= 1024.0 && unit < UNITS.len() - 1 {
value /= 1024.0;
unit += 1;
}
if unit == 0 {
format!("{} {}", bytes, UNITS[0])
} else {
format!("{:.1} {}", value, UNITS[unit])
}
}
#[cfg(unix)]
pub(super) fn disk_free_bytes(path: &Path) -> Option<u64> {
use std::ffi::CString;
use std::os::unix::ffi::OsStrExt;
use std::path::PathBuf;
// Walk up to the first existing ancestor (for fresh installs where the
// state dir hasn't been created yet).
let mut probe: PathBuf = path.to_path_buf();
while !probe.exists() {
match probe.parent() {
Some(p) => probe = p.to_path_buf(),
None => return None,
}
}
let c_path = CString::new(probe.as_os_str().as_bytes()).ok()?;
let mut stat: libc::statvfs = unsafe { std::mem::zeroed() };
if unsafe { libc::statvfs(c_path.as_ptr(), &mut stat) } != 0 {
return None;
}
Some(stat.f_bavail as u64 * stat.f_frsize)
}
#[cfg(windows)]
pub(super) fn disk_free_bytes(_path: &Path) -> Option<u64> {
None
}
#[cfg(not(any(unix, windows)))]
pub(super) fn disk_free_bytes(_path: &Path) -> Option<u64> {
None
}
pub(super) fn which_exists(name: &str) -> bool {
let probe = if cfg!(target_os = "windows") {
"where"
} else {
"which"
};
std::process::Command::new(probe)
.arg(name)
.stdout(std::process::Stdio::null())
.stderr(std::process::Stdio::null())
.status()
.map(|s| s.success())
.unwrap_or(false)
}
pub(super) fn parse_json_file(path: &Path) -> Result<(), String> {
let content = fs::read_to_string(path).map_err(|e| format!("read failed: {}", e))?;
serde_json::from_str::<Value>(&content).map_err(|e| format!("invalid JSON: {}", e))?;
Ok(())
}
/// Generate a unique `doctor-<pid>-<micros>-<sequence>` id for JSON command envelopes.
pub(super) fn new_id() -> String {
static NEXT_ID: AtomicU64 = AtomicU64::new(0);
let sequence = NEXT_ID.fetch_add(1, Ordering::Relaxed);
format!(
"doctor-{}-{}-{}",
std::process::id(),
SystemTime::now()
.duration_since(SystemTime::UNIX_EPOCH)
.map(|d| d.as_micros())
.unwrap_or(0),
sequence
)
}
#[cfg(test)]
mod tests {
use super::*;
use tempfile::TempDir;
#[test]
fn test_human_size_units() {
assert_eq!(human_size(0), "0 B");
assert_eq!(human_size(512), "512 B");
assert_eq!(human_size(1024), "1.0 KB");
assert_eq!(human_size(1024 * 1024), "1.0 MB");
assert_eq!(human_size(1024 * 1024 * 1024), "1.0 GB");
assert_eq!(human_size(1_500_000), "1.4 MB");
}
#[test]
fn test_disk_free_walks_up_to_existing_ancestor() {
let dir = TempDir::new().unwrap();
let nested = dir.path().join("a/b/c/d");
let bytes = disk_free_bytes(&nested);
if cfg!(unix) {
assert!(bytes.is_some());
assert!(bytes.unwrap() > 0);
}
}
#[test]
fn test_is_writable_dir_matches_metadata() {
let dir = TempDir::new().unwrap();
assert!(is_writable_dir(dir.path()));
let missing = dir.path().join("does-not-exist");
assert!(!is_writable_dir(&missing));
}
#[test]
fn test_which_exists_matches_common_binaries() {
// `sh` exists on every unix; `cmd` exists on windows.
let probe = if cfg!(target_os = "windows") {
"cmd"
} else {
"sh"
};
assert!(which_exists(probe));
assert!(!which_exists(
"agent-browser-this-does-not-exist-please-dont-install-it"
));
}
#[test]
fn test_parse_json_file_valid_and_invalid() {
let dir = TempDir::new().unwrap();
let valid = dir.path().join("ok.json");
fs::write(&valid, r#"{"k": 1}"#).unwrap();
assert!(parse_json_file(&valid).is_ok());
let invalid = dir.path().join("bad.json");
fs::write(&invalid, "{not json}").unwrap();
let err = parse_json_file(&invalid).unwrap_err();
assert!(err.contains("invalid JSON"));
let missing = dir.path().join("nope.json");
let err = parse_json_file(&missing).unwrap_err();
assert!(err.contains("read failed"));
}
#[test]
fn test_parse_json_file_accepts_arrays() {
// The config parser rejects arrays at the Config type level, but
// doctor only checks syntactic JSON validity so it should accept
// both arrays and objects.
let dir = TempDir::new().unwrap();
let path = dir.path().join("arr.json");
fs::write(&path, r#"[1, 2, 3]"#).unwrap();
assert!(parse_json_file(&path).is_ok());
}
#[test]
fn test_new_id_is_unique_per_call() {
let a = new_id();
let b = new_id();
assert_ne!(a, b);
assert!(a.starts_with("doctor-"));
}
}
+188
View File
@@ -0,0 +1,188 @@
//! Live launch test: spawn a scratch daemon session, launch headless
//! Chrome, navigate to `about:blank`, then close. Skipped under `--quick`.
//!
//! A `LaunchGuard` Drop impl ensures the scratch session is closed and its
//! sidecar files cleaned even on panic or early return.
use std::env;
use std::time::{Duration, Instant, SystemTime};
use serde_json::{json, Value};
use super::helpers::new_id;
use super::{Check, Status};
use crate::connection::{cleanup_stale_files, ensure_daemon, send_command, DaemonOptions};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Launch test";
if env::var("AGENT_BROWSER_PROVIDER").is_ok() {
checks.push(Check::new(
"launch.skipped.provider",
category,
Status::Info,
"Skipped (AGENT_BROWSER_PROVIDER is set; would consume cloud quota)",
));
return;
}
if env::var("AGENT_BROWSER_CDP").is_ok() {
checks.push(Check::new(
"launch.skipped.cdp",
category,
Status::Info,
"Skipped (AGENT_BROWSER_CDP is set; would attach to a real browser)",
));
return;
}
let session = format!(
"doctor-{}-{}",
std::process::id(),
SystemTime::now()
.duration_since(SystemTime::UNIX_EPOCH)
.map(|d| d.as_millis())
.unwrap_or(0)
);
// Armed after `ensure_daemon` succeeds so we don't send a stray `close`
// or delete sidecar files for a daemon that never started. On every early
// return past the `Some(...)` assignment below, Drop runs one close and
// one `cleanup_stale_files`.
let mut _guard: Option<LaunchGuard> = None;
let opts = DaemonOptions {
headed: false,
debug: false,
executable_path: None,
extensions: &[],
init_scripts: &[],
enable: &[],
args: None,
user_agent: None,
proxy: None,
proxy_bypass: None,
proxy_username: None,
proxy_password: None,
ignore_https_errors: false,
allow_file_access: false,
hide_scrollbars: true,
profile: None,
state: None,
provider: None,
device: None,
session_name: None,
download_path: None,
allowed_domains: None,
action_policy: None,
confirm_actions: None,
engine: None,
auto_connect: false,
force_launch: false,
idle_timeout: None,
default_timeout: None,
cdp: None,
no_auto_dialog: false,
};
let started = Instant::now();
if let Err(e) = ensure_daemon(&session, &opts) {
checks.push(
Check::new(
"launch.daemon",
category,
Status::Fail,
format!("Could not start daemon: {}", e),
)
.with_fix("check Chrome install and re-run with --debug"),
);
return;
}
_guard = Some(LaunchGuard {
session: session.clone(),
});
let launch_cmd = json!({
"id": new_id(),
"action": "launch",
"headless": true,
});
if let Err(e) = send_json(launch_cmd, &session) {
checks.push(
Check::new(
"launch.launch",
category,
Status::Fail,
format!("Browser launch failed: {}", e),
)
.with_fix("agent-browser install # or check --debug output"),
);
return;
}
let open_cmd = json!({
"id": new_id(),
"action": "navigate",
"url": "about:blank",
});
if let Err(e) = send_json(open_cmd, &session) {
checks.push(
Check::new(
"launch.navigate",
category,
Status::Fail,
format!("Navigation to about:blank failed: {}", e),
)
.with_fix("re-run with --debug for full launch logs"),
);
return;
}
// Close + stale-file cleanup happen exactly once via LaunchGuard::drop at
// end of scope; no explicit close here.
let elapsed = started.elapsed();
let secs = elapsed.as_secs_f64();
if elapsed > Duration::from_secs(5) {
checks.push(Check::new(
"launch.elapsed",
category,
Status::Warn,
format!(
"Headless launch + about:blank in {:.2}s (slow; expected < 5s)",
secs
),
));
} else {
checks.push(Check::new(
"launch.elapsed",
category,
Status::Pass,
format!("Headless launch + about:blank in {:.2}s", secs),
));
}
}
fn send_json(cmd: Value, session: &str) -> Result<(), String> {
match send_command(cmd, session) {
Ok(resp) => {
if resp.success {
Ok(())
} else {
Err(resp.error.unwrap_or_else(|| "unknown error".to_string()))
}
}
Err(e) => Err(e),
}
}
/// Best-effort cleanup when the launch test panics or returns early.
struct LaunchGuard {
session: String,
}
impl Drop for LaunchGuard {
fn drop(&mut self) {
let close_cmd = json!({ "id": new_id(), "action": "close" });
let _ = send_command(close_cmd, &self.session);
cleanup_stale_files(&self.session);
}
}
+289
View File
@@ -0,0 +1,289 @@
//! Diagnose an agent-browser installation.
//!
//! Runs a battery of checks across environment, Chrome install, daemon
//! state, config files, encryption, providers, network reachability, and
//! a live headless browser launch test.
//!
//! Auto-cleans stale daemon socket/pid/version sidecar files. Destructive
//! repairs (reinstalling Chrome, purging old state files, generating a
//! missing encryption key) are gated behind `--fix`.
mod chrome;
mod config;
mod daemon;
mod environment;
mod fix;
mod helpers;
mod launch;
mod network;
mod providers;
mod security;
use serde_json::{json, Value};
use crate::color;
#[derive(Default, Clone, Copy)]
pub struct DoctorOptions {
pub offline: bool,
pub quick: bool,
pub fix: bool,
pub json: bool,
}
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
#[repr(u8)]
pub(crate) enum Status {
Pass,
Warn,
Fail,
Info,
}
impl Status {
fn as_str(&self) -> &'static str {
match self {
Status::Pass => "pass",
Status::Warn => "warn",
Status::Fail => "fail",
Status::Info => "info",
}
}
fn label(&self) -> String {
match self {
Status::Pass => color::green("pass"),
Status::Warn => color::yellow("warn"),
Status::Fail => color::red("fail"),
Status::Info => color::dim("info"),
}
}
}
#[derive(Clone)]
pub(crate) struct Check {
pub id: String,
pub category: &'static str,
pub status: Status,
pub message: String,
pub fix: Option<String>,
}
impl Check {
fn new(
id: impl Into<String>,
category: &'static str,
status: Status,
message: impl Into<String>,
) -> Self {
Self {
id: id.into(),
category,
status,
message: message.into(),
fix: None,
}
}
fn with_fix(mut self, fix: impl Into<String>) -> Self {
self.fix = Some(fix.into());
self
}
}
/// Run the doctor command. Returns the process exit code.
pub fn run_doctor(opts: DoctorOptions) -> i32 {
let mut checks: Vec<Check> = Vec::new();
let mut fixed: Vec<String> = Vec::new();
environment::check(&mut checks);
chrome::check(&mut checks);
daemon::check(&mut checks);
config::check(&mut checks);
security::check(&mut checks);
providers::check(&mut checks);
if !opts.offline {
network::check(&mut checks);
}
if !opts.quick {
launch::check(&mut checks);
}
if opts.fix {
fix::run(&mut checks, &mut fixed);
}
let summary = summarize(&checks);
let exit_code = if summary.fail > 0 { 1 } else { 0 };
if opts.json {
print_json(&checks, &summary, &fixed, exit_code == 0);
} else {
print_text(&checks, &summary, &fixed, opts.fix);
}
exit_code
}
struct Summary {
pass: usize,
warn: usize,
fail: usize,
}
fn summarize(checks: &[Check]) -> Summary {
let mut s = Summary {
pass: 0,
warn: 0,
fail: 0,
};
for c in checks {
match c.status {
Status::Pass => s.pass += 1,
Status::Warn => s.warn += 1,
Status::Fail => s.fail += 1,
Status::Info => {}
}
}
s
}
fn print_text(checks: &[Check], summary: &Summary, fixed: &[String], fix_ran: bool) {
println!("{}", color::bold("agent-browser doctor"));
let mut current_category = "";
for c in checks {
if c.category != current_category {
current_category = c.category;
println!();
println!("{}", color::bold(current_category));
}
println!(" {} {}", c.status.label(), c.message);
if let Some(fix) = &c.fix {
println!(" {} {}", color::dim("fix:"), fix);
}
}
if !fixed.is_empty() {
println!();
println!("{}", color::bold("Fixed"));
for line in fixed {
println!(" {} {}", color::green("done"), line);
}
}
println!();
let line = format!(
"Summary: {} pass, {} warn, {} fail",
summary.pass, summary.warn, summary.fail
);
if summary.fail > 0 {
println!("{}", color::red(&line));
} else if summary.warn > 0 {
println!("{}", color::yellow(&line));
} else {
println!("{}", color::green(&line));
}
if !fix_ran && checks.iter().any(|c| c.fix.is_some()) {
println!();
println!(
"{} Run with {} to attempt repairs.",
color::dim("tip:"),
color::bold("--fix")
);
}
}
fn print_json(checks: &[Check], summary: &Summary, fixed: &[String], success: bool) {
let checks_json: Vec<Value> = checks
.iter()
.map(|c| {
let mut obj = json!({
"id": c.id,
"category": c.category,
"status": c.status.as_str(),
"message": c.message,
});
if let Some(fix) = &c.fix {
obj["fix"] = json!(fix);
}
obj
})
.collect();
let payload = json!({
"success": success,
"summary": {
"pass": summary.pass,
"warn": summary.warn,
"fail": summary.fail,
},
"checks": checks_json,
"fixed": fixed,
});
println!("{}", payload);
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_summary_counts_each_status() {
let checks = vec![
Check::new("a", "Cat", Status::Pass, "ok"),
Check::new("b", "Cat", Status::Pass, "ok"),
Check::new("c", "Cat", Status::Warn, "meh"),
Check::new("d", "Cat", Status::Fail, "no"),
Check::new("e", "Cat", Status::Info, "fyi"),
];
let s = summarize(&checks);
assert_eq!(s.pass, 2);
assert_eq!(s.warn, 1);
assert_eq!(s.fail, 1);
}
#[test]
fn test_summary_zeroes_when_only_info() {
let checks = vec![Check::new("a", "Cat", Status::Info, "ignored")];
let s = summarize(&checks);
assert_eq!(s.pass, 0);
assert_eq!(s.warn, 0);
assert_eq!(s.fail, 0);
}
#[test]
fn test_status_label_does_not_panic() {
for s in &[Status::Pass, Status::Warn, Status::Fail, Status::Info] {
assert!(!s.label().is_empty());
assert!(!s.as_str().is_empty());
}
}
#[test]
fn test_status_as_str_values() {
assert_eq!(Status::Pass.as_str(), "pass");
assert_eq!(Status::Warn.as_str(), "warn");
assert_eq!(Status::Fail.as_str(), "fail");
assert_eq!(Status::Info.as_str(), "info");
}
#[test]
fn test_check_new_and_with_fix() {
let c = Check::new("id", "cat", Status::Warn, "msg").with_fix("do thing");
assert_eq!(c.id, "id");
assert_eq!(c.category, "cat");
assert_eq!(c.status, Status::Warn);
assert_eq!(c.message, "msg");
assert_eq!(c.fix.as_deref(), Some("do thing"));
}
#[test]
fn test_check_new_no_fix_by_default() {
let c = Check::new("id", "cat", Status::Pass, "msg");
assert!(c.fix.is_none());
}
}
+154
View File
@@ -0,0 +1,154 @@
//! Probe reachability of the Chrome for Testing CDN, AI Gateway (if
//! configured), and the currently-selected provider endpoint. Each probe
//! has a 3-second timeout.
use std::env;
use std::time::{Duration, Instant};
use super::{Check, Status};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Network";
let rt = match tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
{
Ok(r) => r,
Err(e) => {
checks.push(Check::new(
"net.runtime",
category,
Status::Fail,
format!("Could not start tokio runtime for probes: {}", e),
));
return;
}
};
let client = match reqwest::Client::builder()
.user_agent(format!("agent-browser/{}", env!("CARGO_PKG_VERSION")))
.timeout(Duration::from_secs(3))
.connect_timeout(Duration::from_secs(3))
.build()
{
Ok(c) => c,
Err(e) => {
checks.push(Check::new(
"net.client",
category,
Status::Fail,
format!("Could not build HTTP client: {}", e),
));
return;
}
};
let chrome_url =
"https://googlechromelabs.github.io/chrome-for-testing/last-known-good-versions-with-downloads.json";
probe_url(
&rt,
&client,
checks,
category,
"net.chrome_cdn",
chrome_url,
"Chrome for Testing CDN",
);
if env::var("AI_GATEWAY_API_KEY").is_ok() {
let url = env::var("AI_GATEWAY_URL")
.unwrap_or_else(|_| "https://ai-gateway.vercel.sh".to_string());
probe_url(
&rt,
&client,
checks,
category,
"net.ai_gateway",
&url,
"AI Gateway",
);
}
if let Ok(provider) = env::var("AGENT_BROWSER_PROVIDER") {
let url: Option<String> = match provider.to_lowercase().as_str() {
"browserbase" => Some("https://api.browserbase.com".to_string()),
"browserless" => Some(
env::var("BROWSERLESS_API_URL")
.unwrap_or_else(|_| "https://production-sfo.browserless.io".to_string()),
),
"browseruse" | "browser-use" => Some("https://api.browser-use.com".to_string()),
"kernel" => Some(
env::var("KERNEL_ENDPOINT")
.unwrap_or_else(|_| "https://api.onkernel.com".to_string()),
),
_ => None,
};
if let Some(url) = url {
probe_url(
&rt,
&client,
checks,
category,
"net.provider",
&url,
&format!("Provider {}", provider),
);
}
}
}
fn probe_url(
rt: &tokio::runtime::Runtime,
client: &reqwest::Client,
checks: &mut Vec<Check>,
category: &'static str,
id: &'static str,
url: &str,
label: &str,
) {
let started = Instant::now();
let result = rt.block_on(async { client.head(url).send().await });
let elapsed_ms = started.elapsed().as_millis();
match result {
Ok(resp) => {
let status = resp.status();
if status.is_success() || status.is_redirection() || status.as_u16() == 405 {
checks.push(Check::new(
id,
category,
Status::Pass,
format!(
"{} reachable ({}ms, HTTP {})",
label,
elapsed_ms,
status.as_u16()
),
));
} else {
checks.push(Check::new(
id,
category,
Status::Warn,
format!(
"{} returned HTTP {} after {}ms",
label,
status.as_u16(),
elapsed_ms
),
));
}
}
Err(e) => {
checks.push(
Check::new(
id,
category,
Status::Fail,
format!("{} unreachable after {}ms: {}", label, elapsed_ms, e),
)
.with_fix("check network connectivity / firewall / proxy settings"),
);
}
}
}
+128
View File
@@ -0,0 +1,128 @@
//! Check remote browser providers: API key presence for Browserless,
//! Browserbase, Browser Use, Kernel, AgentCore (AWS), Appium for iOS, and
//! the AI Gateway chat key. Info-level unless the provider is selected
//! via `AGENT_BROWSER_PROVIDER`.
use std::env;
use super::helpers::which_exists;
use super::{Check, Status};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Providers";
let active = env::var("AGENT_BROWSER_PROVIDER").ok();
let normalized = active
.as_ref()
.map(|s| s.to_lowercase())
.unwrap_or_default();
let active_status = |provider: &str, ok: bool| -> Status {
if normalized == provider {
if ok {
Status::Pass
} else {
Status::Fail
}
} else {
Status::Info
}
};
let providers: &[(&str, &[&str], &str)] = &[
("browserless", &["BROWSERLESS_API_KEY"], "Browserless"),
("browserbase", &["BROWSERBASE_API_KEY"], "Browserbase"),
("browseruse", &["BROWSER_USE_API_KEY"], "Browser Use"),
("kernel", &["KERNEL_API_KEY"], "Kernel"),
];
for (id, env_keys, label) in providers {
let present = env_keys.iter().any(|k| env::var(k).is_ok());
let provider_id = *id;
let status = active_status(provider_id, present);
let msg = if present {
format!("{}: API key present", label)
} else {
format!("{}: {} not set", label, env_keys.join(" / "))
};
let mut check = Check::new(format!("providers.{}", provider_id), category, status, msg);
if status == Status::Fail {
check = check.with_fix(format!(
"set {} (or unset AGENT_BROWSER_PROVIDER={})",
env_keys.first().copied().unwrap_or(""),
provider_id
));
}
checks.push(check);
}
let aws_present = env::var("AWS_ACCESS_KEY_ID").is_ok()
|| env::var("AWS_PROFILE").is_ok()
|| env::var("AWS_SESSION_TOKEN").is_ok();
let agentcore_status = active_status("agentcore", aws_present);
let mut agentcore_check = Check::new(
"providers.agentcore",
category,
agentcore_status,
if aws_present {
"AgentCore: AWS credentials resolvable".to_string()
} else {
"AgentCore: no AWS credentials in env (AWS_ACCESS_KEY_ID / AWS_PROFILE)".to_string()
},
);
if agentcore_status == Status::Fail {
agentcore_check = agentcore_check
.with_fix("export AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY or AWS_PROFILE");
}
checks.push(agentcore_check);
if normalized == "ios" {
if which_exists("appium") {
checks.push(Check::new(
"providers.ios",
category,
Status::Pass,
"iOS: appium binary on PATH",
));
} else {
checks.push(
Check::new(
"providers.ios",
category,
Status::Fail,
"iOS: appium binary not found on PATH",
)
.with_fix("npm install -g appium && appium driver install xcuitest"),
);
}
}
let chat_key_present = env::var("AI_GATEWAY_API_KEY").is_ok();
if chat_key_present {
checks.push(Check::new(
"providers.chat",
category,
Status::Info,
"AI_GATEWAY_API_KEY present (chat enabled)",
));
} else {
checks.push(
Check::new(
"providers.chat",
category,
Status::Info,
"AI_GATEWAY_API_KEY not set (chat command disabled)",
)
.with_fix("export AI_GATEWAY_API_KEY=gw_..."),
);
}
if let Some(active) = active {
checks.push(Check::new(
"providers.active",
category,
Status::Info,
format!("AGENT_BROWSER_PROVIDER = {}", active),
));
}
}
+167
View File
@@ -0,0 +1,167 @@
//! Check security posture: encryption key presence / permissions, saved
//! state file age, and the optional action policy file.
use std::env;
use std::fs;
use std::path::PathBuf;
use std::time::{Duration, SystemTime};
#[cfg(unix)]
use std::os::unix::fs::PermissionsExt;
use super::helpers::parse_json_file;
use super::{Check, Status};
use crate::native::state::{get_sessions_dir, get_state_dir};
pub(super) fn check(checks: &mut Vec<Check>) {
let category = "Security";
let key_env = env::var("AGENT_BROWSER_ENCRYPTION_KEY").ok();
let key_file = get_state_dir().join(".encryption-key");
if let Some(hex) = &key_env {
if hex.len() == 64 && hex.chars().all(|c| c.is_ascii_hexdigit()) {
checks.push(Check::new(
"security.encryption_key",
category,
Status::Pass,
"AGENT_BROWSER_ENCRYPTION_KEY set (64-char hex)",
));
} else {
checks.push(
Check::new(
"security.encryption_key",
category,
Status::Fail,
"AGENT_BROWSER_ENCRYPTION_KEY is not a 64-char hex string",
)
.with_fix("export AGENT_BROWSER_ENCRYPTION_KEY=$(openssl rand -hex 32)"),
);
}
} else if key_file.exists() {
let mut msg = format!("Encryption key file present: {}", key_file.display());
let mut status = Status::Pass;
let mut fix: Option<String> = None;
#[cfg(unix)]
if let Ok(meta) = fs::metadata(&key_file) {
let mode = meta.permissions().mode() & 0o777;
if mode & 0o077 != 0 {
status = Status::Warn;
msg = format!(
"Encryption key file is too permissive ({:o}): {}",
mode,
key_file.display()
);
fix = Some(format!("chmod 600 {}", key_file.display()));
}
}
let mut check = Check::new("security.encryption_key", category, status, msg);
if let Some(f) = fix {
check = check.with_fix(f);
}
checks.push(check);
} else {
checks.push(
Check::new(
"security.encryption_key",
category,
Status::Info,
"No encryption key set (will be auto-generated on first auth save)",
)
.with_fix("export AGENT_BROWSER_ENCRYPTION_KEY=$(openssl rand -hex 32)"),
);
}
let sessions_dir = get_sessions_dir();
if sessions_dir.exists() {
let expire_days = env::var("AGENT_BROWSER_STATE_EXPIRE_DAYS")
.ok()
.and_then(|s| s.parse::<u64>().ok())
.unwrap_or(30);
let cutoff = SystemTime::now()
.checked_sub(Duration::from_secs(expire_days * 86_400))
.unwrap_or(SystemTime::UNIX_EPOCH);
let mut total = 0usize;
let mut old = 0usize;
if let Ok(entries) = fs::read_dir(&sessions_dir) {
for entry in entries.flatten() {
if entry.file_type().map(|t| t.is_file()).unwrap_or(false) {
total += 1;
if let Ok(meta) = entry.metadata() {
if let Ok(modified) = meta.modified() {
if modified < cutoff {
old += 1;
}
}
}
}
}
}
if total == 0 {
checks.push(Check::new(
"security.state_count",
category,
Status::Info,
"No saved state files",
));
} else if old > 0 {
checks.push(
Check::new(
"security.state_count",
category,
Status::Warn,
format!(
"{} state file(s) older than {} days ({} total)",
old, expire_days, total
),
)
.with_fix(format!(
"agent-browser state clean --older-than {}",
expire_days
)),
);
} else {
checks.push(Check::new(
"security.state_count",
category,
Status::Pass,
format!("{} saved state file(s)", total),
));
}
}
if let Ok(policy_path) = env::var("AGENT_BROWSER_ACTION_POLICY") {
let p = PathBuf::from(&policy_path);
if !p.exists() {
checks.push(
Check::new(
"security.action_policy",
category,
Status::Fail,
format!(
"AGENT_BROWSER_ACTION_POLICY points to missing file: {}",
policy_path
),
)
.with_fix("update or unset AGENT_BROWSER_ACTION_POLICY"),
);
} else {
match parse_json_file(&p) {
Ok(_) => checks.push(Check::new(
"security.action_policy",
category,
Status::Pass,
format!("Action policy: {}", policy_path),
)),
Err(e) => checks.push(
Check::new(
"security.action_policy",
category,
Status::Fail,
format!("Action policy: {}: {}", policy_path, e),
)
.with_fix(format!("edit {}", policy_path)),
),
}
}
}
}
+186 -3
View File
@@ -60,6 +60,8 @@ pub struct Config {
pub session_name: Option<String>,
pub executable_path: Option<String>,
pub extensions: Option<Vec<String>>,
pub init_scripts: Option<Vec<String>>,
pub enable: Option<Vec<String>>,
pub profile: Option<String>,
pub state: Option<String>,
pub proxy: Option<String>,
@@ -68,6 +70,7 @@ pub struct Config {
pub user_agent: Option<String>,
pub provider: Option<String>,
pub device: Option<String>,
pub hide_scrollbars: Option<bool>,
pub ignore_https_errors: Option<bool>,
pub allow_file_access: Option<bool>,
pub cdp: Option<String>,
@@ -88,6 +91,7 @@ pub struct Config {
pub screenshot_format: Option<String>,
pub idle_timeout: Option<String>,
pub no_auto_dialog: Option<bool>,
pub model: Option<String>,
}
impl Config {
@@ -106,6 +110,20 @@ impl Config {
}
(a, b) => b.or(a),
},
init_scripts: match (self.init_scripts, other.init_scripts) {
(Some(mut a), Some(b)) => {
a.extend(b);
Some(a)
}
(a, b) => b.or(a),
},
enable: match (self.enable, other.enable) {
(Some(mut a), Some(b)) => {
a.extend(b);
Some(a)
}
(a, b) => b.or(a),
},
profile: other.profile.or(self.profile),
state: other.state.or(self.state),
proxy: other.proxy.or(self.proxy),
@@ -114,6 +132,7 @@ impl Config {
user_agent: other.user_agent.or(self.user_agent),
provider: other.provider.or(self.provider),
device: other.device.or(self.device),
hide_scrollbars: other.hide_scrollbars.or(self.hide_scrollbars),
ignore_https_errors: other.ignore_https_errors.or(self.ignore_https_errors),
allow_file_access: other.allow_file_access.or(self.allow_file_access),
cdp: other.cdp.or(self.cdp),
@@ -134,6 +153,7 @@ impl Config {
screenshot_format: other.screenshot_format.or(self.screenshot_format),
idle_timeout: other.idle_timeout.or(self.idle_timeout),
no_auto_dialog: other.no_auto_dialog.or(self.no_auto_dialog),
model: other.model.or(self.model),
}
}
}
@@ -169,6 +189,12 @@ fn env_var_is_truthy(name: &str) -> bool {
}
}
fn env_var_bool(name: &str) -> Option<bool> {
env::var(name)
.ok()
.map(|val| !matches!(val.to_lowercase().as_str(), "0" | "false" | "no" | ""))
}
/// Parse an optional boolean value after a flag. Returns (value, consumed_next_arg).
/// Recognizes "true" as true, "false" as false. Bare flag defaults to true.
fn parse_bool_arg(args: &[String], i: usize) -> (bool, bool) {
@@ -198,6 +224,8 @@ fn extract_config_path(args: &[String]) -> Option<Option<String>> {
"--executable-path",
"--cdp",
"--extension",
"--init-script",
"--enable",
"--profile",
"--state",
"--proxy",
@@ -219,6 +247,7 @@ fn extract_config_path(args: &[String]) -> Option<Option<String>> {
"--screenshot-quality",
"--screenshot-format",
"--idle-timeout",
"--model",
];
let mut i = 0;
while i < args.len() {
@@ -274,6 +303,8 @@ pub struct Flags {
pub executable_path: Option<String>,
pub cdp: Option<String>,
pub extensions: Vec<String>,
pub init_scripts: Vec<String>,
pub enable: Vec<String>,
pub profile: Option<String>,
pub state: Option<String>,
pub proxy: Option<String>,
@@ -283,8 +314,10 @@ pub struct Flags {
pub provider: Option<String>,
pub ignore_https_errors: bool,
pub allow_file_access: bool,
pub hide_scrollbars: bool,
pub device: Option<String>,
pub auto_connect: bool,
pub force_launch: bool,
pub session_name: Option<String>,
pub annotate: bool,
pub color_scheme: Option<String>,
@@ -300,12 +333,18 @@ pub struct Flags {
pub screenshot_quality: Option<u32>,
pub screenshot_format: Option<String>,
pub idle_timeout: Option<String>, // Canonical milliseconds string for AGENT_BROWSER_IDLE_TIMEOUT_MS
pub default_timeout: Option<u64>, // AGENT_BROWSER_DEFAULT_TIMEOUT in ms
pub no_auto_dialog: bool,
pub model: Option<String>,
pub verbose: bool,
pub quiet: bool,
// Track which launch-time options were explicitly passed via CLI
// (as opposed to being set only via environment variables)
pub cli_executable_path: bool,
pub cli_extensions: bool,
pub cli_init_scripts: bool,
pub cli_enable: bool,
pub cli_profile: bool,
pub cli_state: bool,
pub cli_args: bool,
@@ -313,6 +352,7 @@ pub struct Flags {
pub cli_proxy: bool,
pub cli_proxy_bypass: bool,
pub cli_allow_file_access: bool,
pub cli_hide_scrollbars: bool,
pub cli_annotate: bool,
pub cli_download_path: bool,
pub cli_headed: bool,
@@ -340,6 +380,38 @@ pub fn parse_flags(args: &[String]) -> Flags {
config.extensions.unwrap_or_default()
};
let init_scripts_env = env::var("AGENT_BROWSER_INIT_SCRIPTS")
.ok()
.map(|s| {
s.split(',')
.map(|p| p.trim().to_string())
.filter(|p| !p.is_empty())
.collect::<Vec<_>>()
})
.unwrap_or_default();
let init_scripts = if !init_scripts_env.is_empty() {
init_scripts_env
} else {
config.init_scripts.unwrap_or_default()
};
let enable_env = env::var("AGENT_BROWSER_ENABLE")
.ok()
.map(|s| {
s.split(',')
.map(|p| p.trim().to_string())
.filter(|p| !p.is_empty())
.collect::<Vec<_>>()
})
.unwrap_or_default();
let enable = if !enable_env.is_empty() {
enable_env
} else {
config.enable.unwrap_or_default()
};
let mut flags = Flags {
json: env_var_is_truthy("AGENT_BROWSER_JSON") || config.json.unwrap_or(false),
headed: env_var_is_truthy("AGENT_BROWSER_HEADED") || config.headed.unwrap_or(false),
@@ -354,6 +426,8 @@ pub fn parse_flags(args: &[String]) -> Flags {
.or(config.executable_path),
cdp: config.cdp,
extensions,
init_scripts,
enable,
profile: env::var("AGENT_BROWSER_PROFILE").ok().or(config.profile),
state: env::var("AGENT_BROWSER_STATE").ok().or(config.state),
proxy: env::var("AGENT_BROWSER_PROXY")
@@ -379,9 +453,15 @@ pub fn parse_flags(args: &[String]) -> Flags {
|| config.ignore_https_errors.unwrap_or(false),
allow_file_access: env_var_is_truthy("AGENT_BROWSER_ALLOW_FILE_ACCESS")
|| config.allow_file_access.unwrap_or(false),
hide_scrollbars: env_var_bool("AGENT_BROWSER_HIDE_SCROLLBARS")
.or(config.hide_scrollbars)
.unwrap_or(true),
device: env::var("AGENT_BROWSER_IOS_DEVICE").ok().or(config.device),
auto_connect: env_var_is_truthy("AGENT_BROWSER_AUTO_CONNECT")
|| config.auto_connect.unwrap_or(false),
auto_connect: !env_var_is_truthy("AGENT_BROWSER_NO_AUTO_CONNECT")
&& (env_var_is_truthy("AGENT_BROWSER_AUTO_CONNECT")
|| config.auto_connect.unwrap_or(true)),
force_launch: env_var_is_truthy("AGENT_BROWSER_FORCE_LAUNCH")
|| env::var("CI").is_ok(),
session_name: env::var("AGENT_BROWSER_SESSION_NAME")
.ok()
.or(config.session_name),
@@ -432,10 +512,18 @@ pub fn parse_flags(args: &[String]) -> Flags {
"AGENT_BROWSER_IDLE_TIMEOUT_MS",
)
.or(config.idle_timeout),
default_timeout: env::var("AGENT_BROWSER_DEFAULT_TIMEOUT")
.ok()
.and_then(|s| s.parse::<u64>().ok()),
no_auto_dialog: env_var_is_truthy("AGENT_BROWSER_NO_AUTO_DIALOG")
|| config.no_auto_dialog.unwrap_or(false),
model: env::var("AI_GATEWAY_MODEL").ok().or(config.model),
verbose: false,
quiet: false,
cli_executable_path: false,
cli_extensions: false,
cli_init_scripts: false,
cli_enable: false,
cli_profile: false,
cli_state: false,
cli_args: false,
@@ -443,6 +531,7 @@ pub fn parse_flags(args: &[String]) -> Flags {
cli_proxy: false,
cli_proxy_bypass: false,
cli_allow_file_access: false,
cli_hide_scrollbars: false,
cli_annotate: false,
cli_download_path: false,
cli_headed: false,
@@ -512,6 +601,27 @@ pub fn parse_flags(args: &[String]) -> Flags {
i += 1;
}
}
"--init-script" => {
if let Some(s) = args.get(i + 1) {
flags.init_scripts.push(s.clone());
flags.cli_init_scripts = true;
i += 1;
}
}
"--enable" => {
if let Some(s) = args.get(i + 1) {
// Allow either repeated --enable foo --enable bar, or
// a single --enable foo,bar comma-list for convenience.
for item in s.split(',') {
let trimmed = item.trim();
if !trimmed.is_empty() {
flags.enable.push(trimmed.to_string());
}
}
flags.cli_enable = true;
i += 1;
}
}
"--cdp" => {
if let Some(s) = args.get(i + 1) {
flags.cdp = Some(s.clone());
@@ -581,6 +691,14 @@ pub fn parse_flags(args: &[String]) -> Flags {
i += 1;
}
}
"--hide-scrollbars" => {
let (val, consumed) = parse_bool_arg(args, i);
flags.hide_scrollbars = val;
flags.cli_hide_scrollbars = true;
if consumed {
i += 1;
}
}
"--device" => {
if let Some(d) = args.get(i + 1) {
flags.device = Some(d.clone());
@@ -590,10 +708,17 @@ pub fn parse_flags(args: &[String]) -> Flags {
"--auto-connect" => {
let (val, consumed) = parse_bool_arg(args, i);
flags.auto_connect = val;
if !val {
flags.force_launch = true;
}
if consumed {
i += 1;
}
}
"--launch" | "--new" => {
flags.force_launch = true;
flags.auto_connect = false;
}
"--session-name" => {
if let Some(s) = args.get(i + 1) {
flags.session_name = Some(s.clone());
@@ -715,6 +840,18 @@ pub fn parse_flags(args: &[String]) -> Flags {
i += 1;
}
}
"--model" => {
if let Some(s) = args.get(i + 1) {
flags.model = Some(s.clone());
i += 1;
}
}
"-v" | "--verbose" => {
flags.verbose = true;
}
"-q" | "--quiet" => {
flags.quiet = true;
}
"--config" => {
// Already handled by load_config(); skip the value
i += 1;
@@ -737,11 +874,22 @@ pub fn clean_args(args: &[String]) -> Vec<String> {
"--debug",
"--ignore-https-errors",
"--allow-file-access",
"--hide-scrollbars",
"--auto-connect",
"--launch",
"--new",
"--annotate",
"--content-boundaries",
"--confirm-interactive",
"--no-auto-dialog",
"-v",
"--verbose",
"-q",
"--quiet",
// doctor-specific flags; harmless on other commands (ignored)
"--offline",
"--quick",
"--fix",
];
// Global flags that always take a value (need to skip the next arg too)
const GLOBAL_FLAGS_WITH_VALUE: &[&str] = &[
@@ -750,6 +898,8 @@ pub fn clean_args(args: &[String]) -> Vec<String> {
"--executable-path",
"--cdp",
"--extension",
"--init-script",
"--enable",
"--profile",
"--state",
"--proxy",
@@ -772,6 +922,7 @@ pub fn clean_args(args: &[String]) -> Vec<String> {
"--screenshot-quality",
"--screenshot-format",
"--idle-timeout",
"--model",
];
let mut i = 0;
@@ -805,6 +956,7 @@ pub fn clean_args(args: &[String]) -> Vec<String> {
#[cfg(test)]
mod tests {
use super::*;
use crate::test_utils::EnvGuard;
fn args(s: &str) -> Vec<String> {
s.split_whitespace().map(String::from).collect()
@@ -1048,6 +1200,7 @@ mod tests {
"userAgent": "test-agent",
"provider": "ios",
"device": "iPhone 15",
"hideScrollbars": false,
"ignoreHttpsErrors": true,
"allowFileAccess": true,
"cdp": "9222",
@@ -1073,6 +1226,7 @@ mod tests {
assert_eq!(config.user_agent.as_deref(), Some("test-agent"));
assert_eq!(config.provider.as_deref(), Some("ios"));
assert_eq!(config.device.as_deref(), Some("iPhone 15"));
assert_eq!(config.hide_scrollbars, Some(false));
assert_eq!(config.ignore_https_errors, Some(true));
assert_eq!(config.allow_file_access, Some(true));
assert_eq!(config.cdp.as_deref(), Some("9222"));
@@ -1326,6 +1480,33 @@ mod tests {
assert!(flags.cli_allow_file_access);
}
#[test]
fn test_hide_scrollbars_default_true() {
let guard = EnvGuard::new(&["AGENT_BROWSER_HIDE_SCROLLBARS"]);
guard.remove("AGENT_BROWSER_HIDE_SCROLLBARS");
let flags = parse_flags(&args("open example.com"));
assert!(flags.hide_scrollbars);
assert!(!flags.cli_hide_scrollbars);
}
#[test]
fn test_hide_scrollbars_false() {
let guard = EnvGuard::new(&["AGENT_BROWSER_HIDE_SCROLLBARS"]);
guard.remove("AGENT_BROWSER_HIDE_SCROLLBARS");
let flags = parse_flags(&args("--hide-scrollbars false open"));
assert!(!flags.hide_scrollbars);
assert!(flags.cli_hide_scrollbars);
}
#[test]
fn test_hide_scrollbars_bare_defaults_true() {
let guard = EnvGuard::new(&["AGENT_BROWSER_HIDE_SCROLLBARS"]);
guard.remove("AGENT_BROWSER_HIDE_SCROLLBARS");
let flags = parse_flags(&args("--hide-scrollbars open"));
assert!(flags.hide_scrollbars);
assert!(flags.cli_hide_scrollbars);
}
#[test]
fn test_auto_connect_false() {
let flags = parse_flags(&args("--auto-connect false open"));
@@ -1334,7 +1515,9 @@ mod tests {
#[test]
fn test_clean_args_removes_bool_flag_with_value() {
let cleaned = clean_args(&args("--headed false --debug true open example.com"));
let cleaned = clean_args(&args(
"--headed false --debug true --hide-scrollbars false open example.com",
));
assert_eq!(cleaned, vec!["open", "example.com"]);
}
+277 -140
View File
@@ -183,9 +183,12 @@ fn platform_key() -> &'static str {
}
async fn fetch_download_url() -> Result<(String, String), String> {
let resp = reqwest::get(LAST_KNOWN_GOOD_URL)
let client = http_client()?;
let resp = client
.get(LAST_KNOWN_GOOD_URL)
.send()
.await
.map_err(|e| format!("Failed to fetch version info: {}", e))?;
.map_err(|e| format!("Failed to fetch version info: {}", format_reqwest_error(&e)))?;
let body: serde_json::Value = resp
.json()
@@ -223,44 +226,110 @@ async fn fetch_download_url() -> Result<(String, String), String> {
Ok((version, url))
}
fn format_reqwest_error(e: &reqwest::Error) -> String {
let mut msg = e.to_string();
let mut source = std::error::Error::source(e);
while let Some(cause) = source {
msg.push_str(&format!(": {}", cause));
source = std::error::Error::source(cause);
}
msg
}
fn http_client() -> Result<reqwest::Client, String> {
reqwest::Client::builder()
.user_agent(format!("agent-browser/{}", env!("CARGO_PKG_VERSION")))
.timeout(std::time::Duration::from_secs(120))
.connect_timeout(std::time::Duration::from_secs(30))
.build()
.map_err(|e| format!("Failed to create HTTP client: {}", format_reqwest_error(&e)))
}
async fn download_bytes(url: &str) -> Result<Vec<u8>, String> {
let resp = reqwest::get(url)
.await
.map_err(|e| format!("Download failed: {}", e))?;
let client = http_client()?;
let max_retries = 3;
let mut last_err = String::new();
let total = resp.content_length();
let mut bytes = Vec::new();
let mut stream = resp;
let mut downloaded: u64 = 0;
let mut last_pct: u64 = 0;
for attempt in 0..max_retries {
if attempt > 0 {
eprintln!(
" Retrying download (attempt {}/{})",
attempt + 1,
max_retries
);
tokio::time::sleep(std::time::Duration::from_secs(1 << attempt)).await;
}
loop {
let chunk = stream
.chunk()
.await
.map_err(|e| format!("Download error: {}", e))?;
match chunk {
Some(data) => {
downloaded += data.len() as u64;
bytes.extend_from_slice(&data);
let resp = match client.get(url).send().await {
Ok(r) => r,
Err(e) => {
last_err = format!("Download failed: {}", format_reqwest_error(&e));
if e.is_connect() || e.is_timeout() {
continue;
}
return Err(last_err);
}
};
if let Some(total) = total {
let pct = (downloaded * 100) / total;
if pct >= last_pct + 5 {
last_pct = pct;
let mb = downloaded as f64 / 1_048_576.0;
let total_mb = total as f64 / 1_048_576.0;
eprint!("\r {:.0}/{:.0} MB ({pct}%)", mb, total_mb);
let _ = io::stderr().flush();
let status = resp.status();
if !status.is_success() {
last_err = format!(
"Download failed: server returned HTTP {} for {}",
status, url
);
if status.is_server_error() {
continue;
}
return Err(last_err);
}
let total = resp.content_length();
let mut bytes = Vec::new();
let mut stream = resp;
let mut downloaded: u64 = 0;
let mut last_pct: u64 = 0;
let mut chunk_err = None;
loop {
let chunk = stream
.chunk()
.await
.map_err(|e| format!("Download error: {}", format_reqwest_error(&e)));
match chunk {
Ok(Some(data)) => {
downloaded += data.len() as u64;
bytes.extend_from_slice(&data);
if let Some(total) = total {
let pct = (downloaded * 100) / total;
if pct >= last_pct + 5 {
last_pct = pct;
let mb = downloaded as f64 / 1_048_576.0;
let total_mb = total as f64 / 1_048_576.0;
eprint!("\r {:.0}/{:.0} MB ({pct}%)", mb, total_mb);
let _ = io::stderr().flush();
}
}
}
Ok(None) => break,
Err(e) => {
chunk_err = Some(e);
break;
}
}
None => break,
}
eprintln!();
if let Some(e) = chunk_err {
last_err = e;
continue;
}
return Ok(bytes);
}
eprintln!();
Ok(bytes)
Err(last_err)
}
fn extract_zip(bytes: Vec<u8>, dest: &Path) -> Result<(), String> {
@@ -703,123 +772,191 @@ fn package_exists_apt(pkg: &str) -> bool {
.unwrap_or(false)
}
// ---------------------------------------------------------------------------
// Dashboard install
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::*;
use tokio::io::{AsyncReadExt, AsyncWriteExt};
use tokio::net::TcpListener;
pub fn get_dashboard_dir() -> PathBuf {
dirs::home_dir()
.unwrap_or_else(|| PathBuf::from("."))
.join(".agent-browser")
.join("dashboard")
}
const DASHBOARD_VERSION: &str = env!("CARGO_PKG_VERSION");
fn dashboard_download_url() -> String {
format!(
"https://github.com/vercel-labs/agent-browser/releases/download/v{}/dashboard.zip",
DASHBOARD_VERSION
)
}
pub fn run_dashboard_install() {
println!("{}", color::cyan("Installing dashboard..."));
let dest = get_dashboard_dir();
if dest.join("index.html").exists() {
println!(
"{} Dashboard is already installed at {}",
color::success_indicator(),
dest.display()
fn http_response(status: u16, reason: &str, body: &[u8]) -> Vec<u8> {
let header = format!(
"HTTP/1.1 {} {}\r\nContent-Length: {}\r\nConnection: close\r\n\r\n",
status,
reason,
body.len()
);
return;
let mut resp = header.into_bytes();
resp.extend_from_slice(body);
resp
}
let url = dashboard_download_url();
println!(" Downloading dashboard v{}", DASHBOARD_VERSION);
println!(" {}", url);
async fn accept_once(listener: &TcpListener, response: &[u8]) {
let (mut s, _) = listener.accept().await.unwrap();
let mut buf = [0u8; 4096];
let _ = s.read(&mut buf).await;
s.write_all(response).await.unwrap();
}
let rt = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.unwrap_or_else(|e| {
eprintln!(
"{} Failed to create runtime: {}",
color::error_indicator(),
e
);
exit(1);
async fn accept_with_ua_check(listener: &TcpListener, response: &[u8]) -> String {
let (mut s, _) = listener.accept().await.unwrap();
let mut buf = [0u8; 4096];
let n = s.read(&mut buf).await.unwrap();
let request = String::from_utf8_lossy(&buf[..n]).to_string();
s.write_all(response).await.unwrap();
request
}
#[tokio::test]
async fn download_bytes_returns_body_on_200() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let body = b"fake-zip-content";
let resp = http_response(200, "OK", body);
let server = tokio::spawn(async move {
accept_once(&listener, &resp).await;
});
let bytes = match rt.block_on(download_bytes(&url)) {
Ok(b) => b,
Err(e) => {
eprintln!("{} {}", color::error_indicator(), e);
eprintln!(" The dashboard may not be available for this version yet.");
eprintln!(" You can build it locally: cd packages/dashboard && pnpm build");
exit(1);
}
};
match extract_dashboard_zip(bytes, &dest) {
Ok(()) => {
println!(
"{} Dashboard v{} installed successfully",
color::success_indicator(),
DASHBOARD_VERSION
);
println!(" Location: {}", dest.display());
}
Err(e) => {
let _ = fs::remove_dir_all(&dest);
eprintln!("{} {}", color::error_indicator(), e);
exit(1);
}
}
}
fn extract_dashboard_zip(bytes: Vec<u8>, dest: &Path) -> Result<(), String> {
fs::create_dir_all(dest).map_err(|e| format!("Failed to create directory: {}", e))?;
let cursor = io::Cursor::new(bytes);
let mut archive =
zip::ZipArchive::new(cursor).map_err(|e| format!("Failed to read zip archive: {}", e))?;
for i in 0..archive.len() {
let mut file = archive
.by_index(i)
.map_err(|e| format!("Failed to read zip entry: {}", e))?;
let enclosed = match file.enclosed_name() {
Some(name) => name.to_owned(),
None => continue,
};
let rel_path = enclosed.to_string_lossy().to_string();
if rel_path.is_empty() || file.is_dir() {
if file.is_dir() {
let out_dir = dest.join(&rel_path);
let _ = fs::create_dir_all(&out_dir);
}
continue;
}
let out_path = dest.join(&rel_path);
if !out_path.starts_with(dest) {
continue;
}
if let Some(parent) = out_path.parent() {
fs::create_dir_all(parent)
.map_err(|e| format!("Failed to create parent dir {}: {}", parent.display(), e))?;
}
let mut out_file = fs::File::create(&out_path)
.map_err(|e| format!("Failed to create file {}: {}", out_path.display(), e))?;
io::copy(&mut file, &mut out_file)
.map_err(|e| format!("Failed to write {}: {}", out_path.display(), e))?;
let url = format!("http://127.0.0.1:{}/test.zip", port);
let result = download_bytes(&url).await;
assert!(result.is_ok());
assert_eq!(result.unwrap(), body);
server.await.unwrap();
}
Ok(())
#[tokio::test]
async fn download_bytes_returns_error_on_404() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let resp = http_response(404, "Not Found", b"not found");
let server = tokio::spawn(async move {
accept_once(&listener, &resp).await;
});
let url = format!("http://127.0.0.1:{}/test.zip", port);
let result = download_bytes(&url).await;
assert!(result.is_err());
let err = result.unwrap_err();
assert!(
err.contains("HTTP 404"),
"expected HTTP 404 in error, got: {}",
err
);
server.await.unwrap();
}
#[tokio::test]
async fn download_bytes_retries_on_500() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let server = tokio::spawn(async move {
// First two attempts: 500
let r500 = http_response(500, "Internal Server Error", b"error");
accept_once(&listener, &r500).await;
accept_once(&listener, &r500).await;
// Third attempt: 200
let r200 = http_response(200, "OK", b"ok-data");
accept_once(&listener, &r200).await;
});
let url = format!("http://127.0.0.1:{}/test.zip", port);
let result = download_bytes(&url).await;
assert!(
result.is_ok(),
"expected success after retries: {:?}",
result
);
assert_eq!(result.unwrap(), b"ok-data");
server.await.unwrap();
}
#[tokio::test]
async fn download_bytes_gives_up_after_max_retries() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let server = tokio::spawn(async move {
let r500 = http_response(500, "Internal Server Error", b"error");
// All 3 attempts get 500
accept_once(&listener, &r500).await;
accept_once(&listener, &r500).await;
accept_once(&listener, &r500).await;
});
let url = format!("http://127.0.0.1:{}/test.zip", port);
let result = download_bytes(&url).await;
assert!(result.is_err());
let err = result.unwrap_err();
assert!(
err.contains("HTTP 500"),
"expected HTTP 500 in error, got: {}",
err
);
server.await.unwrap();
}
#[tokio::test]
async fn download_bytes_does_not_retry_on_403() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let resp = http_response(403, "Forbidden", b"forbidden");
let server = tokio::spawn(async move {
// Only one request should arrive (no retries for 4xx)
accept_once(&listener, &resp).await;
});
let url = format!("http://127.0.0.1:{}/test.zip", port);
let result = download_bytes(&url).await;
assert!(result.is_err());
assert!(result.unwrap_err().contains("HTTP 403"));
server.await.unwrap();
}
#[tokio::test]
async fn http_client_sends_user_agent() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let resp = http_response(200, "OK", b"ok");
let server = tokio::spawn(async move {
let req = accept_with_ua_check(&listener, &resp).await;
req
});
let client = http_client().unwrap();
let url = format!("http://127.0.0.1:{}/test", port);
let _ = client.get(&url).send().await;
let request_text = server.await.unwrap();
let expected_ua = format!("agent-browser/{}", env!("CARGO_PKG_VERSION"));
assert!(
request_text.contains(&expected_ua),
"expected User-Agent '{}' in request:\n{}",
expected_ua,
request_text
);
}
#[test]
fn download_bytes_connection_refused_includes_details() {
// Use a port that nothing is listening on
let rt = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.unwrap();
let result = rt.block_on(download_bytes("http://127.0.0.1:1/test.zip"));
assert!(result.is_err());
let err = result.unwrap_err();
// The new code should include the root cause (connection refused)
// not just the vague "error sending request for url"
assert!(
err.contains("Connection refused")
|| err.contains("connection refused")
|| err.contains("actively refused it"),
"expected 'connection refused' in error, got: {}",
err
);
}
}
+246 -155
View File
@@ -1,10 +1,13 @@
mod chat;
mod color;
mod commands;
mod connection;
mod doctor;
mod flags;
mod install;
mod native;
mod output;
mod skills;
#[cfg(test)]
mod test_utils;
mod upgrade;
@@ -18,10 +21,13 @@ use std::process::exit;
#[cfg(windows)]
use windows_sys::Win32::Foundation::CloseHandle;
#[cfg(windows)]
use windows_sys::Win32::System::Threading::{OpenProcess, PROCESS_QUERY_LIMITED_INFORMATION};
use windows_sys::Win32::System::Threading::OpenProcess;
use commands::{gen_id, parse_command, ParseError};
use connection::{ensure_daemon, get_socket_dir, send_command, DaemonOptions};
use connection::{
cleanup_stale_files, ensure_daemon, get_socket_dir, is_pid_alive, send_command, walk_daemons,
DaemonOptions,
};
use flags::{clean_args, parse_flags, Flags};
use install::run_install;
use output::{
@@ -54,6 +60,23 @@ fn print_json_error_with_type(message: impl AsRef<str>, error_type: &str) {
}));
}
fn should_send_hide_scrollbars_launch_option(
cli_hide_scrollbars: bool,
hide_scrollbars: bool,
) -> bool {
cli_hide_scrollbars || !hide_scrollbars
}
fn apply_hide_scrollbars_launch_option(
launch_cmd: &mut serde_json::Value,
cli_hide_scrollbars: bool,
hide_scrollbars: bool,
) {
if should_send_hide_scrollbars_launch_option(cli_hide_scrollbars, hide_scrollbars) {
launch_cmd["hideScrollbars"] = json!(hide_scrollbars);
}
}
struct ParsedProxy {
server: String,
username: Option<String>,
@@ -117,51 +140,74 @@ fn parse_proxy(proxy_str: &str) -> ParsedProxy {
}
}
fn run_profiles(json_mode: bool) {
use crate::native::cdp::chrome::{find_chrome_user_data_dir, list_chrome_profiles};
let user_data_dir = match find_chrome_user_data_dir() {
Some(dir) => dir,
None => {
if json_mode {
print_json_error("No Chrome user data directory found");
} else {
eprintln!("{}", color::red("No Chrome user data directory found"));
}
exit(1);
}
};
let profiles = list_chrome_profiles(&user_data_dir);
if profiles.is_empty() {
if json_mode {
print_json_value(json!({
"success": true,
"data": []
}));
} else {
println!("No Chrome profiles found");
}
return;
}
if json_mode {
let items: Vec<serde_json::Value> = profiles
.iter()
.map(|p| {
json!({
"directory": p.directory,
"name": p.name
})
})
.collect();
print_json_value(json!({
"success": true,
"data": items
}));
} else {
println!(
"{} ({}):\n",
color::bold("Chrome profiles"),
user_data_dir.display()
);
for p in &profiles {
println!(
" {} {}",
color::bold(&p.directory),
color::dim(&format!("({})", p.name))
);
}
}
}
fn run_session(args: &[String], session: &str, json_mode: bool) {
let subcommand = args.get(1).map(|s| s.as_str());
match subcommand {
Some("list") => {
let socket_dir = get_socket_dir();
let mut sessions: Vec<String> = Vec::new();
if let Ok(entries) = fs::read_dir(&socket_dir) {
for entry in entries.flatten() {
let name = entry.file_name().to_string_lossy().to_string();
// Look for pid files in socket directory
if name.ends_with(".pid") {
let session_name = name.strip_suffix(".pid").unwrap_or("");
if !session_name.is_empty() {
// Check if session is actually running
let pid_path = socket_dir.join(&name);
if let Ok(pid_str) = fs::read_to_string(&pid_path) {
if let Ok(pid) = pid_str.trim().parse::<u32>() {
#[cfg(unix)]
let running = unsafe {
libc::kill(pid as i32, 0) == 0
|| std::io::Error::last_os_error().raw_os_error()
!= Some(libc::ESRCH)
};
#[cfg(windows)]
let running = unsafe {
let handle =
OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION, 0, pid);
if handle != 0 {
CloseHandle(handle);
true
} else {
false
}
};
if running {
sessions.push(session_name.to_string());
}
}
}
}
}
}
}
let sessions: Vec<String> = walk_daemons()
.sessions
.into_iter()
.map(|s| s.name)
.collect();
if json_mode {
println!(
@@ -202,25 +248,6 @@ fn get_dashboard_pid_path() -> std::path::PathBuf {
get_socket_dir().join("dashboard.pid")
}
fn is_pid_alive(pid: u32) -> bool {
#[cfg(unix)]
{
unsafe { libc::kill(pid as i32, 0) == 0 }
}
#[cfg(windows)]
{
unsafe {
let handle = OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION, 0, pid);
if handle != 0 {
CloseHandle(handle);
true
} else {
false
}
}
}
}
fn run_dashboard_start(port: u16, json_mode: bool) {
let pid_path = get_dashboard_pid_path();
@@ -379,43 +406,15 @@ fn run_dashboard_stop(json_mode: bool) {
}
fn run_close_all(flags: &Flags) {
let socket_dir = get_socket_dir();
let mut sessions: Vec<String> = Vec::new();
if let Ok(entries) = fs::read_dir(&socket_dir) {
for entry in entries.flatten() {
let name = entry.file_name().to_string_lossy().to_string();
if let Some(session_name) = name.strip_suffix(".pid") {
if session_name.is_empty() {
continue;
}
let pid_path = socket_dir.join(&name);
if let Ok(pid_str) = fs::read_to_string(&pid_path) {
if let Ok(pid) = pid_str.trim().parse::<u32>() {
#[cfg(unix)]
let running = unsafe {
libc::kill(pid as i32, 0) == 0
|| std::io::Error::last_os_error().raw_os_error()
!= Some(libc::ESRCH)
};
#[cfg(windows)]
let running = unsafe {
let handle = OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION, 0, pid);
if handle != 0 {
CloseHandle(handle);
true
} else {
false
}
};
if running {
sessions.push(session_name.to_string());
}
}
}
}
}
}
// walk_daemons auto-cleans stale .pid / .sock / .stream sidecar files and
// separates out the standalone dashboard. We only want to send `close` to
// real session daemons; the dashboard has its own `dashboard stop`.
let inventory = walk_daemons();
let sessions: Vec<(String, u32)> = inventory
.sessions
.iter()
.map(|s| (s.name.clone(), s.pid))
.collect();
if sessions.is_empty() {
if flags.json {
@@ -432,7 +431,7 @@ fn run_close_all(flags: &Flags) {
let mut closed: Vec<String> = Vec::new();
let mut failed: Vec<(String, String)> = Vec::new();
for session in &sessions {
for (session, pid) in &sessions {
let cmd = json!({ "id": gen_id(), "action": "close" });
match send_command(cmd, session) {
Ok(resp) if resp.success => closed.push(session.clone()),
@@ -440,7 +439,25 @@ fn run_close_all(flags: &Flags) {
let err = resp.error.unwrap_or_else(|| "Unknown error".to_string());
failed.push((session.clone(), err));
}
Err(e) => failed.push((session.clone(), e.to_string())),
Err(_) => {
// Daemon is unreachable despite its process existing.
// Force-kill the process and clean up stale files so future
// sessions are not poisoned.
#[cfg(unix)]
unsafe {
libc::kill(*pid as i32, libc::SIGKILL);
}
#[cfg(windows)]
unsafe {
let handle = OpenProcess(1, 0, *pid); // PROCESS_TERMINATE = 1
if handle != 0 {
windows_sys::Win32::System::Threading::TerminateProcess(handle, 1);
CloseHandle(handle);
}
}
cleanup_stale_files(session);
closed.push(session.clone());
}
}
}
@@ -511,7 +528,7 @@ fn main() {
}
let args: Vec<String> = env::args().skip(1).collect();
let flags = parse_flags(&args);
let mut flags = parse_flags(&args);
let clean = clean_args(&args);
let has_help = args.iter().any(|a| a == "--help" || a == "-h");
@@ -550,13 +567,21 @@ fn main() {
return;
}
// Handle doctor separately (doesn't need daemon; spawns its own scratch
// session for the live launch test).
if clean.first().map(|s| s.as_str()) == Some("doctor") {
let opts = doctor::DoctorOptions {
offline: args.iter().any(|a| a == "--offline"),
quick: args.iter().any(|a| a == "--quick"),
fix: args.iter().any(|a| a == "--fix"),
json: flags.json,
};
exit(doctor::run_doctor(opts));
}
// Handle dashboard subcommand
if clean.first().map(|s| s.as_str()) == Some("dashboard") {
match clean.get(1).map(|s| s.as_str()) {
Some("install") => {
install::run_dashboard_install();
return;
}
Some("start") | None => {
let port = clean
.iter()
@@ -582,6 +607,18 @@ fn main() {
}
}
// Handle profiles command (doesn't need daemon)
if clean.first().map(|s| s.as_str()) == Some("profiles") {
run_profiles(flags.json);
return;
}
// Handle skills command (doesn't need daemon)
if clean.first().map(|s| s.as_str()) == Some("skills") {
skills::run_skills(&clean, flags.json);
return;
}
// Handle session separately (doesn't need daemon)
if clean.first().map(|s| s.as_str()) == Some("session") {
run_session(&clean, &flags.session, flags.json);
@@ -598,6 +635,17 @@ fn main() {
return;
}
// Handle chat command
if clean.first().map(|s| s.as_str()) == Some("chat") {
let message = if clean.len() > 1 {
Some(clean[1..].join(" "))
} else {
None
};
chat::run_chat(&flags, message);
return;
}
let mut cmd = match parse_command(&clean, &flags) {
Ok(c) => c,
Err(e) => {
@@ -700,6 +748,8 @@ fn main() {
debug: flags.debug,
executable_path: flags.executable_path.as_deref(),
extensions: &flags.extensions,
init_scripts: &flags.init_scripts,
enable: &flags.enable,
args: flags.args.as_deref(),
user_agent: flags.user_agent.as_deref(),
proxy: proxy_server.as_deref(),
@@ -708,6 +758,7 @@ fn main() {
proxy_password: proxy_password.as_deref(),
ignore_https_errors: flags.ignore_https_errors,
allow_file_access: flags.allow_file_access,
hide_scrollbars: flags.hide_scrollbars,
profile: flags.profile.as_deref(),
state: flags.state.as_deref(),
provider: flags.provider.as_deref(),
@@ -719,7 +770,9 @@ fn main() {
confirm_actions: flags.confirm_actions.as_deref(),
engine: flags.engine.as_deref(),
auto_connect: flags.auto_connect,
force_launch: flags.force_launch,
idle_timeout: flags.idle_timeout.as_deref(),
default_timeout: flags.default_timeout,
cdp: flags.cdp.as_deref(),
no_auto_dialog: flags.no_auto_dialog,
};
@@ -779,6 +832,7 @@ fn main() {
},
flags.ignore_https_errors.then_some("--ignore-https-errors"),
flags.cli_allow_file_access.then_some("--allow-file-access"),
flags.cli_hide_scrollbars.then_some("--hide-scrollbars"),
flags.cli_download_path.then_some("--download-path"),
flags.cli_headed.then_some("--headed"),
]
@@ -787,11 +841,24 @@ fn main() {
.collect();
if !ignored_flags.is_empty() && !flags.json {
eprintln!(
"{} {} ignored: daemon already running. Use 'agent-browser close' first to restart with new options.",
color::warning_indicator(),
ignored_flags.join(", ")
);
// Special case: --headed is irrelevant in CDP-attach mode
// (your existing Chrome is always already visible). The
// "agent-browser close + reopen" advice doesn't help because
// the new daemon will attach right back to the same Chrome.
// Don't suggest a useless workaround.
if ignored_flags == ["--headed"] {
eprintln!(
"{} --headed has no effect when attached to your running Chrome (it's already visible). \
Pass --launch to spawn a separate browser if you need to control headedness.",
color::warning_indicator(),
);
} else {
eprintln!(
"{} {} ignored: daemon already running. Use 'agent-browser close' first to restart with new options.",
color::warning_indicator(),
ignored_flags.join(", ")
);
}
}
}
@@ -806,24 +873,9 @@ fn main() {
exit(1);
}
if flags.auto_connect && flags.cdp.is_some() {
let msg = "Cannot use --auto-connect and --cdp together";
if flags.json {
print_json_error(msg);
} else {
eprintln!("{} {}", color::error_indicator(), msg);
}
exit(1);
}
if flags.auto_connect && flags.provider.is_some() {
let msg = "Cannot use --auto-connect and -p/--provider together";
if flags.json {
print_json_error(msg);
} else {
eprintln!("{} {}", color::error_indicator(), msg);
}
exit(1);
// Explicit --cdp or --provider disables auto-connect (they specify the connection)
if flags.cdp.is_some() || flags.provider.is_some() {
flags.auto_connect = false;
}
if flags.provider.is_some() && !flags.extensions.is_empty() {
@@ -1029,13 +1081,17 @@ fn main() {
|| flags.args.is_some()
|| flags.user_agent.is_some()
|| flags.allow_file_access
|| should_send_hide_scrollbars_launch_option(
flags.cli_hide_scrollbars,
flags.hide_scrollbars,
)
|| flags.color_scheme.is_some()
|| flags.download_path.is_some()
|| flags.engine.is_some()
|| !flags.extensions.is_empty())
&& flags.cdp.is_none()
&& flags.provider.is_none()
&& !flags.auto_connect
&& (flags.force_launch || !flags.auto_connect)
{
let mut launch_cmd = json!({
"id": gen_id(),
@@ -1103,6 +1159,12 @@ fn main() {
launch_cmd["allowFileAccess"] = json!(true);
}
apply_hide_scrollbars_launch_option(
&mut launch_cmd,
flags.cli_hide_scrollbars,
flags.hide_scrollbars,
);
if let Some(ref cs) = flags.color_scheme {
launch_cmd["colorScheme"] = json!(cs);
}
@@ -1150,10 +1212,16 @@ fn main() {
}
}
// Handle batch command: read commands from stdin, execute sequentially
// Handle batch command: from args or stdin
if cmd.get("action").and_then(|v| v.as_str()) == Some("batch") {
let bail = cmd.get("bail").and_then(|v| v.as_bool()).unwrap_or(false);
run_batch(&flags, bail);
let arg_commands = cmd.get("commands").and_then(|v| v.as_array()).map(|arr| {
arr.iter()
.filter_map(|v| v.as_str())
.map(commands::shell_words_split)
.collect::<Vec<Vec<String>>>()
});
run_batch(&flags, bail, arg_commands);
return;
}
@@ -1233,36 +1301,40 @@ fn main() {
}
}
fn run_batch(flags: &Flags, bail: bool) {
use std::io::Read as _;
fn run_batch(flags: &Flags, bail: bool, arg_commands: Option<Vec<Vec<String>>>) {
let commands: Vec<Vec<String>> = if let Some(cmds) = arg_commands {
cmds
} else {
use std::io::Read as _;
let mut input = String::new();
if let Err(e) = std::io::stdin().read_to_string(&mut input) {
if flags.json {
print_json_error(format!("Failed to read stdin: {}", e));
} else {
eprintln!("{} Failed to read stdin: {}", color::error_indicator(), e);
}
exit(1);
}
let commands: Vec<Vec<String>> = match serde_json::from_str(&input) {
Ok(c) => c,
Err(e) => {
let mut input = String::new();
if let Err(e) = std::io::stdin().read_to_string(&mut input) {
if flags.json {
print_json_error(format!(
"Invalid JSON input: {}. Expected an array of string arrays, e.g. [[\"open\", \"https://example.com\"], [\"snapshot\"]]",
e
));
print_json_error(format!("Failed to read stdin: {}", e));
} else {
eprintln!(
"{} Invalid JSON input: {}. Expected an array of string arrays.",
color::error_indicator(),
e
);
eprintln!("{} Failed to read stdin: {}", color::error_indicator(), e);
}
exit(1);
}
match serde_json::from_str(&input) {
Ok(c) => c,
Err(e) => {
if flags.json {
print_json_error(format!(
"Invalid JSON input: {}. Expected an array of string arrays, e.g. [[\"open\", \"https://example.com\"], [\"snapshot\"]]",
e
));
} else {
eprintln!(
"{} Invalid JSON input: {}. Expected an array of string arrays.",
color::error_indicator(),
e
);
}
exit(1);
}
}
};
if commands.is_empty() {
@@ -1445,4 +1517,23 @@ mod tests {
"Daemon process exited during startup:\nline \"quoted\"\u{001b}[2mansi\u{001b}[22m"
);
}
#[test]
fn test_hide_scrollbars_launch_option_serialization() {
assert!(!should_send_hide_scrollbars_launch_option(false, true));
assert!(should_send_hide_scrollbars_launch_option(false, false));
assert!(should_send_hide_scrollbars_launch_option(true, true));
let mut default_cmd = json!({ "action": "launch" });
apply_hide_scrollbars_launch_option(&mut default_cmd, false, true);
assert!(default_cmd.get("hideScrollbars").is_none());
let mut config_false_cmd = json!({ "action": "launch" });
apply_hide_scrollbars_launch_option(&mut config_false_cmd, false, false);
assert_eq!(config_false_cmd["hideScrollbars"], false);
let mut cli_true_cmd = json!({ "action": "launch" });
apply_hide_scrollbars_launch_option(&mut cli_true_cmd, true, true);
assert_eq!(cli_true_cmd["hideScrollbars"], true);
}
}
+1104 -153
View File
File diff suppressed because it is too large Load Diff
+495 -64
View File
@@ -1,5 +1,5 @@
use serde_json::{json, Value};
use std::collections::HashSet;
use std::collections::{HashMap, HashSet};
use std::future::Future;
use std::sync::Arc;
use std::time::{Duration, Instant};
@@ -10,6 +10,7 @@ use super::cdp::client::CdpClient;
use super::cdp::discovery::discover_cdp_url;
use super::cdp::lightpanda::{launch_lightpanda, LightpandaLaunchOptions, LightpandaProcess};
use super::cdp::types::*;
use super::element::{resolve_element_object_id, RefMap};
// ---------------------------------------------------------------------------
// Launch validation
@@ -110,6 +111,26 @@ fn update_page_target_info_in_pages(pages: &mut [PageInfo], target: &TargetInfo)
false
}
fn active_page_index_after_removal(
active_page_index: usize,
removed_index: usize,
remaining_pages: usize,
) -> usize {
if remaining_pages == 0 {
return 0;
}
if removed_index < active_page_index {
return active_page_index - 1;
}
if active_page_index >= remaining_pages {
return remaining_pages - 1;
}
active_page_index
}
/// Converts common error messages into AI-friendly, actionable descriptions.
pub fn to_ai_friendly_error(error: &str) -> String {
let lower = error.to_lowercase();
@@ -137,6 +158,13 @@ pub fn to_ai_friendly_error(error: &str) -> String {
#[derive(Debug, Clone)]
pub struct PageInfo {
pub tab_id: u32,
/// Optional user-assigned label (e.g. "docs", "app"). Set via
/// `tab new --label <name>`. Labels are agent-assigned and never
/// auto-generated, never rewritten on navigation, and unique within a
/// session. Agents use labels instead of `t<N>` for readable multi-tab
/// workflows.
pub label: Option<String>,
pub target_id: String,
pub session_id: String,
pub url: String,
@@ -144,6 +172,77 @@ pub struct PageInfo {
pub target_type: String, // "page" or "webview"
}
/// Canonical string form of a stable tab id: `t1`, `t2`, ... The `t` prefix
/// disambiguates stable ids from positional indices (which the CLI no longer
/// accepts) and matches the `@e<N>` convention used for element refs.
pub fn format_tab_id(tab_id: u32) -> String {
format!("t{}", tab_id)
}
/// A tab reference as parsed from CLI/JSON input. Either a stable id like
/// `t2` or a user-assigned label like `docs`.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum TabRef {
Id(u32),
Label(String),
}
impl TabRef {
/// Parse a user-supplied string tab reference. Rejects bare integers
/// with a teaching error so agents and scripts don't silently confuse
/// stable ids with positional indices.
pub fn parse(input: &str) -> Result<Self, String> {
let input = input.trim();
if input.is_empty() {
return Err("Empty tab reference; expected `t<N>` (e.g. `t2`) or a label".to_string());
}
if let Some(digits) = input.strip_prefix('t').or_else(|| input.strip_prefix('T')) {
if !digits.is_empty() && digits.chars().all(|c| c.is_ascii_digit()) {
let id: u32 = digits.parse().map_err(|_| {
format!(
"Tab id `{}` out of range; ids are incrementing positive integers",
input
)
})?;
if id == 0 {
return Err(format!(
"Tab id `{}` is invalid; tab ids start at t1",
input
));
}
return Ok(TabRef::Id(id));
}
}
if input.chars().all(|c| c.is_ascii_digit()) {
return Err(format!(
"Expected a tab id like `t{}` or a label; positional integers are not accepted \
(run `agent-browser tab` to list stable tab ids)",
input
));
}
if !is_valid_label(input) {
return Err(format!(
"Invalid tab label `{}`; labels must start with a letter and contain only \
letters, digits, `-`, and `_`",
input
));
}
Ok(TabRef::Label(input.to_string()))
}
}
/// Labels must look like identifiers: start with a letter, contain only
/// letters/digits/dashes/underscores. This keeps them distinguishable from
/// `t<N>` ids at a glance and safe to pass through shells without quoting.
pub fn is_valid_label(s: &str) -> bool {
let mut chars = s.chars();
match chars.next() {
Some(c) if c.is_ascii_alphabetic() => {}
_ => return false,
}
chars.all(|c| c.is_ascii_alphanumeric() || c == '-' || c == '_')
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum WaitUntil {
Load,
@@ -201,14 +300,53 @@ pub struct BrowserManager {
default_timeout_ms: u64,
/// Stored download path from launch options, re-applied to new contexts (e.g., recording)
pub download_path: Option<String>,
/// Whether to ignore HTTPS certificate errors, re-applied to new contexts (e.g., recording)
pub ignore_https_errors: bool,
/// Origins visited during this session, used by save_state to collect cross-origin localStorage.
visited_origins: HashSet<String>,
next_tab_id: u32,
}
const LIGHTPANDA_CDP_CONNECT_TIMEOUT: Duration = Duration::from_secs(5);
const LIGHTPANDA_CDP_CONNECT_POLL_INTERVAL: Duration = Duration::from_millis(100);
const LIGHTPANDA_TARGET_INIT_TIMEOUT: Duration = Duration::from_secs(10);
/// Outcome of a single `Browser.getVersion` liveness probe.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum LivenessProbe {
/// Chrome answered — the connection is definitely alive.
Responded,
/// The CDP transport errored (WebSocket closed/reset) — the socket is gone.
TransportError,
/// The probe timed out with no response.
TimedOut,
}
/// Decide whether a CDP connection should be considered alive from one probe.
///
/// The subtle case is [`LivenessProbe::TimedOut`]. For a browser we launched
/// ourselves (`is_external_attach == false`) a hung CDP socket is a real
/// problem and the daemon should reconnect. But for an *externally attached*
/// browser — the stealth fork's default, where we attach to the user's real
/// Chrome — a slow/no response is almost always Chrome being briefly busy or,
/// critically, showing the Chrome 136+ "Allow remote debugging?" consent modal,
/// which blocks CDP responses until the user clicks Allow.
///
/// Treating that timeout as "dead" tears down the already-consented connection
/// and forces a reconnect, which re-pops the consent prompt; repeated on every
/// command it produces an endless prompt loop and a connection storm that can
/// freeze Chrome. So for external attaches we keep the connection alive on
/// timeout. A genuinely dead external socket instead surfaces as
/// [`LivenessProbe::TransportError`] (and Chrome being closed by the user is a
/// transport error, not a timeout), so zombie-socket detection is preserved.
fn connection_alive_from_probe(probe: LivenessProbe, is_external_attach: bool) -> bool {
match probe {
LivenessProbe::Responded => true,
LivenessProbe::TransportError => false,
LivenessProbe::TimedOut => is_external_attach,
}
}
impl BrowserManager {
pub async fn launch(options: LaunchOptions, engine: Option<&str>) -> Result<Self, String> {
let engine = engine.unwrap_or("chrome");
@@ -272,7 +410,9 @@ impl BrowserManager {
active_page_index: 0,
default_timeout_ms: 25_000,
download_path: download_path.clone(),
ignore_https_errors,
visited_origins: HashSet::new(),
next_tab_id: 1,
};
manager.discover_and_attach_targets().await?;
manager
@@ -359,11 +499,16 @@ impl BrowserManager {
active_page_index: 0,
default_timeout_ms: 25_000,
download_path: None,
ignore_https_errors: false,
visited_origins: HashSet::new(),
next_tab_id: 1,
};
if direct_page {
let tab_id = manager.assign_tab_id();
manager.pages.push(PageInfo {
tab_id,
label: None,
target_id: "provider-page".to_string(),
session_id: String::new(),
url: String::new(),
@@ -428,7 +573,11 @@ impl BrowserManager {
)
.await?;
let tab_id = self.next_tab_id;
self.next_tab_id += 1;
self.pages.push(PageInfo {
tab_id,
label: None,
target_id: result.target_id,
session_id: attach_result.session_id.clone(),
url: "about:blank".to_string(),
@@ -451,7 +600,11 @@ impl BrowserManager {
)
.await?;
let tab_id = self.next_tab_id;
self.next_tab_id += 1;
self.pages.push(PageInfo {
tab_id,
label: None,
target_id: target.target_id.clone(),
session_id: attach_result.session_id.clone(),
url: target.url.clone(),
@@ -479,6 +632,13 @@ impl BrowserManager {
self.client
.send_command_no_params("Runtime.enable", Some(session_id))
.await?;
// Resume the target if it is paused waiting for the debugger.
// This is needed for real browser sessions (Chrome 144+) where targets
// are paused after attach until explicitly resumed. No-op otherwise.
let _ = self
.client
.send_command_no_params("Runtime.runIfWaitingForDebugger", Some(session_id))
.await;
self.client
.send_command_no_params("Network.enable", Some(session_id))
.await?;
@@ -508,6 +668,10 @@ impl BrowserManager {
self.client
.send_command_no_params("Runtime.enable", None)
.await?;
let _ = self
.client
.send_command_no_params("Runtime.runIfWaitingForDebugger", None)
.await;
self.client
.send_command_no_params("Network.enable", None)
.await?;
@@ -701,21 +865,27 @@ impl BrowserManager {
self.default_timeout_ms
}
/// Checks if the CDP connection is alive by sending a simple command.
/// Returns false if the command times out or fails.
/// Checks if the CDP connection is alive by sending a `Browser.getVersion`
/// probe. See [`connection_alive_from_probe`] for how the outcome maps to a
/// liveness verdict — in particular why a timeout does NOT tear down an
/// externally-attached browser.
pub async fn is_connection_alive(&self) -> bool {
let timeout = tokio::time::Duration::from_secs(3);
let result = tokio::time::timeout(
let probe = match tokio::time::timeout(
timeout,
self.client
.send_command_no_params("Browser.getVersion", None),
)
.await;
match result {
Ok(Ok(_)) => true,
Ok(Err(_)) | Err(_) => false,
}
.await
{
Ok(Ok(_)) => LivenessProbe::Responded,
Ok(Err(_)) => LivenessProbe::TransportError,
Err(_) => LivenessProbe::TimedOut,
};
// No child process => we attached to an external browser (the user's
// real Chrome — the stealth fork's default).
let is_external_attach = self.browser_process.is_none();
connection_alive_from_probe(probe, is_external_attach)
}
/// Non-blocking check whether the locally-launched browser process has exited
@@ -785,7 +955,11 @@ impl BrowserManager {
)
.await?;
let tab_id = self.next_tab_id;
self.next_tab_id += 1;
self.pages.push(PageInfo {
tab_id,
label: None,
target_id: result.target_id,
session_id: attach_result.session_id.clone(),
url: "about:blank".to_string(),
@@ -814,13 +988,22 @@ impl BrowserManager {
}
}
fn update_active_page_after_removal(&mut self, removed_index: usize) {
self.active_page_index = active_page_index_after_removal(
self.active_page_index,
removed_index,
self.pages.len(),
);
}
pub fn tab_list(&self) -> Vec<Value> {
self.pages
.iter()
.enumerate()
.map(|(i, p)| {
json!({
"index": i,
"tabId": format_tab_id(p.tab_id),
"label": p.label,
"title": p.title,
"url": p.url,
"type": p.target_type,
@@ -830,7 +1013,61 @@ impl BrowserManager {
.collect()
}
pub async fn tab_new(&mut self, url: Option<&str>) -> Result<Value, String> {
/// Resolve a user-supplied `TabRef` (either `t<N>` or a label) to the
/// stable numeric `tab_id`. Returns a teaching error for unknown tabs.
pub fn resolve_tab_ref(&self, tab_ref: &TabRef) -> Result<u32, String> {
match tab_ref {
TabRef::Id(id) => {
if self.has_tab_id(*id) {
Ok(*id)
} else {
Err(format!(
"Tab {} not found; run `agent-browser tab` to list open tabs",
format_tab_id(*id)
))
}
}
TabRef::Label(name) => self
.pages
.iter()
.find(|p| p.label.as_deref() == Some(name.as_str()))
.map(|p| p.tab_id)
.ok_or_else(|| {
format!(
"No tab with label `{}`; run `agent-browser tab` to list open tabs",
name
)
}),
}
}
/// Returns true iff a tab already carries the given label.
pub fn has_label(&self, label: &str) -> bool {
self.pages.iter().any(|p| p.label.as_deref() == Some(label))
}
pub async fn tab_new(
&mut self,
url: Option<&str>,
label: Option<&str>,
) -> Result<Value, String> {
if let Some(label) = label {
if !is_valid_label(label) {
return Err(format!(
"Invalid tab label `{}`; labels must start with a letter and contain only \
letters, digits, `-`, and `_`",
label
));
}
if self.has_label(label) {
return Err(format!(
"Label `{}` is already used by another tab; labels must be unique within a \
session",
label
));
}
}
let target_url = url.unwrap_or("about:blank");
let result: CreateTargetResult = self
@@ -858,8 +1095,13 @@ impl BrowserManager {
self.enable_domains(&attach.session_id).await?;
let tab_id = self.next_tab_id;
self.next_tab_id += 1;
let index = self.pages.len();
let label = label.map(|s| s.to_string());
self.pages.push(PageInfo {
tab_id,
label: label.clone(),
target_id: result.target_id,
session_id: attach.session_id,
url: target_url.to_string(),
@@ -868,7 +1110,12 @@ impl BrowserManager {
});
self.active_page_index = index;
Ok(json!({ "index": index, "url": target_url }))
Ok(json!({
"tabId": format_tab_id(tab_id),
"label": label,
"url": target_url,
"total": self.pages.len(),
}))
}
pub async fn tab_switch(&mut self, index: usize) -> Result<Value, String> {
@@ -898,7 +1145,13 @@ impl BrowserManager {
page.title = title.clone();
}
Ok(json!({ "index": index, "url": url, "title": title }))
let page = &self.pages[index];
Ok(json!({
"tabId": format_tab_id(page.tab_id),
"label": page.label,
"url": url,
"title": title,
}))
}
pub async fn tab_close(&mut self, index: Option<usize>) -> Result<Value, String> {
@@ -913,6 +1166,9 @@ impl BrowserManager {
}
let page = self.pages.remove(target_index);
self.update_active_page_after_removal(target_index);
let closed_tab_id = page.tab_id;
let closed_label = page.label.clone();
let _ = self
.client
.send_command_typed::<_, Value>(
@@ -924,14 +1180,14 @@ impl BrowserManager {
)
.await;
if self.active_page_index >= self.pages.len() {
self.active_page_index = self.pages.len() - 1;
}
let session_id = self.pages[self.active_page_index].session_id.clone();
self.enable_domains(&session_id).await?;
Ok(json!({ "closed": target_index, "activeIndex": self.active_page_index }))
Ok(json!({
"tabId": format_tab_id(closed_tab_id),
"label": closed_label,
"closed": true,
}))
}
// -----------------------------------------------------------------------
@@ -958,6 +1214,39 @@ impl BrowserManager {
Some(session_id),
)
.await?;
// Screencast captures the actual content area, not the emulated CSS
// viewport, so resize the content area to match.
if let Ok(target_id) = self.active_target_id() {
if let Ok(window_info) = self
.client
.send_command(
"Browser.getWindowForTarget",
Some(json!({ "targetId": target_id })),
None,
)
.await
{
if let Some(window_id) = window_info.get("windowId").and_then(|v| v.as_i64()) {
if let Err(e) = self
.client
.send_command(
"Browser.setContentsSize",
Some(json!({
"windowId": window_id,
"width": width,
"height": height,
})),
None,
)
.await
{
eprintln!("Browser.setContentsSize failed (experimental CDP): {e}");
}
}
}
}
Ok(())
}
@@ -1080,50 +1369,25 @@ impl BrowserManager {
Ok(())
}
pub async fn upload_files(&self, selector: &str, files: &[String]) -> Result<(), String> {
pub async fn upload_files(
&self,
selector: &str,
files: &[String],
ref_map: &RefMap,
iframe_sessions: &HashMap<String, String>,
) -> Result<(), String> {
let session_id = self.active_session_id()?;
let node_result = self
.client
.send_command(
"DOM.querySelector",
Some(json!({
"nodeId": 1,
"selector": selector,
})),
Some(session_id),
)
.await;
let (object_id, effective_session_id) =
resolve_element_object_id(&self.client, session_id, ref_map, selector, iframe_sessions)
.await?;
// Alternative: resolve via JS
let result: EvaluateResult = self
.client
.send_command_typed(
"Runtime.evaluate",
&EvaluateParams {
expression: format!(
"document.querySelector({})",
serde_json::to_string(selector).unwrap_or_default()
),
return_by_value: Some(false),
await_promise: Some(false),
},
Some(session_id),
)
.await?;
let object_id = result
.result
.object_id
.ok_or("File input element not found")?;
// Get the DOM node from the remote object
let describe: Value = self
.client
.send_command(
"DOM.describeNode",
Some(json!({ "objectId": object_id })),
Some(session_id),
Some(&effective_session_id),
)
.await?;
@@ -1133,9 +1397,6 @@ impl BrowserManager {
.and_then(|v| v.as_i64())
.ok_or("Could not get backendNodeId for file input")?;
// Suppress unused variable warning
let _ = node_result;
self.client
.send_command(
"DOM.setFileInputFiles",
@@ -1143,7 +1404,7 @@ impl BrowserManager {
"files": files,
"backendNodeId": backend_node_id,
})),
Some(session_id),
Some(&effective_session_id),
)
.await?;
@@ -1167,6 +1428,46 @@ impl BrowserManager {
.to_string())
}
pub async fn remove_script_to_evaluate(&self, identifier: &str) -> Result<(), String> {
let session_id = self.active_session_id()?;
self.client
.send_command(
"Page.removeScriptToEvaluateOnNewDocument",
Some(json!({ "identifier": identifier })),
Some(session_id),
)
.await?;
Ok(())
}
pub async fn tab_switch_by_id(&mut self, tab_id: u32) -> Result<Value, String> {
let index = self
.pages
.iter()
.position(|p| p.tab_id == tab_id)
.ok_or_else(|| format!("Tab ID {} not found", tab_id))?;
self.tab_switch(index).await
}
pub async fn tab_close_by_id(&mut self, tab_id: Option<u32>) -> Result<Value, String> {
let index = match tab_id {
Some(id) => Some(
self.pages
.iter()
.position(|p| p.tab_id == id)
.ok_or_else(|| format!("Tab ID {} not found", id))?,
),
None => None,
};
self.tab_close(index).await
}
pub fn assign_tab_id(&mut self) -> u32 {
let id = self.next_tab_id;
self.next_tab_id += 1;
id
}
pub fn add_page(&mut self, page: PageInfo) {
let index = self.pages.len();
self.pages.push(page);
@@ -1180,7 +1481,7 @@ impl BrowserManager {
pub fn remove_page_by_target_id(&mut self, target_id: &str) {
if let Some(pos) = self.pages.iter().position(|p| p.target_id == target_id) {
self.pages.remove(pos);
self.update_active_page_if_needed();
self.update_active_page_after_removal(pos);
}
}
@@ -1192,6 +1493,16 @@ impl BrowserManager {
self.pages.len()
}
/// Returns the stable `tab_id` of the currently active page, if any.
pub fn active_tab_id(&self) -> Option<u32> {
self.pages.get(self.active_page_index).map(|p| p.tab_id)
}
/// Returns true if a tab with the given stable `tab_id` is still open.
pub fn has_tab_id(&self, tab_id: u32) -> bool {
self.pages.iter().any(|p| p.tab_id == tab_id)
}
pub fn pages_list(&self) -> Vec<PageInfo> {
self.pages.clone()
}
@@ -1256,10 +1567,8 @@ async fn poll_network_idle(
}
}
}
"Page.loadEventFired" => {
if p.is_empty() {
idle_start = Some(tokio::time::Instant::now());
}
"Page.loadEventFired" if p.is_empty() => {
idle_start = Some(tokio::time::Instant::now());
}
_ => {}
}
@@ -1347,7 +1656,9 @@ async fn initialize_lightpanda_manager(
active_page_index: 0,
default_timeout_ms: 25_000,
download_path: None,
ignore_https_errors: false,
visited_origins: HashSet::new(),
next_tab_id: 1,
};
match discover_and_attach_lightpanda_targets(&mut manager, deadline).await {
@@ -1453,6 +1764,104 @@ mod tests {
use super::*;
use tokio::time::sleep;
#[test]
fn test_format_tab_id() {
assert_eq!(format_tab_id(1), "t1");
assert_eq!(format_tab_id(42), "t42");
}
#[test]
fn liveness_responded_is_alive_for_both_kinds() {
assert!(connection_alive_from_probe(LivenessProbe::Responded, true));
assert!(connection_alive_from_probe(LivenessProbe::Responded, false));
}
#[test]
fn liveness_transport_error_is_dead_for_both_kinds() {
// A closed/reset WebSocket is a genuine death — reconnect in both cases.
assert!(!connection_alive_from_probe(LivenessProbe::TransportError, true));
assert!(!connection_alive_from_probe(LivenessProbe::TransportError, false));
}
#[test]
fn liveness_timeout_keeps_external_attach_alive() {
// Regression guard for the remote-debugging consent storm: a timed-out
// probe must NOT tear down an externally-attached browser, otherwise the
// daemon reconnects and re-pops Chrome's "Allow remote debugging?" modal
// on every command (endless prompts + browser freeze).
assert!(connection_alive_from_probe(LivenessProbe::TimedOut, true));
}
#[test]
fn liveness_timeout_marks_launched_browser_dead() {
// A browser we launched that stops responding is a real problem worth a
// reconnect (and has no consent modal to worry about).
assert!(!connection_alive_from_probe(LivenessProbe::TimedOut, false));
}
#[test]
fn test_parse_tab_ref_id() {
assert_eq!(TabRef::parse("t1"), Ok(TabRef::Id(1)));
assert_eq!(TabRef::parse("t42"), Ok(TabRef::Id(42)));
assert_eq!(TabRef::parse("T7"), Ok(TabRef::Id(7)));
}
#[test]
fn test_parse_tab_ref_label() {
assert_eq!(TabRef::parse("docs"), Ok(TabRef::Label("docs".to_string())));
assert_eq!(
TabRef::parse("app-2"),
Ok(TabRef::Label("app-2".to_string()))
);
assert_eq!(
TabRef::parse("my_tab"),
Ok(TabRef::Label("my_tab".to_string()))
);
}
#[test]
fn test_parse_tab_ref_rejects_bare_integer() {
let err = TabRef::parse("2").unwrap_err();
assert!(
err.contains("positional integers are not accepted"),
"error should teach the user to use `t<N>`: {}",
err
);
assert!(err.contains("t2"));
}
#[test]
fn test_parse_tab_ref_rejects_empty() {
assert!(TabRef::parse("").is_err());
assert!(TabRef::parse(" ").is_err());
}
#[test]
fn test_parse_tab_ref_rejects_zero() {
let err = TabRef::parse("t0").unwrap_err();
assert!(err.contains("start at t1"));
}
#[test]
fn test_parse_tab_ref_rejects_invalid_label() {
assert!(TabRef::parse("2docs").is_err());
assert!(TabRef::parse("-docs").is_err());
assert!(TabRef::parse("docs!").is_err());
assert!(TabRef::parse("docs space").is_err());
}
#[test]
fn test_is_valid_label() {
assert!(is_valid_label("docs"));
assert!(is_valid_label("Docs"));
assert!(is_valid_label("app-2"));
assert!(is_valid_label("my_tab"));
assert!(!is_valid_label(""));
assert!(!is_valid_label("2docs"));
assert!(!is_valid_label("-docs"));
assert!(!is_valid_label("docs!"));
}
#[test]
fn test_should_track_popup_target_with_empty_url() {
let target = TargetInfo {
@@ -1484,6 +1893,8 @@ mod tests {
#[test]
fn test_update_page_target_info_in_pages_updates_existing_page() {
let mut pages = vec![PageInfo {
tab_id: 1,
label: None,
target_id: "popup-1".to_string(),
session_id: "session-1".to_string(),
url: String::new(),
@@ -1504,6 +1915,26 @@ mod tests {
assert_eq!(pages[0].title, "Popup");
}
#[test]
fn test_active_page_index_after_removal_shifts_when_earlier_tab_is_removed() {
assert_eq!(active_page_index_after_removal(2, 0, 3), 1);
}
#[test]
fn test_active_page_index_after_removal_keeps_same_slot_when_later_tab_is_removed() {
assert_eq!(active_page_index_after_removal(1, 2, 3), 1);
}
#[test]
fn test_active_page_index_after_removal_clamps_when_active_last_tab_is_removed() {
assert_eq!(active_page_index_after_removal(3, 3, 3), 2);
}
#[test]
fn test_active_page_index_after_removal_resets_when_last_page_disappears() {
assert_eq!(active_page_index_after_removal(0, 0, 0), 0);
}
#[test]
fn test_validate_launch_options_extensions_and_cdp() {
let ext = vec!["/path/to/ext".to_string()];
File diff suppressed because it is too large Load Diff
+2 -2
View File
@@ -87,8 +87,8 @@ impl CdpClient {
let ws_tx = Arc::new(Mutex::new(ws_tx));
let pending: PendingMap = Arc::new(Mutex::new(HashMap::new()));
let (event_tx, _) = broadcast::channel(256);
let (raw_tx, _) = broadcast::channel(512);
let (event_tx, _) = broadcast::channel(4096);
let (raw_tx, _) = broadcast::channel(4096);
let pending_clone = pending.clone();
let event_tx_clone = event_tx.clone();
+123 -19
View File
@@ -9,7 +9,7 @@ use std::time::Duration;
use tokio::io::{AsyncBufReadExt, AsyncWriteExt, BufReader};
use tokio::signal;
use tokio::sync::{mpsc, RwLock};
use tokio::sync::{mpsc, Notify, RwLock};
use super::actions::{execute_command, DaemonState};
use super::cdp::client::CdpClient;
@@ -41,11 +41,30 @@ pub async fn run_daemon(session: &str) {
session
);
}
} else {
// Redirect stderr to /dev/null to prevent daemon crash when the
// parent CLI drops the piped stderr handle after startup. Cloud
// providers (AgentCore, Browserbase, etc.) may write to stderr
// during connection setup; a broken pipe would kill the daemon.
#[cfg(unix)]
{
use std::os::unix::io::IntoRawFd;
if let Ok(devnull) = fs::File::create("/dev/null") {
let fd = devnull.into_raw_fd();
unsafe {
libc::dup2(fd, 2);
libc::close(fd);
}
}
}
}
let pid_path = socket_dir.join(format!("{}.pid", session));
let _ = fs::write(&pid_path, process::id().to_string());
let version_path = socket_dir.join(format!("{}.version", session));
let _ = fs::write(&version_path, env!("CARGO_PKG_VERSION"));
// On Unix the daemon listens on a Unix domain socket; on Windows it uses
// TCP, so there is no .sock file — only a .port file written by the server.
let socket_path = socket_dir.join(format!("{}.sock", session));
@@ -118,6 +137,7 @@ pub async fn run_daemon(session: &str) {
let _ = fs::remove_file(socket_dir.join(format!("{}.port", session)));
}
let _ = fs::remove_file(&pid_path);
let _ = fs::remove_file(&version_path);
let _ = fs::remove_file(&stream_path);
let _ = fs::remove_file(socket_dir.join(format!("{}.engine", session)));
let _ = fs::remove_file(socket_dir.join(format!("{}.provider", session)));
@@ -156,13 +176,18 @@ async fn run_socket_server(
let (reset_tx, mut reset_rx) = mpsc::channel::<()>(64);
let reset_tx = idle_timeout_ms.map(|_| Arc::new(reset_tx));
let mut drain_interval = tokio::time::interval(Duration::from_millis(500));
// Notifier used by handle_connection to signal the daemon loop to exit
// after a "close" command, instead of calling process::exit() which skips
// destructors and can leave Chrome processes orphaned (issue #1113).
let close_notify = Arc::new(Notify::new());
let mut drain_interval = tokio::time::interval(Duration::from_millis(100));
drain_interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop {
let sleep_future = idle_timeout_ms.map(|ms| tokio::time::sleep(Duration::from_millis(ms)));
let mut sleep_pin = sleep_future.map(Box::pin);
let idle_sleep = idle_timeout_ms.map(|ms| tokio::time::sleep(Duration::from_millis(ms)));
let mut idle_sleep_pin = idle_sleep.map(Box::pin);
loop {
tokio::select! {
accept_result = listener.accept() => {
match accept_result {
@@ -170,8 +195,9 @@ async fn run_socket_server(
let state = state.clone();
let reset_tx = reset_tx.clone();
let sf = stream_file.clone();
let cn = close_notify.clone();
tokio::spawn(async move {
handle_connection(stream, state, reset_tx, sf).await;
handle_connection(stream, state, reset_tx, sf, cn).await;
});
}
Err(e) => {
@@ -193,10 +219,9 @@ async fn run_socket_server(
}
}
_ = async {
if let Some(ref mut s) = sleep_pin {
s.as_mut().await
} else {
std::future::pending::<()>().await
match idle_sleep_pin {
Some(ref mut s) => s.as_mut().await,
None => std::future::pending::<()>().await,
}
}, if idle_timeout_ms.is_some() => {
let mut s = state.lock().await;
@@ -206,8 +231,16 @@ async fn run_socket_server(
break;
}
_ = reset_rx.recv(), if idle_timeout_ms.is_some() => {
idle_sleep_pin = idle_timeout_ms
.map(|ms| Box::pin(tokio::time::sleep(Duration::from_millis(ms))));
continue;
}
_ = close_notify.notified() => {
// "close" command was handled; browser already closed by
// handle_close(). Break to run cleanup and exit gracefully
// so destructors fire.
break;
}
_ = shutdown_signal() => {
let mut s = state.lock().await;
if let Some(ref mut mgr) = s.browser {
@@ -262,10 +295,12 @@ async fn run_socket_server(
let (reset_tx, mut reset_rx) = mpsc::channel::<()>(64);
let reset_tx = idle_timeout_ms.map(|_| Arc::new(reset_tx));
loop {
let sleep_future = idle_timeout_ms.map(|ms| tokio::time::sleep(Duration::from_millis(ms)));
let mut sleep_pin = sleep_future.map(Box::pin);
let close_notify = Arc::new(Notify::new());
let idle_sleep = idle_timeout_ms.map(|ms| tokio::time::sleep(Duration::from_millis(ms)));
let mut idle_sleep_pin = idle_sleep.map(Box::pin);
loop {
tokio::select! {
accept_result = listener.accept() => {
match accept_result {
@@ -273,8 +308,9 @@ async fn run_socket_server(
let state = state.clone();
let reset_tx = reset_tx.clone();
let sf = stream_file.clone();
let cn = close_notify.clone();
tokio::spawn(async move {
handle_connection(stream, state, reset_tx, sf).await;
handle_connection(stream, state, reset_tx, sf, cn).await;
});
}
Err(e) => {
@@ -283,10 +319,9 @@ async fn run_socket_server(
}
}
_ = async {
if let Some(ref mut s) = sleep_pin {
s.as_mut().await
} else {
std::future::pending::<()>().await
match idle_sleep_pin {
Some(ref mut s) => s.as_mut().await,
None => std::future::pending::<()>().await,
}
}, if idle_timeout_ms.is_some() => {
let mut s = state.lock().await;
@@ -297,8 +332,14 @@ async fn run_socket_server(
break;
}
_ = reset_rx.recv(), if idle_timeout_ms.is_some() => {
idle_sleep_pin = idle_timeout_ms
.map(|ms| Box::pin(tokio::time::sleep(Duration::from_millis(ms))));
continue;
}
_ = close_notify.notified() => {
let _ = fs::remove_file(&port_path);
break;
}
_ = shutdown_signal() => {
let mut s = state.lock().await;
if let Some(ref mut mgr) = s.browser {
@@ -318,6 +359,7 @@ async fn handle_connection<S>(
state: std::sync::Arc<tokio::sync::Mutex<DaemonState>>,
idle_reset_tx: Option<Arc<mpsc::Sender<()>>>,
stream_file_cleanup: Option<PathBuf>,
close_notify: Arc<Notify>,
) where
S: tokio::io::AsyncRead + tokio::io::AsyncWrite + Unpin,
{
@@ -374,8 +416,12 @@ async fn handle_connection<S>(
if let Some(ref path) = stream_file_cleanup {
let _ = fs::remove_file(path);
}
// Signal the daemon loop to exit gracefully instead of
// calling process::exit(), which skips destructors and
// can leave Chrome processes orphaned (issue #1113).
tokio::time::sleep(tokio::time::Duration::from_millis(100)).await;
process::exit(0);
close_notify.notify_one();
return;
}
}
Err(_) => break,
@@ -534,6 +580,64 @@ mod tests {
}
}
/// Regression test for #1101: idle timeout must fire even while the
/// drain interval ticks every 500 ms. The bug was that `sleep_future`
/// was created **inside** the loop, so each drain tick dropped the
/// in-progress sleep and replaced it with a fresh one the timer
/// could never reach its deadline.
#[tokio::test]
async fn test_idle_timeout_fires_despite_drain_interval() {
use tokio::sync::mpsc;
let idle_timeout_ms: u64 = 1000;
let mut drain_interval = tokio::time::interval(Duration::from_millis(500));
drain_interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
let (_reset_tx, mut reset_rx) = mpsc::channel::<()>(64);
let start = tokio::time::Instant::now();
let exited = tokio::time::timeout(Duration::from_secs(5), async {
let mut idle_sleep_pin = Some(Box::pin(tokio::time::sleep(Duration::from_millis(
idle_timeout_ms,
))));
loop {
tokio::select! {
_ = drain_interval.tick() => {}
_ = async {
match idle_sleep_pin {
Some(ref mut s) => s.as_mut().await,
None => std::future::pending::<()>().await,
}
} => {
break;
}
_ = reset_rx.recv() => {
idle_sleep_pin = Some(Box::pin(
tokio::time::sleep(Duration::from_millis(idle_timeout_ms)),
));
continue;
}
}
}
})
.await;
let elapsed = start.elapsed();
assert!(
exited.is_ok(),
"idle timeout never fired loop ran for >5 s (bug #1101)"
);
assert!(
elapsed < Duration::from_millis(idle_timeout_ms + 500),
"idle timeout took too long: {:?} (expected ~{} ms)",
elapsed,
idle_timeout_ms,
);
}
/// Verify that `ChromeProcess::has_exited()` (which uses `Child::try_wait()`)
/// correctly detects a killed child, the same way the drain interval does
/// in the fixed daemon code. This ensures crash detection works without
File diff suppressed because it is too large Load Diff
+299 -1
View File
@@ -103,6 +103,10 @@ impl RefMap {
entries
}
pub fn remove(&mut self, ref_id: &str) {
self.map.remove(ref_id);
}
pub fn clear(&mut self) {
self.map.clear();
self.next_ref = 1;
@@ -159,6 +163,30 @@ pub async fn resolve_element_center(
// Try cached backend_node_id first (fast path)
if let Some(backend_node_id) = entry.backend_node_id {
// Identity check: React often re-uses the same DOM node when
// re-rendering — backendNodeId stays the same but accessibleName
// / role changes. Without this verification, `click @e20` (saved
// when the button said "Add post") happily clicks the *same*
// node that now says "Post all", silently submitting the thread.
//
// Set AGENT_BROWSER_VERIFY_REF=0 to skip (saves one CDP
// roundtrip per ref-based interaction; only safe if you know
// the page is static between snapshot and click).
if std::env::var("AGENT_BROWSER_VERIFY_REF").as_deref() != Ok("0") {
if let Err(e) = verify_ref_identity(
client,
effective_session_id,
backend_node_id,
&ref_id,
&entry.role,
&entry.name,
)
.await
{
return Err(e);
}
}
let result: Result<DomGetBoxModelResult, String> = client
.send_command_typed(
"DOM.getBoxModel",
@@ -173,6 +201,31 @@ pub async fn resolve_element_center(
if let Ok(r) = result {
let (x, y) = box_model_center(&r.model);
// Occlusion check: a transient overlay (X.com's "click
// outside to close" mask, modal backdrop, sticky banner,
// etc.) can land on top of our target between snapshot
// and click. Coordinates are correct, but
// `document.elementFromPoint(x, y)` returns the overlay
// — and the click goes to the overlay's handler, not
// ours. Catch it here so the user gets "occluded by
// DIV[testid=mask]" instead of "modal silently closed +
// thread submitted by accident".
//
// Set AGENT_BROWSER_VERIFY_CLICK_TARGET=0 to skip.
if std::env::var("AGENT_BROWSER_VERIFY_CLICK_TARGET").as_deref() != Ok("0") {
if let Err(e) = verify_click_target(
client,
effective_session_id,
backend_node_id,
&ref_id,
x,
y,
)
.await
{
return Err(e);
}
}
return Ok((x, y, effective_session_id.to_string()));
}
// backend_node_id is stale; re-query the accessibility tree below
@@ -226,6 +279,24 @@ pub async fn resolve_element_object_id(
// Try cached backend_node_id first (fast path)
if let Some(backend_node_id) = entry.backend_node_id {
// Same identity guard as resolve_element_center — see that
// function for why React DOM-node-reuse breaks ref-based
// interactions if we skip this.
if std::env::var("AGENT_BROWSER_VERIFY_REF").as_deref() != Ok("0") {
if let Err(e) = verify_ref_identity(
client,
effective_session_id,
backend_node_id,
&ref_id,
&entry.role,
&entry.name,
)
.await
{
return Err(e);
}
}
let result: Result<DomResolveNodeResult, String> = client
.send_command_typed(
"DOM.resolveNode",
@@ -329,6 +400,233 @@ fn resolve_frame_session<'a>(
.unwrap_or(session_id)
}
/// Verify that the cached backendNodeId still has the same accessible role
/// and name it had when the snapshot ran. Catches the case where React (or
/// any reconciler) reused the DOM node for a different component instance
/// — same physical node, different semantics.
///
/// On mismatch, returns an actionable error naming both the snapshot label
/// and the current label so the agent can re-snapshot intelligently.
/// On any CDP failure (e.g. node deleted), returns Ok(()) so the caller's
/// existing fallback (`find_node_id_by_role_name`) takes over.
async fn verify_ref_identity(
client: &CdpClient,
session_id: &str,
backend_node_id: i64,
ref_id: &str,
expected_role: &str,
expected_name: &str,
) -> Result<(), String> {
let params = serde_json::json!({
"backendNodeId": backend_node_id,
"fetchRelatives": false,
});
// Tight 1s timeout: this is a defensive guard, not a critical path.
// The default 30s CDP timeout was the dominant factor in the
// "click hangs 5+ minutes" report — three CDP calls (verify +
// resolveNode + paint-settle) at 30s each, multiplied by parallel
// click invocations queueing on the daemon, totalled multi-minute
// user-visible hangs. Cap our own helper so a stuck AX query
// doesn't make `click` worse than the no-guard version was.
let resp: Result<GetFullAXTreeResult, String> = match tokio::time::timeout(
std::time::Duration::from_secs(1),
client.send_command_typed("Accessibility.getPartialAXTree", &params, Some(session_id)),
)
.await
{
Ok(r) => r,
// Timeout: skip identity verification rather than block the click.
Err(_) => return Ok(()),
};
let Ok(tree) = resp else {
// Node likely gone; let the box-model call fail and trigger fallback.
return Ok(());
};
// Find the AXNode for our backendNodeId. fetchRelatives=false still
// returns ancestors; the target node has the matching backendNodeId.
let Some(node) = tree
.nodes
.iter()
.find(|n| n.backend_d_o_m_node_id == Some(backend_node_id))
else {
return Ok(());
};
let actual_role = extract_ax_string(&node.role);
let actual_name = extract_ax_string(&node.name);
if actual_role == expected_role && actual_name == expected_name {
return Ok(());
}
Err(format!(
"Ref {} no longer matches its snapshot. Was [{} \"{}\"], now [{} \"{}\"].\n\
The DOM mutated between snapshot and interaction (typical with React/Vue \
reusing nodes during re-render). Take a fresh snapshot, then re-target.\n\
To bypass this guard set AGENT_BROWSER_VERIFY_REF=0.",
ref_id, expected_role, expected_name, actual_role, actual_name,
))
}
/// At the moment we'd dispatch the click, ask the page itself which element
/// occupies (x, y). If it's not our target (and not a descendant or
/// ancestor), an overlay has appeared between snapshot and click — we'd
/// silently click the overlay otherwise. Returns Err with details about
/// the occluding element so the caller can wait + re-snapshot.
///
/// Implemented as a single Runtime.callFunctionOn: resolve the cached
/// backendNodeId to a remote object, then run a function on it that
/// compares with elementFromPoint. The function returns null when the
/// click is safe and a JSON string with diagnostic info when it isn't.
async fn verify_click_target(
client: &CdpClient,
session_id: &str,
backend_node_id: i64,
ref_id: &str,
x: f64,
y: f64,
) -> Result<(), String> {
use serde::Deserialize;
// Resolve once. backendNodeId is stable across renders; only the
// element under (x, y) is what changes when an overlay flickers.
let resolve_params = DomResolveNodeParams {
backend_node_id: Some(backend_node_id),
node_id: None,
object_group: Some("agent-browser-occlusion".to_string()),
};
let resolve_fut = client.send_command_typed::<_, serde_json::Value>(
"DOM.resolveNode",
&resolve_params,
Some(session_id),
);
let Ok(resolve_resp) =
tokio::time::timeout(std::time::Duration::from_millis(500), resolve_fut).await
else {
return Ok(());
};
let Ok(resolved) = resolve_resp else { return Ok(()) };
let Some(object_id) = resolved
.get("object")
.and_then(|o| o.get("objectId"))
.and_then(|v| v.as_str())
else {
return Ok(());
};
// Auto-retry on transient occlusion. Many real-world overlays
// (modal backdrops, focus rings, click-outside masks) blink in for
// a frame or two during state transitions and clear on their own.
// Without retries the user gets an "occluded" error and has to
// wrap every click in their own retry loop. With retries the
// common case is invisible — only persistent overlays surface.
//
// AGENT_BROWSER_OCCLUSION_RETRIES (default 3, 0 disables)
// AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS (default 200)
let max_retries: u32 = std::env::var("AGENT_BROWSER_OCCLUSION_RETRIES")
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(3);
let retry_delay_ms: u64 = std::env::var("AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS")
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(200);
#[derive(Deserialize)]
struct Occluder {
tag: Option<String>,
testid: Option<String>,
role: Option<String>,
#[serde(rename = "ariaLabel")]
aria_label: Option<String>,
text: Option<String>,
reason: Option<String>,
}
// function(x, y) { ... } where `this` is the target element.
// Return null → click is safe.
// Return JSON → describes the occluding element.
let function_decl = "function(x, y) { \
const at = document.elementFromPoint(x, y); \
if (!at) return JSON.stringify({reason:'no-element-at-point'}); \
if (at === this || this.contains(at) || at.contains(this)) return null; \
return JSON.stringify({ \
tag: at.tagName, \
testid: (at.dataset && at.dataset.testid) || null, \
role: at.getAttribute('role'), \
ariaLabel: at.getAttribute('aria-label'), \
text: ((at.textContent||'').trim().slice(0, 60)) \
}); \
}";
let mut last_occ: Option<Occluder> = None;
for attempt in 0..=max_retries {
if attempt > 0 {
tokio::time::sleep(std::time::Duration::from_millis(retry_delay_ms)).await;
}
let call_params = serde_json::json!({
"objectId": object_id,
"functionDeclaration": function_decl,
"arguments": [{"value": x}, {"value": y}],
"returnByValue": true,
});
let call_fut = client.send_command_typed::<_, serde_json::Value>(
"Runtime.callFunctionOn",
&call_params,
Some(session_id),
);
let Ok(call_resp) =
tokio::time::timeout(std::time::Duration::from_millis(500), call_fut).await
else {
return Ok(()); // probe itself stalled — fall through to click
};
let Ok(call_result) = call_resp else {
return Ok(());
};
let value = call_result.get("result").and_then(|r| r.get("value"));
let json_str = match value {
Some(serde_json::Value::String(s)) => s.clone(),
// null / undefined → element at point IS our target. Safe.
_ => return Ok(()),
};
let occ: Occluder = match serde_json::from_str(&json_str) {
Ok(v) => v,
Err(_) => return Ok(()),
};
last_occ = Some(occ);
}
// All retries exhausted — overlay is sticky. Build the descriptive error.
let occ = last_occ.expect("loop ran at least once");
if let Some(reason) = occ.reason {
return Err(format!(
"Ref {} cannot be clicked at its computed position: {}. \
The element may have moved off-screen re-run snapshot.",
ref_id, reason
));
}
let mut desc = occ.tag.unwrap_or_else(|| "unknown".to_string());
if let Some(t) = occ.testid {
desc.push_str(&format!("[testid={}]", t));
}
if let Some(r) = occ.role {
desc.push_str(&format!("[role={}]", r));
}
if let Some(a) = occ.aria_label {
desc.push_str(&format!("[aria-label=\"{}\"]", a));
}
if let Some(t) = occ.text {
if !t.is_empty() {
desc.push_str(&format!(" text=\"{}\"", t));
}
}
let waited_ms = (max_retries as u64) * retry_delay_ms;
Err(format!(
"Ref {} is occluded by {} at the click point (still occluded after \
{} retries / {}ms). A persistent overlay is in the way \
re-run snapshot, dismiss the overlay, or set \
AGENT_BROWSER_VERIFY_CLICK_TARGET=0 to bypass.",
ref_id, desc, max_retries, waited_ms,
))
}
/// Re-query the accessibility tree to find a node matching role+name+nth,
/// returning its fresh backendDOMNodeId. This uses the same data source
/// (Accessibility.getFullAXTree) that built the ref map during snapshot,
@@ -380,7 +678,7 @@ async fn find_node_id_by_role_name(
))
}
fn extract_ax_string(value: &Option<AXValue>) -> String {
pub(super) fn extract_ax_string(value: &Option<AXValue>) -> String {
match value {
Some(v) => match &v.value {
Some(Value::String(s)) => s.clone(),
+41
View File
@@ -884,6 +884,46 @@ pub async fn tap_touch(
Ok(())
}
/// After a click is dispatched, give the page two animation frames + a
/// microtask boundary to let React/Vue/Svelte commit any state update
/// scheduled by the click handler. Without this wait, follow-up commands
/// (e.g. `inserttext` against the textbox the click was supposed to mount)
/// race the renderer and can land on stale or wrong elements.
///
/// The wait is bounded to ~33ms in the common case (two RAFs at 60fps) and
/// returns immediately on any error — never an exception path.
///
/// Set `AGENT_BROWSER_CLICK_WAIT_STABLE=0` to disable for perf-sensitive
/// scripts that don't drive SPA UIs.
async fn wait_for_paint_settled(client: &CdpClient, session_id: &str) {
if std::env::var("AGENT_BROWSER_CLICK_WAIT_STABLE").as_deref() == Ok("0") {
return;
}
let script = "new Promise(resolve => \
requestAnimationFrame(() => \
requestAnimationFrame(() => \
queueMicrotask(() => resolve(true)))))";
// Tight 500ms timeout. RAF normally fires at 16ms, two RAFs total ~33ms.
// If the tab is hidden / throttled / page is doing something pathological
// and RAF doesn't fire in 500ms, we'd rather return now than stall the
// user's click. Without this cap, a stuck RAF inherited the default 30s
// CDP timeout and was the main contributor to the "click hangs 5+ min"
// user report.
let _ = tokio::time::timeout(
std::time::Duration::from_millis(500),
client.send_command_typed::<_, Value>(
"Runtime.evaluate",
&EvaluateParams {
expression: script.to_string(),
return_by_value: Some(true),
await_promise: Some(true),
},
Some(session_id),
),
)
.await;
}
async fn dispatch_click(
client: &CdpClient,
session_id: &str,
@@ -955,6 +995,7 @@ async fn dispatch_click(
)
.await?;
wait_for_paint_settled(client, session_id).await;
Ok(())
}
+4
View File
@@ -25,6 +25,8 @@ pub mod policy;
#[allow(dead_code)]
pub mod providers;
#[allow(dead_code)]
pub mod react;
#[allow(dead_code)]
pub mod recording;
#[allow(dead_code)]
pub mod screenshot;
@@ -33,6 +35,8 @@ pub mod snapshot;
#[allow(dead_code)]
pub mod state;
#[allow(dead_code)]
pub mod stealth;
#[allow(dead_code)]
pub mod storage;
#[allow(dead_code)]
pub mod stream;
File diff suppressed because one or more lines are too long
+31
View File
@@ -0,0 +1,31 @@
//! React/web introspection primitives.
//!
//! Scripts and handlers for the `react` subcommands (tree, inspect, renders,
//! suspense) plus the universal `vitals` verb and the generic `pushstate`
//! SPA-navigation action. These primitives are framework-agnostic: React-side
//! commands only require the `__REACT_DEVTOOLS_GLOBAL_HOOK__` to be installed,
//! and `vitals` / `pushstate` are pure web-standard APIs.
//!
//! The React DevTools `installHook.js` is vendored from the React DevTools
//! Chrome extension (MIT, facebook/react). It's registered via
//! `addScriptToEvaluateOnNewDocument` before any page JS runs when the user
//! passes `--enable react-devtools` at launch.
pub mod scripts;
mod renders;
mod suspense;
mod tree;
mod vitals;
pub use renders::{format_renders_report, RendersData};
pub use suspense::{format_suspense_report, Boundary};
pub use tree::{format_tree, TreeNode};
pub use vitals::{format_vitals_report, VitalsData};
/// React DevTools hook script (MIT, from facebook/react).
/// Registered via `addScriptToEvaluateOnNewDocument` to install
/// `window.__REACT_DEVTOOLS_GLOBAL_HOOK__` before any page JS runs. React
/// detects the hook on boot and registers its renderers against it, which
/// enables every `react …` command.
pub const INSTALL_HOOK_JS: &str = include_str!("installHook.js");
+169
View File
@@ -0,0 +1,169 @@
//! React fiber render profiler report formatter.
//!
//! Default output is the
//! full agent-readable report (summary, FPS, component table, per-component
//! "change details (prev -> next)"). `--json` emits the raw structured data
//! instead.
use serde::{Deserialize, Serialize};
#[derive(Debug, Deserialize, Serialize)]
pub struct RendersData {
pub elapsed: f64,
pub fps: FpsStats,
#[serde(rename = "totalRenders")]
pub total_renders: i64,
#[serde(rename = "totalMounts")]
pub total_mounts: i64,
#[serde(rename = "totalReRenders")]
pub total_re_renders: i64,
#[serde(rename = "totalComponents")]
pub total_components: i64,
pub components: Vec<Component>,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct FpsStats {
pub avg: i64,
pub min: i64,
pub max: i64,
pub drops: i64,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct Component {
pub name: String,
pub count: i64,
pub mounts: i64,
#[serde(rename = "reRenders")]
pub re_renders: i64,
#[serde(rename = "instanceCount")]
pub instance_count: i64,
#[serde(rename = "totalTime")]
pub total_time: f64,
#[serde(rename = "selfTime")]
pub self_time: f64,
#[serde(rename = "domMutations")]
pub dom_mutations: i64,
pub changes: Vec<Change>,
#[serde(rename = "changeSummary")]
pub change_summary: std::collections::HashMap<String, i64>,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct Change {
#[serde(rename = "type")]
pub change_type: String,
pub name: Option<String>,
pub prev: Option<String>,
pub next: Option<String>,
}
pub fn format_renders_report(d: &RendersData) -> String {
if d.components.is_empty() {
return "(no renders captured)".to_string();
}
let mut lines: Vec<String> = Vec::new();
lines.push(format!("# Render Profile - {}s recording", d.elapsed));
lines.push(format!(
"# {} renders ({} mounts + {} re-renders) across {} components",
d.total_renders, d.total_mounts, d.total_re_renders, d.total_components
));
lines.push(format!(
"# FPS: avg {}, min {}, max {}, drops (<30fps): {}",
d.fps.avg, d.fps.min, d.fps.max, d.fps.drops
));
lines.push(String::new());
lines.push("## Components by total render time".to_string());
let top: Vec<&Component> = d.components.iter().take(50).collect();
let name_w = top.iter().map(|c| c.name.len()).max().unwrap_or(9).max(9);
lines.push(format!(
"| {:<name_w$} | Insts | Mounts | Re-renders | Total | Self | DOM | Top change reason |",
"Component",
name_w = name_w
));
lines.push(format!(
"| {:-<name_w$} | ----- | ------ | ---------- | -------- | -------- | ----- | -------------------------- |",
"",
name_w = name_w
));
for c in &top {
let total = if c.total_time > 0.0 {
format!("{}ms", c.total_time)
} else {
"-".to_string()
};
let self_time = if c.self_time > 0.0 {
format!("{}ms", c.self_time)
} else {
"-".to_string()
};
let dom = format!("{}/{}", c.dom_mutations, c.count);
let top_change = c
.change_summary
.iter()
.max_by_key(|(_, v)| *v)
.map(|(k, _)| k.as_str())
.unwrap_or("-");
lines.push(format!(
"| {:<name_w$} | {:>5} | {:>6} | {:>10} | {:>8} | {:>8} | {:>5} | {:<26} |",
c.name,
c.instance_count,
c.mounts,
c.re_renders,
total,
self_time,
dom,
top_change,
name_w = name_w
));
}
if d.components.len() > 50 {
lines.push(format!("... and {} more", d.components.len() - 50));
}
let detailed: Vec<&Component> = d
.components
.iter()
.filter(|c| {
c.changes
.iter()
.any(|ch| ch.change_type != "mount" && ch.change_type != "parent")
})
.take(15)
.collect();
if !detailed.is_empty() {
lines.push(String::new());
lines.push("## Change details (prev -> next)".to_string());
for c in &detailed {
lines.push(format!(" {}", c.name));
let mut seen = std::collections::HashSet::new();
for ch in &c.changes {
if ch.change_type == "mount" || ch.change_type == "parent" {
continue;
}
let name = ch.name.clone().unwrap_or_default();
let key = format!("{}:{}", ch.change_type, name);
if !seen.insert(key) {
continue;
}
let label = match ch.change_type.as_str() {
"props" => format!("props.{}", name),
"state" => format!("state ({})", name),
_ => format!("context ({})", name),
};
lines.push(format!(
" {}: {} -> {}",
label,
ch.prev.clone().unwrap_or_else(|| "?".into()),
ch.next.clone().unwrap_or_else(|| "?".into())
));
}
}
}
lines.join("\n")
}
+745
View File
@@ -0,0 +1,745 @@
//! Browser-side evaluation scripts for React/web introspection.
//!
//! These are JavaScript strings evaluated in the page context via
//! `Runtime.evaluate`. They assume the React DevTools hook is already
//! installed (via `--enable react-devtools`) except for `VITALS_INIT` and
//! `PUSHSTATE`, which only use standard Web APIs.
//!
//! Kept as raw strings rather than TS/JS files because the daemon is a single
//! Rust binary with no filesystem vendor step at runtime.
/// Build a no-argument async IIFE page-eval that returns the component tree as
/// JSON.
pub const TREE_SNAPSHOT: &str = r#"
(async () => {
const hook = window.__REACT_DEVTOOLS_GLOBAL_HOOK__;
if (!hook) throw new Error("React DevTools hook not installed - relaunch with --enable react-devtools");
const ri = hook.rendererInterfaces && hook.rendererInterfaces.get && hook.rendererInterfaces.get(1);
if (!ri) throw new Error("No React renderer attached - the page has not booted React yet");
const batches = await new Promise((resolve) => {
const out = [];
const origEmit = hook.emit;
hook.emit = function (event, payload) {
if (event === "operations") out.push(Array.from(payload));
return origEmit.apply(hook, arguments);
};
ri.flushInitialOperations();
setTimeout(() => {
hook.emit = origEmit;
resolve(out);
}, 50);
});
const nodes = batches.flatMap((ops) => {
let i = 2;
const strings = [null];
const tableEnd = ++i + ops[i - 1];
while (i < tableEnd) {
const len = ops[i++];
strings.push(String.fromCodePoint(...ops.slice(i, i + len)));
i += len;
}
const out = [];
while (i < ops.length) {
const op = ops[i];
if (op === 1) {
const id = ops[i + 1];
const type = ops[i + 2];
i += 3;
if (type === 11) {
out.push({ id, type, name: null, key: null, parent: 0 });
i += 4;
} else {
out.push({
id,
type,
name: strings[ops[i + 2]] || null,
key: strings[ops[i + 3]] || null,
parent: ops[i],
});
i += 5;
}
} else {
i += skip(op, ops, i);
}
}
return out;
function skip(op, ops, i) {
if (op === 2) return 2 + ops[i + 1];
if (op === 3) return 3 + ops[i + 2];
if (op === 4) return 3;
if (op === 5) return 4;
if (op === 6) return 1;
if (op === 7) return 3;
if (op === 8) return 6 + rects(ops[i + 5]);
if (op === 9) return 2 + ops[i + 1];
if (op === 10) return 3 + ops[i + 2];
if (op === 11) return 3 + rects(ops[i + 2]);
if (op === 12) return suspenders(ops, i);
if (op === 13) return 2;
return 1;
}
function rects(n) {
return n === -1 ? 0 : n * 4;
}
function suspenders(ops, i) {
let j = i + 2;
for (let c = 0; c < ops[i + 1]; c++) j += 5 + ops[j + 4];
return j - i;
}
});
return JSON.stringify(nodes);
})()
"#;
/// Template for `inspect` — replace {{ID}} with the numeric fiber id.
pub const TREE_INSPECT: &str = r#"
(() => {
const id = {{ID}};
const hook = window.__REACT_DEVTOOLS_GLOBAL_HOOK__;
const ri = hook && hook.rendererInterfaces && hook.rendererInterfaces.get && hook.rendererInterfaces.get(1);
if (!ri) throw new Error("No React renderer attached");
if (!ri.hasElementWithId(id)) throw new Error("element " + id + " not found (page reloaded?)");
const result = ri.inspectElement(1, id, null, true);
if (!result || result.type !== "full-data") {
throw new Error("inspect failed: " + (result && result.type));
}
const v = result.value;
const name = ri.getDisplayNameForElementID(id);
const lines = [name + " #" + id];
if (v.key != null) lines.push("key: " + JSON.stringify(v.key));
section("props", v.props);
section("hooks", v.hooks);
section("state", v.state);
section("context", v.context);
if (v.owners && v.owners.length) {
lines.push("rendered by: " + v.owners.map((o) => o.displayName).join(" > "));
}
const source = Array.isArray(v.source)
? [v.source[1], v.source[2], v.source[3]]
: null;
return JSON.stringify({ text: lines.join("\n"), source });
function section(label, payload) {
const data = (payload && payload.data) || payload;
if (data == null) return;
if (Array.isArray(data)) {
if (data.length === 0) return;
lines.push(label + ":");
for (const h of data) lines.push(" " + hookLine(h));
} else if (typeof data === "object") {
const entries = Object.entries(data);
if (entries.length === 0) return;
lines.push(label + ":");
for (const [k, val] of entries) lines.push(" " + k + ": " + preview(val));
}
}
function hookLine(h) {
const idx = h.id != null ? "[" + h.id + "] " : "";
const sub = h.subHooks && h.subHooks.length ? " (" + h.subHooks.length + " sub)" : "";
return idx + h.name + ": " + preview(h.value) + sub;
}
function preview(v) {
if (v == null) return String(v);
if (typeof v !== "object") return JSON.stringify(v);
if (v.type === "undefined") return "undefined";
if (v.preview_long) return v.preview_long;
if (v.preview_short) return v.preview_short;
if (Array.isArray(v)) return "[" + v.map(preview).join(", ") + "]";
const entries = Object.entries(v).map((e) => e[0] + ": " + preview(e[1]));
return "{" + entries.join(", ") + "}";
}
})()
"#;
/// Fiber profiler init script. Registered via `addScriptToEvaluateOnNewDocument`
/// so it survives navigations; also evaluated immediately on the current page
/// by `react renders start`.
pub const RENDERS_INIT: &str = r#"
(() => {
const hook = window.__REACT_DEVTOOLS_GLOBAL_HOOK__;
if (!hook || window.__AB_RENDERS_ACTIVE__) return;
const MAX_COMPONENTS = 200;
const data = {};
const fps = { frames: [], last: 0, rafId: 0 };
window.__AB_RENDERS__ = data;
window.__AB_RENDERS_FPS__ = fps;
window.__AB_RENDERS_START__ = performance.now();
window.__AB_RENDERS_ACTIVE__ = true;
function fpsLoop(now) {
if (fps.last > 0) fps.frames.push(now - fps.last);
fps.last = now;
fps.rafId = requestAnimationFrame(fpsLoop);
}
fps.rafId = requestAnimationFrame(fpsLoop);
const origOnCommit = hook.onCommitFiberRoot;
window.__AB_RENDERS_ORIG_COMMIT__ = origOnCommit;
hook.onCommitFiberRoot = function (rendererID, root) {
try { walkFiber(root.current); } catch {}
if (typeof origOnCommit === "function") {
return origOnCommit.apply(hook, arguments);
}
};
function getName(fiber) {
if (!fiber.type || typeof fiber.type === "string") return null;
return fiber.type.displayName || fiber.type.name || null;
}
function brief(val) {
if (val === undefined) return "undefined";
if (val === null) return "null";
if (typeof val === "function") return "fn()";
if (typeof val === "string") return val.length > 60 ? '"' + val.slice(0, 57) + '..."' : '"' + val + '"';
if (typeof val === "number" || typeof val === "boolean") return String(val);
if (Array.isArray(val)) return "Array(" + val.length + ")";
if (typeof val === "object") {
try {
const keys = Object.keys(val);
return keys.length <= 3 ? "{" + keys.join(", ") + "}" : "{" + keys.slice(0, 3).join(", ") + ", ...}";
} catch { return "{...}"; }
}
return String(val).slice(0, 40);
}
function getChanges(fiber) {
const changes = [];
const alt = fiber.alternate;
if (!alt) { changes.push({ type: "mount" }); return changes; }
if (fiber.memoizedProps !== alt.memoizedProps) {
const curr = fiber.memoizedProps || {};
const prev = alt.memoizedProps || {};
const allKeys = new Set([...Object.keys(curr), ...Object.keys(prev)]);
for (const k of allKeys) {
if (k !== "children" && curr[k] !== prev[k]) {
changes.push({ type: "props", name: k, prev: brief(prev[k]), next: brief(curr[k]) });
}
}
}
if (fiber.memoizedState !== alt.memoizedState) {
let curr = fiber.memoizedState;
let prev = alt.memoizedState;
let hookIdx = 0;
while (curr || prev) {
if ((curr && curr.memoizedState) !== (prev && prev.memoizedState)) {
changes.push({
type: "state",
name: "hook #" + hookIdx,
prev: brief(prev && prev.memoizedState),
next: brief(curr && curr.memoizedState),
});
}
curr = curr && curr.next;
prev = prev && prev.next;
hookIdx++;
}
}
if (fiber.dependencies && fiber.dependencies.firstContext) {
let ctx = fiber.dependencies.firstContext;
let altCtx = alt.dependencies && alt.dependencies.firstContext;
while (ctx) {
if (!altCtx || ctx.memoizedValue !== (altCtx && altCtx.memoizedValue)) {
const ctxName =
(ctx.context && ctx.context.displayName) ||
(ctx.context && ctx.context.Provider && ctx.context.Provider.displayName) ||
"unknown";
changes.push({
type: "context",
name: ctxName,
prev: brief(altCtx && altCtx.memoizedValue),
next: brief(ctx.memoizedValue),
});
}
ctx = ctx.next;
altCtx = altCtx && altCtx.next;
}
}
if (changes.length === 0) {
let parent = fiber.return;
while (parent) {
const pName = getName(parent);
if (pName) {
const suffix = !parent.alternate ? " (mount)" : "";
changes.push({ type: "parent", name: pName + suffix });
break;
}
parent = parent.return;
}
if (changes.length === 0) changes.push({ type: "parent", name: "unknown" });
}
return changes;
}
function childrenTime(fiber) {
let t = 0;
let child = fiber.child;
while (child) {
if (typeof child.actualDuration === "number") t += child.actualDuration;
child = child.sibling;
}
return t;
}
function hasDomMutation(fiber) {
if (!fiber.alternate) return true;
let child = fiber.child;
while (child) {
if (typeof child.type === "string" && (child.flags & 6) > 0) return true;
child = child.sibling;
}
return false;
}
function walkFiber(fiber) {
if (!fiber) return;
const tag = fiber.tag;
if (tag === 0 || tag === 1 || tag === 2 || tag === 11 || tag === 15) {
const didRender =
fiber.alternate === null ||
fiber.flags > 0 ||
fiber.memoizedProps !== (fiber.alternate && fiber.alternate.memoizedProps) ||
fiber.memoizedState !== (fiber.alternate && fiber.alternate.memoizedState);
if (didRender) {
const name = getName(fiber);
if (name) {
if (!(name in data) && Object.keys(data).length >= MAX_COMPONENTS) {
// at cap - skip
} else {
if (!data[name]) {
data[name] = {
count: 0, mounts: 0, totalTime: 0, selfTime: 0,
domMutations: 0, changes: [], _instances: new Set(),
};
}
data[name].count++;
if (!fiber.alternate) data[name].mounts++;
if (!data[name]._instances.has(fiber)) {
data[name]._instances.add(fiber);
if (fiber.alternate) data[name]._instances.add(fiber.alternate);
}
if (typeof fiber.actualDuration === "number") {
data[name].totalTime += fiber.actualDuration;
data[name].selfTime += Math.max(0, fiber.actualDuration - childrenTime(fiber));
}
if (hasDomMutation(fiber)) data[name].domMutations++;
const ch = getChanges(fiber);
for (const c of ch) {
if (data[name].changes.length < 50) data[name].changes.push(c);
}
}
}
}
}
walkFiber(fiber.child);
walkFiber(fiber.sibling);
}
})()
"#;
/// Stop script for fiber profiler. Returns the collected profile as JSON.
pub const RENDERS_STOP: &str = r#"
(() => {
const active = window.__AB_RENDERS_ACTIVE__;
if (!active) throw new Error("renders recording not active - run `react renders start` first");
const data = window.__AB_RENDERS__;
const startTime = window.__AB_RENDERS_START__;
const elapsed = performance.now() - startTime;
const fpsData = window.__AB_RENDERS_FPS__;
let fpsStats = { avg: 0, min: 0, max: 0, drops: 0 };
if (fpsData) {
cancelAnimationFrame(fpsData.rafId);
if (fpsData.frames.length > 0) {
const fpsSamples = fpsData.frames.map((dt) => (dt > 0 ? 1000 / dt : 0));
const sum = fpsSamples.reduce((a, b) => a + b, 0);
fpsStats = {
avg: Math.round(sum / fpsSamples.length),
min: Math.round(Math.min(...fpsSamples)),
max: Math.round(Math.max(...fpsSamples)),
drops: fpsSamples.filter((f) => f < 30).length,
};
}
}
const hook = window.__REACT_DEVTOOLS_GLOBAL_HOOK__;
const orig = window.__AB_RENDERS_ORIG_COMMIT__;
if (hook) hook.onCommitFiberRoot = orig || undefined;
delete window.__AB_RENDERS__;
delete window.__AB_RENDERS_START__;
delete window.__AB_RENDERS_ACTIVE__;
delete window.__AB_RENDERS_ORIG_COMMIT__;
delete window.__AB_RENDERS_FPS__;
if (!data) {
return JSON.stringify({
elapsed: 0, fps: fpsStats, totalRenders: 0, totalMounts: 0,
totalReRenders: 0, totalComponents: 0, components: [],
});
}
const round = (n) => Math.round(n * 100) / 100;
const components = Object.entries(data)
.map(([name, entry]) => {
const summary = {};
for (const c of entry.changes) {
const key = c.type === "props" ? "props." + c.name
: c.type === "state" ? "state (" + c.name + ")"
: c.type === "context" ? "context (" + c.name + ")"
: c.type === "parent" ? "parent (" + c.name + ")"
: c.type;
summary[key] = (summary[key] || 0) + 1;
}
return {
name,
count: entry.count,
mounts: entry.mounts,
reRenders: entry.count - entry.mounts,
instanceCount: entry._instances.size,
totalTime: round(entry.totalTime),
selfTime: round(entry.selfTime),
domMutations: entry.domMutations,
changes: entry.changes,
changeSummary: summary,
};
})
.sort((a, b) => b.totalTime - a.totalTime || b.count - a.count);
return JSON.stringify({
elapsed: round(elapsed / 1000),
fps: fpsStats,
totalRenders: components.reduce((s, c) => s + c.count, 0),
totalMounts: components.reduce((s, c) => s + c.mounts, 0),
totalReRenders: components.reduce((s, c) => s + c.reRenders, 0),
totalComponents: components.length,
components,
});
})()
"#;
/// Suspense boundary walker. Returns boundaries with suspendedBy metadata as JSON.
pub const SUSPENSE_WALK: &str = r#"
(async () => {
const hook = window.__REACT_DEVTOOLS_GLOBAL_HOOK__;
if (!hook) throw new Error("React DevTools hook not installed - relaunch with --enable react-devtools");
const ri = hook.rendererInterfaces && hook.rendererInterfaces.get && hook.rendererInterfaces.get(1);
if (!ri) throw new Error("No React renderer attached");
const batches = await new Promise((resolve) => {
const out = [];
const origEmit = hook.emit;
hook.emit = function (event, payload) {
if (event === "operations") out.push(payload);
return origEmit.apply(this, arguments);
};
ri.flushInitialOperations();
setTimeout(() => {
hook.emit = origEmit;
resolve(out);
}, 50);
});
const boundaryMap = new Map();
for (const ops of batches) decodeSuspenseOps(ops, boundaryMap);
const results = [];
for (const b of boundaryMap.values()) {
if (b.parentID === 0) continue;
const boundary = {
id: b.id,
parentID: b.parentID,
name: b.name,
isSuspended: b.isSuspended,
environments: b.environments,
suspendedBy: [],
unknownSuspenders: null,
owners: [],
jsxSource: null,
};
if (ri.hasElementWithId(b.id)) {
const displayName = ri.getDisplayNameForElementID(b.id);
if (displayName) boundary.name = displayName;
const result = ri.inspectElement(1, b.id, null, true);
if (result && result.type === "full-data") {
parseInspection(boundary, result.value);
}
}
results.push(boundary);
}
return JSON.stringify(results);
function decodeSuspenseOps(ops, map) {
let i = 2;
const strings = [null];
const tableEnd = ++i + ops[i - 1];
while (i < tableEnd) {
const len = ops[i++];
strings.push(String.fromCodePoint(...ops.slice(i, i + len)));
i += len;
}
while (i < ops.length) {
const op = ops[i];
if (op === 1) {
const type = ops[i + 2];
i += 3 + (type === 11 ? 4 : 5);
} else if (op === 2) {
i += 2 + ops[i + 1];
} else if (op === 3) {
i += 3 + ops[i + 2];
} else if (op === 4) {
i += 3;
} else if (op === 5) {
i += 4;
} else if (op === 6) {
i++;
} else if (op === 7) {
i += 3;
} else if (op === 8) {
const id = ops[i + 1];
const parentID = ops[i + 2];
const nameStrID = ops[i + 3];
const isSuspended = ops[i + 4] === 1;
const numRects = ops[i + 5];
i += 6;
if (numRects !== -1) i += numRects * 4;
map.set(id, { id, parentID, name: strings[nameStrID] || null, isSuspended, environments: [] });
} else if (op === 9) {
i += 2 + ops[i + 1];
} else if (op === 10) {
i += 3 + ops[i + 2];
} else if (op === 11) {
const numRects = ops[i + 2];
i += 3;
if (numRects !== -1) i += numRects * 4;
} else if (op === 12) {
i++;
const changeLen = ops[i++];
for (let c = 0; c < changeLen; c++) {
const id = ops[i++];
i++;
i++;
const isSuspended = ops[i++] === 1;
const envLen = ops[i++];
const envs = [];
for (let e = 0; e < envLen; e++) {
const n = strings[ops[i++]];
if (n != null) envs.push(n);
}
const node = map.get(id);
if (node) {
node.isSuspended = isSuspended;
for (const env of envs) {
if (!node.environments.includes(env)) node.environments.push(env);
}
}
}
} else if (op === 13) {
i += 2;
} else {
i++;
}
}
}
function parseInspection(boundary, data) {
const rawSuspendedBy = data.suspendedBy;
const rawSuspenders = Array.isArray(rawSuspendedBy)
? rawSuspendedBy
: rawSuspendedBy && Array.isArray(rawSuspendedBy.data) ? rawSuspendedBy.data : null;
if (rawSuspenders) {
for (const entry of rawSuspenders) {
const awaited = entry && entry.awaited;
if (!awaited) continue;
const desc = preview(awaited.description) || preview(awaited.value);
boundary.suspendedBy.push({
name: awaited.name || "unknown",
description: desc,
duration: awaited.end && awaited.start ? Math.round(awaited.end - awaited.start) : 0,
env: awaited.env || (entry && entry.env) || null,
ownerName: (awaited.owner && awaited.owner.displayName) || null,
ownerStack: parseStack((awaited.owner && awaited.owner.stack) || awaited.stack),
awaiterName: (entry && entry.owner && entry.owner.displayName) || null,
awaiterStack: parseStack((entry && entry.owner && entry.owner.stack) || (entry && entry.stack)),
});
}
}
if (data.unknownSuspenders && data.unknownSuspenders !== 0) {
const reasons = {
1: "production build (no debug info)",
2: "old React version (missing tracking)",
3: "thrown Promise (library using throw instead of use())",
};
boundary.unknownSuspenders = reasons[data.unknownSuspenders] || "unknown reason";
}
if (Array.isArray(data.owners)) {
for (const o of data.owners) {
if (o && o.displayName) {
const src = Array.isArray(o.stack) && o.stack.length > 0 && Array.isArray(o.stack[0])
? [o.stack[0][1] || "(unknown)", o.stack[0][2], o.stack[0][3]]
: null;
boundary.owners.push({ name: o.displayName, env: o.env || null, source: src });
}
}
}
if (Array.isArray(data.stack) && data.stack.length > 0) {
const frame = data.stack[0];
if (Array.isArray(frame) && frame.length >= 4) {
boundary.jsxSource = [frame[1] || "(unknown)", frame[2], frame[3]];
}
}
}
function parseStack(raw) {
if (!Array.isArray(raw) || raw.length === 0) return null;
return raw
.filter((f) => Array.isArray(f) && f.length >= 4)
.map((f) => [f[0] || "", f[1] || "", f[2] || 0, f[3] || 0]);
}
function preview(v) {
if (v == null) return "";
if (typeof v === "string") return v;
if (typeof v !== "object") return String(v);
if (typeof v.preview_long === "string") return v.preview_long;
if (typeof v.preview_short === "string") return v.preview_short;
if (typeof v.value === "string") return v.value;
try {
const s = JSON.stringify(v);
return s.length > 80 ? s.slice(0, 77) + "..." : s;
} catch {
return "";
}
}
})()
"#;
/// Init script for Core Web Vitals + React hydration timing capture. Installs
/// PerformanceObservers for LCP/CLS and intercepts `console.timeStamp` to
/// capture React's profiling reconciler timings. Idempotent.
pub const VITALS_INIT: &str = r#"
(() => {
if (window.__AB_VITALS_INSTALLED__) return;
window.__AB_VITALS_INSTALLED__ = true;
const cwv = { lcp: null, cls: 0, clsEntries: [], fcp: null, inp: null };
window.__AB_VITALS__ = cwv;
try {
new PerformanceObserver((list) => {
const entries = list.getEntries();
if (entries.length > 0) {
const last = entries[entries.length - 1];
cwv.lcp = {
startTime: Math.round(last.startTime * 100) / 100,
size: last.size,
element: last.element && last.element.tagName ? last.element.tagName.toLowerCase() : null,
url: last.url || null,
};
}
}).observe({ type: "largest-contentful-paint", buffered: true });
} catch {}
try {
new PerformanceObserver((list) => {
for (const entry of list.getEntries()) {
if (!entry.hadRecentInput) {
cwv.cls += entry.value;
cwv.clsEntries.push({
value: Math.round(entry.value * 10000) / 10000,
startTime: Math.round(entry.startTime * 100) / 100,
});
}
}
}).observe({ type: "layout-shift", buffered: true });
} catch {}
try {
new PerformanceObserver((list) => {
for (const entry of list.getEntries()) {
if (entry.name === "first-contentful-paint") {
cwv.fcp = Math.round(entry.startTime * 100) / 100;
}
}
}).observe({ type: "paint", buffered: true });
} catch {}
try {
new PerformanceObserver((list) => {
let worst = cwv.inp || 0;
for (const entry of list.getEntries()) {
if (entry.duration > worst) worst = entry.duration;
}
if (worst > 0) cwv.inp = Math.round(worst * 100) / 100;
}).observe({ type: "event", buffered: true, durationThreshold: 40 });
} catch {}
// React profiling build emits console.timeStamp(label, start, end, track, trackGroup, color)
// for reconciler phases and per-component hydration timing. Intercept and collect.
const timing = [];
window.__AB_REACT_TIMING__ = timing;
const orig = console.timeStamp;
console.timeStamp = function (label) {
const args = arguments;
if (typeof label === "string" && args.length >= 3 && typeof args[1] === "number") {
timing.push({
label,
startTime: args[1],
endTime: args[2],
track: args[3] || "",
trackGroup: args[4] || "",
color: args[5] || "",
});
}
return orig.apply(console, args);
};
})()
"#;
/// Read script for vitals — collects observed metrics plus Navigation Timing
/// TTFB and any React hydration phases. Returns JSON.
pub const VITALS_READ: &str = r#"
(() => {
const cwv = window.__AB_VITALS__ || {};
const timing = window.__AB_REACT_TIMING__ || [];
const nav = performance.getEntriesByType("navigation")[0];
const ttfb = nav
? Math.round((nav.responseStart - nav.requestStart) * 100) / 100
: null;
return JSON.stringify({ cwv, timing, ttfb });
})()
"#;
/// SPA client-side navigation. Tries the framework router first so Next.js
/// app/pages router triggers an RSC fetch (pure `history.pushState` would
/// be shallow routing and bypass data loading). Falls back to
/// `history.pushState` + popstate/navigate events for vanilla pages and
/// routers that listen to history events (React Router, TanStack Router,
/// Solid Router, Vue Router).
pub const PUSHSTATE: &str = r#"
((url) => {
const before = location.href;
const absolute = new URL(url, before).href;
if (absolute === before) return before;
// Next.js pages + app router expose window.next.router with a `push`
// method that triggers the RSC fetch and re-render pipeline.
const r = typeof window.next === "object" && window.next && window.next.router;
if (r && typeof r.push === "function") {
try { r.push(url); return location.href; } catch {}
}
history.pushState(null, "", absolute);
try { dispatchEvent(new PopStateEvent("popstate", { state: null })); } catch {}
try { dispatchEvent(new Event("navigate")); } catch {}
return location.href;
})({{URL}})
"#;
+633
View File
@@ -0,0 +1,633 @@
//! React Suspense boundary introspection: walker data types, classifier, and
//! human-readable report.
//!
//! The classifier labels and recommendations are React-Suspense-general —
//! they describe what kind of thing is making a boundary suspend (`client-hook`,
//! `request-api`, `server-fetch`, `cache`, `stream`, `framework`, `unknown`)
//! and a high-level direction for fixing it. Framework-specific reasoning
//! (e.g. Next.js PPR push vs goto semantics) is left to the caller.
use serde::{Deserialize, Serialize};
use std::collections::HashMap;
pub type StackFrame = (String, String, i64, i64);
#[derive(Debug, Deserialize, Serialize, Clone)]
pub struct Boundary {
pub id: i64,
#[serde(rename = "parentID")]
pub parent_id: i64,
pub name: Option<String>,
#[serde(rename = "isSuspended")]
pub is_suspended: bool,
pub environments: Vec<String>,
#[serde(rename = "suspendedBy")]
pub suspended_by: Vec<Suspender>,
#[serde(rename = "unknownSuspenders")]
pub unknown_suspenders: Option<String>,
pub owners: Vec<Owner>,
#[serde(rename = "jsxSource")]
pub jsx_source: Option<(String, i64, i64)>,
}
#[derive(Debug, Deserialize, Serialize, Clone)]
pub struct Owner {
pub name: String,
pub env: Option<String>,
pub source: Option<(String, i64, i64)>,
}
#[derive(Debug, Deserialize, Serialize, Clone)]
pub struct Suspender {
pub name: String,
pub description: String,
pub duration: i64,
pub env: Option<String>,
#[serde(rename = "ownerName")]
pub owner_name: Option<String>,
#[serde(rename = "ownerStack")]
pub owner_stack: Option<Vec<StackFrame>>,
#[serde(rename = "awaiterName")]
pub awaiter_name: Option<String>,
#[serde(rename = "awaiterStack")]
pub awaiter_stack: Option<Vec<StackFrame>>,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum BlockerKind {
ClientHook,
RequestApi,
ServerFetch,
Stream,
Cache,
Framework,
Unknown,
}
impl BlockerKind {
fn label(self) -> &'static str {
match self {
Self::ClientHook => "client-hook",
Self::RequestApi => "request-api",
Self::ServerFetch => "server-fetch",
Self::Stream => "stream",
Self::Cache => "cache",
Self::Framework => "framework",
Self::Unknown => "unknown",
}
}
fn weight(self) -> i32 {
match self {
Self::ClientHook => 7,
Self::RequestApi => 6,
Self::ServerFetch => 5,
Self::Cache => 4,
Self::Stream => 3,
Self::Unknown => 2,
Self::Framework => 1,
}
}
fn actionability(self) -> i32 {
match self {
Self::ClientHook => 90,
Self::RequestApi => 88,
Self::ServerFetch => 82,
Self::Cache => 74,
Self::Stream => 60,
Self::Unknown => 35,
Self::Framework => 18,
}
}
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum BoundaryKind {
RouteSegment,
ExplicitSuspense,
Component,
}
impl BoundaryKind {
fn label(self) -> &'static str {
match self {
Self::RouteSegment => "route-segment",
Self::ExplicitSuspense => "explicit-suspense",
Self::Component => "component",
}
}
fn weight(self) -> i32 {
match self {
Self::RouteSegment => 3,
Self::ExplicitSuspense => 2,
Self::Component => 1,
}
}
}
#[derive(Debug, Clone)]
pub struct ActionableBlocker {
pub key: String,
pub name: String,
pub kind: BlockerKind,
pub env: Option<String>,
pub description: String,
pub owner_name: Option<String>,
pub awaiter_name: Option<String>,
pub source_frame: Option<StackFrame>,
pub owner_frame: Option<StackFrame>,
pub awaiter_frame: Option<StackFrame>,
pub actionability: i32,
pub suggestion: String,
}
#[derive(Debug, Clone)]
pub struct BoundaryInsight {
pub id: i64,
pub name: Option<String>,
pub boundary_kind: BoundaryKind,
pub environments: Vec<String>,
pub source: Option<(String, i64, i64)>,
pub rendered_by: Vec<Owner>,
pub primary_blocker: Option<ActionableBlocker>,
pub blockers: Vec<ActionableBlocker>,
pub unknown_suspenders: Option<String>,
pub actionability: i32,
pub recommendation: String,
}
#[derive(Debug, Clone)]
pub struct RootCauseGroup {
pub kind: BlockerKind,
pub name: String,
pub source_frame: Option<StackFrame>,
pub boundary_names: Vec<String>,
pub count: usize,
pub actionability: i32,
pub suggestion: String,
}
pub struct AnalysisReport {
pub total_boundaries: usize,
pub dynamic_hole_count: usize,
pub static_count: usize,
pub holes: Vec<BoundaryInsight>,
pub statics: Vec<StaticBoundarySummary>,
pub root_causes: Vec<RootCauseGroup>,
pub files_to_read: Vec<String>,
}
#[derive(Debug, Clone)]
pub struct StaticBoundarySummary {
pub name: Option<String>,
pub source: Option<(String, i64, i64)>,
pub rendered_by: Vec<Owner>,
}
pub fn format_suspense_report(boundaries: &[Boundary], only_dynamic: bool) -> String {
let report = analyze_boundaries(boundaries);
format_report(&report, only_dynamic)
}
fn analyze_boundaries(boundaries: &[Boundary]) -> AnalysisReport {
let mut holes: Vec<&Boundary> = Vec::new();
let mut statics_raw: Vec<&Boundary> = Vec::new();
for b in boundaries {
if b.parent_id == 0 {
continue;
}
let has_blocker = !b.suspended_by.is_empty() || b.unknown_suspenders.is_some();
if b.is_suspended || has_blocker {
holes.push(b);
} else {
statics_raw.push(b);
}
}
let mut hole_insights: Vec<BoundaryInsight> = holes.iter().map(|b| build_insight(b)).collect();
hole_insights.sort_by(|a, b| {
b.actionability.cmp(&a.actionability).then_with(|| {
b.boundary_kind
.weight()
.cmp(&a.boundary_kind.weight())
.then_with(|| b.blockers.len().cmp(&a.blockers.len()))
.then_with(|| {
a.name
.as_deref()
.unwrap_or("")
.cmp(b.name.as_deref().unwrap_or(""))
})
})
});
let static_summaries: Vec<StaticBoundarySummary> = statics_raw
.iter()
.map(|b| StaticBoundarySummary {
name: b.name.clone(),
source: b.jsx_source.clone(),
rendered_by: b.owners.clone(),
})
.collect();
let root_causes = build_root_causes(&hole_insights);
let files_to_read = collect_files_to_read(&hole_insights, &root_causes);
AnalysisReport {
total_boundaries: hole_insights.len() + static_summaries.len(),
dynamic_hole_count: hole_insights.len(),
static_count: static_summaries.len(),
holes: hole_insights,
statics: static_summaries,
root_causes,
files_to_read,
}
}
fn build_insight(b: &Boundary) -> BoundaryInsight {
let boundary_kind = infer_boundary_kind(b);
let mut blockers: Vec<ActionableBlocker> = b
.suspended_by
.iter()
.map(build_actionable_blocker)
.collect();
blockers.sort_by(|a, b| {
b.actionability.cmp(&a.actionability).then_with(|| {
b.kind
.weight()
.cmp(&a.kind.weight())
.then_with(|| a.name.cmp(&b.name))
})
});
let primary = blockers.first().cloned();
let recommendation = recommend_fix(
boundary_kind,
primary.as_ref(),
b.unknown_suspenders.as_deref(),
);
let primary_action = primary.as_ref().map(|p| p.actionability).unwrap_or(0);
let base_action = if boundary_kind == BoundaryKind::RouteSegment {
55
} else {
0
};
BoundaryInsight {
id: b.id,
name: b.name.clone(),
boundary_kind,
environments: b.environments.clone(),
source: b.jsx_source.clone(),
rendered_by: b.owners.clone(),
primary_blocker: primary,
blockers,
unknown_suspenders: b.unknown_suspenders.clone(),
actionability: primary_action.max(base_action),
recommendation,
}
}
fn build_actionable_blocker(s: &Suspender) -> ActionableBlocker {
let owner_frame = pick_preferred_frame(s.owner_stack.as_deref());
let awaiter_frame = pick_preferred_frame(s.awaiter_stack.as_deref());
let source_frame = owner_frame.clone().or_else(|| awaiter_frame.clone());
let kind = classify_blocker(s, source_frame.as_ref());
let suggestion = suggest_blocker_fix(kind);
let mut actionability = kind.actionability();
if let Some(ref frame) = source_frame {
if !is_frameworkish_path(&frame.1) {
actionability += 8;
}
}
if s.owner_name.is_some() || s.awaiter_name.is_some() {
actionability += 4;
}
if actionability > 100 {
actionability = 100;
}
let key = build_blocker_key(&s.name, kind, source_frame.as_ref());
ActionableBlocker {
key,
name: s.name.clone(),
kind,
env: s.env.clone(),
description: s.description.clone(),
owner_name: s.owner_name.clone(),
awaiter_name: s.awaiter_name.clone(),
source_frame,
owner_frame,
awaiter_frame,
actionability,
suggestion,
}
}
fn infer_boundary_kind(b: &Boundary) -> BoundaryKind {
let owner_names: Vec<&str> = b.owners.iter().map(|o| o.name.as_str()).collect();
let name_ends_slash = b.name.as_ref().is_some_and(|n| n.ends_with('/'));
if name_ends_slash
|| owner_names.contains(&"LoadingBoundary")
|| owner_names.contains(&"OuterLayoutRouter")
{
return BoundaryKind::RouteSegment;
}
let name_has_suspense = b.name.as_ref().is_some_and(|n| n.contains("Suspense"));
if name_has_suspense || owner_names.iter().any(|n| n.contains("Suspense")) {
return BoundaryKind::ExplicitSuspense;
}
BoundaryKind::Component
}
fn classify_blocker(s: &Suspender, source_frame: Option<&StackFrame>) -> BlockerKind {
let name = s.name.to_lowercase();
match name.as_str() {
"usepathname"
| "useparams"
| "usesearchparams"
| "useselectedlayoutsegments"
| "useselectedlayoutsegment"
| "userouter" => return BlockerKind::ClientHook,
"cookies" | "headers" | "connection" | "params" | "searchparams" | "draftmode" => {
return BlockerKind::RequestApi
}
_ => {}
}
if name == "rsc stream" {
return BlockerKind::Stream;
}
if name.contains("fetch") {
return BlockerKind::ServerFetch;
}
if name.contains("cache") || s.description.to_lowercase().contains("cache") {
return BlockerKind::Cache;
}
if name.starts_with("use") {
return BlockerKind::ClientHook;
}
if let Some(frame) = source_frame {
if is_frameworkish_path(&frame.1) {
return BlockerKind::Framework;
}
}
BlockerKind::Unknown
}
fn suggest_blocker_fix(kind: BlockerKind) -> String {
match kind {
BlockerKind::ClientHook => "Move route hooks behind a smaller client Suspense or provide a real non-null loading fallback for this segment.",
BlockerKind::RequestApi => "Push request-bound reads to a smaller server leaf, or cache around them so the parent shell can stay static.",
BlockerKind::ServerFetch => "Split static shell content from data widgets, then push the fetch into smaller Suspense leaves or cache it.",
BlockerKind::Cache => "This looks cache-related; check whether \"use cache\" or runtime prefetch can eliminate the suspension.",
BlockerKind::Stream => "A stream is still pending here; extract static siblings outside the boundary and push the stream consumer deeper.",
BlockerKind::Framework => "This currently looks framework-driven; find the nearest user-owned caller above it before changing code.",
BlockerKind::Unknown => "Inspect the nearest user-owned owner/awaiter frame and verify whether this suspender really belongs at this boundary.",
}.to_string()
}
fn recommend_fix(
boundary_kind: BoundaryKind,
primary: Option<&ActionableBlocker>,
unknown_suspenders: Option<&str>,
) -> String {
if boundary_kind == BoundaryKind::RouteSegment
&& primary.is_some_and(|p| p.kind == BlockerKind::ClientHook)
{
return "This route segment is suspending on client hooks. Check loading.tsx first; if it is null or visually empty, fix the fallback before chasing deeper push-down work.".to_string();
}
if let Some(p) = primary {
match p.kind {
BlockerKind::ClientHook => {
return "Push the hook-using client UI behind a smaller local Suspense boundary so the parent shell can prerender.".to_string();
}
BlockerKind::RequestApi | BlockerKind::ServerFetch => {
return "Push the request-bound async work into a smaller leaf or split static siblings out of this boundary.".to_string();
}
BlockerKind::Cache => {
return "Check whether caching or runtime prefetch can move this personalized content into the shell.".to_string();
}
BlockerKind::Stream => {
return "Keep the stream behind Suspense, but extract any static shell content outside the boundary.".to_string();
}
BlockerKind::Framework => {
return "The top blocker still looks framework-heavy. Find the nearest user-owned caller before changing boundary placement.".to_string();
}
_ => {}
}
}
if let Some(reason) = unknown_suspenders {
return format!(
"React could not identify the suspender ({}). Investigate the nearest user-owned owner or awaiter frame.",
reason
);
}
"No primary blocker was identified. Inspect the boundary source and owner chain directly."
.to_string()
}
fn pick_preferred_frame(stack: Option<&[StackFrame]>) -> Option<StackFrame> {
let s = stack?;
if s.is_empty() {
return None;
}
s.iter()
.find(|f| !is_frameworkish_path(&f.1))
.cloned()
.or_else(|| s.first().cloned())
}
fn is_frameworkish_path(file: &str) -> bool {
file.contains("/node_modules/")
}
fn build_blocker_key(name: &str, kind: BlockerKind, source_frame: Option<&StackFrame>) -> String {
match source_frame {
None => format!("{}:{}:unknown", kind.label(), name),
Some(f) => format!("{}:{}:{}:{}", kind.label(), name, f.1, f.2),
}
}
fn build_root_causes(holes: &[BoundaryInsight]) -> Vec<RootCauseGroup> {
let mut groups: HashMap<String, RootCauseGroup> = HashMap::new();
for hole in holes {
let Some(blocker) = &hole.primary_blocker else {
continue;
};
let display_name = hole
.name
.clone()
.unwrap_or_else(|| format!("boundary-{}", hole.id));
groups
.entry(blocker.key.clone())
.and_modify(|existing| {
existing.boundary_names.push(display_name.clone());
existing.count += 1;
if blocker.actionability > existing.actionability {
existing.actionability = blocker.actionability;
}
})
.or_insert_with(|| RootCauseGroup {
kind: blocker.kind,
name: blocker.name.clone(),
source_frame: blocker.source_frame.clone(),
boundary_names: vec![display_name],
count: 1,
actionability: blocker.actionability,
suggestion: blocker.suggestion.clone(),
});
}
let mut out: Vec<RootCauseGroup> = groups.into_values().collect();
out.sort_by(|a, b| {
let score_a = (a.count as i32) * a.actionability;
let score_b = (b.count as i32) * b.actionability;
score_b.cmp(&score_a).then_with(|| a.name.cmp(&b.name))
});
out
}
fn collect_files_to_read(holes: &[BoundaryInsight], root_causes: &[RootCauseGroup]) -> Vec<String> {
let mut counts: HashMap<String, i32> = HashMap::new();
let mut add = |f: Option<&str>| {
if let Some(path) = f {
if !path.is_empty() {
*counts.entry(path.to_string()).or_insert(0) += 1;
}
}
};
for hole in holes {
add(hole.source.as_ref().map(|s| s.0.as_str()));
if let Some(pb) = &hole.primary_blocker {
add(pb.source_frame.as_ref().map(|f| f.1.as_str()));
}
for owner in &hole.rendered_by {
add(owner.source.as_ref().map(|s| s.0.as_str()));
}
}
for cause in root_causes {
add(cause.source_frame.as_ref().map(|f| f.1.as_str()));
}
let mut entries: Vec<(String, i32)> = counts.into_iter().collect();
entries.sort_by(|a, b| b.1.cmp(&a.1).then_with(|| a.0.cmp(&b.0)));
entries.into_iter().take(12).map(|(f, _)| f).collect()
}
fn escape_cell(s: &str) -> String {
s.replace('|', "\\|")
}
fn format_report(report: &AnalysisReport, only_dynamic: bool) -> String {
let mut lines: Vec<String> = Vec::new();
lines.push("# Suspense Boundary Analysis".to_string());
if only_dynamic {
lines.push(format!(
"# {} dynamic holes (static boundaries hidden; pass without --only-dynamic to see them)",
report.dynamic_hole_count
));
} else {
lines.push(format!(
"# {} boundaries: {} dynamic holes, {} static",
report.total_boundaries, report.dynamic_hole_count, report.static_count
));
}
lines.push(String::new());
if !report.holes.is_empty() {
lines.push("## Summary".to_string());
if let Some(top) = report.holes.first() {
if let Some(blocker) = &top.primary_blocker {
lines.push(format!(
"- Top actionable hole: {} - {} ({})",
top.name.clone().unwrap_or_else(|| "(unnamed)".into()),
blocker.name,
blocker.kind.label()
));
lines.push(format!("- Suggested next step: {}", top.recommendation));
}
}
if let Some(root) = report.root_causes.first() {
lines.push(format!(
"- Most common root cause: {} ({}) affecting {} boundar{}",
root.name,
root.kind.label(),
root.count,
if root.count == 1 { "y" } else { "ies" }
));
}
lines.push(String::new());
lines.push("## Quick Reference".to_string());
lines.push(
"| Boundary | Type | Primary blocker | Source | Suggested next step |".to_string(),
);
lines.push("| --- | --- | --- | --- | --- |".to_string());
for hole in &report.holes {
let blocker = &hole.primary_blocker;
let source = match blocker.as_ref().and_then(|b| b.source_frame.as_ref()) {
Some(f) => format!("{}:{}", f.1, f.2),
None => match &hole.source {
Some((f, l, _)) => format!("{}:{}", f, l),
None => "unknown".to_string(),
},
};
let blocker_text = match blocker {
Some(b) => format!("{} ({})", b.name, b.kind.label()),
None => "unknown".to_string(),
};
lines.push(format!(
"| {} | {} | {} | {} | {} |",
escape_cell(hole.name.as_deref().unwrap_or("(unnamed)")),
hole.boundary_kind.label(),
escape_cell(&blocker_text),
escape_cell(&source),
escape_cell(&hole.recommendation),
));
}
lines.push(String::new());
if !report.files_to_read.is_empty() {
lines.push("## Files to Read".to_string());
for file in &report.files_to_read {
lines.push(format!("- {}", file));
}
lines.push(String::new());
}
if !report.root_causes.is_empty() {
lines.push("## Root Causes".to_string());
for cause in &report.root_causes {
let source = match &cause.source_frame {
Some(f) => format!("{}:{}", f.1, f.2),
None => "unknown".to_string(),
};
lines.push(format!(
"- {} ({}) at {} - affects {} boundar{}",
cause.name,
cause.kind.label(),
source,
cause.count,
if cause.count == 1 { "y" } else { "ies" }
));
lines.push(format!(" next step: {}", cause.suggestion));
lines.push(format!(" boundaries: {}", cause.boundary_names.join(", ")));
}
lines.push(String::new());
}
}
if !only_dynamic && !report.statics.is_empty() {
lines.push("## Static (not suspended)".to_string());
for b in &report.statics {
let name = b.name.clone().unwrap_or_else(|| "(unnamed)".into());
let src = match &b.source {
Some(s) => format!(" at {}:{}:{}", s.0, s.1, s.2),
None => String::new(),
};
lines.push(format!(" {}{}", name, src));
}
}
lines.join("\n")
}
+67
View File
@@ -0,0 +1,67 @@
//! React component tree snapshot and formatter.
use serde::Deserialize;
#[derive(Debug, Deserialize)]
pub struct TreeNode {
pub id: i64,
#[serde(rename = "type")]
pub node_type: i64,
pub name: Option<String>,
pub key: Option<String>,
pub parent: i64,
}
const HEADER: &str = "# React component tree\n# Columns: depth id parent name [key=...]\n# Use `react inspect <id>` for props/hooks/state. IDs valid until next navigation.";
pub fn format_tree(nodes: &[TreeNode]) -> String {
use std::collections::HashMap;
let mut children: HashMap<i64, Vec<&TreeNode>> = HashMap::new();
for n in nodes {
children.entry(n.parent).or_default().push(n);
}
let mut lines: Vec<String> = vec![HEADER.to_string()];
if let Some(roots) = children.get(&0) {
for root in roots {
walk(root, 0, &children, &mut lines);
}
}
lines.join("\n")
}
fn walk<'a>(
node: &'a TreeNode,
depth: usize,
children: &std::collections::HashMap<i64, Vec<&'a TreeNode>>,
lines: &mut Vec<String>,
) {
let name = node
.name
.clone()
.unwrap_or_else(|| type_name(node.node_type));
let key = match &node.key {
Some(k) => format!(" key={:?}", k),
None => String::new(),
};
let parent = if node.parent == 0 {
"-".to_string()
} else {
node.parent.to_string()
};
lines.push(format!("{} {} {} {}{}", depth, node.id, parent, name, key));
if let Some(cs) = children.get(&node.id) {
for c in cs {
walk(c, depth + 1, children, lines);
}
}
}
fn type_name(t: i64) -> String {
match t {
11 => "Root".to_string(),
12 => "Suspense".to_string(),
13 => "SuspenseList".to_string(),
_ => format!("({})", t),
}
}
+160
View File
@@ -0,0 +1,160 @@
//! Core Web Vitals + React hydration timing report.
//!
//! Universal web-standard metrics (LCP/CLS/TTFB/FCP/INP) via PerformanceObserver
//! and Navigation Timing. When the React profiling build is detected (via
//! `console.timeStamp` entries), also reports hydration phases and per-component
//! hydration timing.
use serde::{Deserialize, Serialize};
#[derive(Debug, Deserialize, Serialize)]
pub struct VitalsData {
pub url: String,
pub ttfb: Option<f64>,
pub lcp: Option<Lcp>,
pub cls: Cls,
pub fcp: Option<f64>,
pub inp: Option<f64>,
pub hydration: Option<HydrationRange>,
pub phases: Vec<Phase>,
#[serde(rename = "hydratedComponents")]
pub hydrated_components: Vec<HydratedComponent>,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct Lcp {
#[serde(rename = "startTime")]
pub start_time: f64,
pub size: Option<i64>,
pub element: Option<String>,
pub url: Option<String>,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct Cls {
pub score: f64,
pub entries: Vec<ClsEntry>,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct ClsEntry {
pub value: f64,
#[serde(rename = "startTime")]
pub start_time: f64,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct HydrationRange {
#[serde(rename = "startTime")]
pub start_time: f64,
#[serde(rename = "endTime")]
pub end_time: f64,
pub duration: f64,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct Phase {
pub label: String,
#[serde(rename = "startTime")]
pub start_time: f64,
#[serde(rename = "endTime")]
pub end_time: f64,
pub duration: f64,
}
#[derive(Debug, Deserialize, Serialize)]
pub struct HydratedComponent {
pub name: String,
#[serde(rename = "startTime")]
pub start_time: f64,
#[serde(rename = "endTime")]
pub end_time: f64,
pub duration: f64,
}
pub fn format_vitals_report(d: &VitalsData) -> String {
let mut lines: Vec<String> = Vec::new();
lines.push(format!("# Page Load Profile - {}", d.url));
lines.push(String::new());
lines.push("## Core Web Vitals".to_string());
let ttfb_str = match d.ttfb {
Some(t) => format!("{}ms", t),
None => "-".to_string(),
};
lines.push(format!(" TTFB {:>10}", ttfb_str));
match &d.lcp {
Some(lcp) => {
let label = match (&lcp.element, &lcp.url) {
(Some(el), Some(url)) => {
let url_trunc: String = url.chars().take(60).collect();
format!(" ({}: {})", el, url_trunc)
}
(Some(el), None) => format!(" ({})", el),
_ => String::new(),
};
lines.push(format!(
" LCP {:>10}{}",
format!("{}ms", lcp.start_time),
label
));
}
None => lines.push(" LCP -".to_string()),
}
lines.push(format!(" CLS {:>10}", d.cls.score));
if let Some(fcp) = d.fcp {
lines.push(format!(" FCP {:>10}", format!("{}ms", fcp)));
}
if let Some(inp) = d.inp {
lines.push(format!(" INP {:>10}", format!("{}ms", inp)));
}
lines.push(String::new());
match &d.hydration {
Some(h) => lines.push(format!(
"## React Hydration - {}ms ({}ms -> {}ms)",
h.duration, h.start_time, h.end_time
)),
None => {
lines.push("## React Hydration - no data (requires React profiling build)".to_string())
}
}
if !d.phases.is_empty() {
for p in &d.phases {
lines.push(format!(
" {:<28} {:>10} ({} -> {})",
p.label,
format!("{}ms", p.duration),
p.start_time,
p.end_time
));
}
lines.push(String::new());
}
if !d.hydrated_components.is_empty() {
lines.push(format!(
"## Hydrated components ({} total, sorted by duration)",
d.hydrated_components.len()
));
for c in d.hydrated_components.iter().take(30) {
lines.push(format!(
" {:<40} {:>10}",
c.name,
format!("{}ms", c.duration)
));
}
if d.hydrated_components.len() > 30 {
lines.push(format!(
" ... and {} more",
d.hydrated_components.len() - 30
));
}
}
lines.join("\n")
}
+254 -5
View File
@@ -80,6 +80,7 @@ pub struct SnapshotOptions {
pub interactive: bool,
pub compact: bool,
pub depth: Option<usize>,
pub urls: bool,
}
struct TreeNode {
@@ -98,7 +99,8 @@ struct TreeNode {
has_ref: bool,
ref_id: Option<String>,
depth: usize,
cursor_info: Option<CursorElementInfo>, // cursor-interactive information
cursor_info: Option<CursorElementInfo>,
url: Option<String>,
}
impl TreeNode {
@@ -121,10 +123,10 @@ impl TreeNode {
ref_id: None,
depth: 0,
cursor_info: None,
url: None,
}
}
// Clear node content
fn clear(&mut self) {
self.role = String::new();
self.name = String::new();
@@ -139,18 +141,45 @@ impl TreeNode {
self.children.clear();
self.parent_idx = None;
self.has_ref = false;
self.url = None;
self.ref_id = None;
self.depth = 0;
self.cursor_info = None;
}
}
/// The type of a hidden form input found inside a cursor-interactive element.
#[derive(Clone, Copy)]
enum HiddenInputKind {
Radio,
Checkbox,
}
impl HiddenInputKind {
fn parse(s: &str) -> Option<Self> {
match s {
"radio" => Some(Self::Radio),
"checkbox" => Some(Self::Checkbox),
_ => None,
}
}
fn as_role(&self) -> &str {
match self {
Self::Radio => "radio",
Self::Checkbox => "checkbox",
}
}
}
/// Information about a cursor-interactive element (elements with cursor:pointer, onclick, tabindex, etc.)
#[derive(Clone)]
struct CursorElementInfo {
kind: String, // "clickable", "focusable", "editable"
hints: Vec<String>,
text: String, // textContent from the DOM element (fallback when ARIA name is empty)
hidden_input_kind: Option<HiddenInputKind>,
hidden_input_checked: Option<String>, // "true", "false", or "mixed" (tristate)
}
struct RoleNameTracker {
@@ -274,7 +303,7 @@ pub async fn take_snapshot(
)
.await?;
let (tree_nodes, root_indices) = build_tree(&ax_tree.nodes);
let (mut tree_nodes, root_indices) = build_tree(&ax_tree.nodes);
// When a selector is given, find AX nodes whose backendDOMNodeId falls
// within the target DOM subtree and pick the top-level ones as roots.
@@ -320,6 +349,8 @@ pub async fn take_snapshot(
.await
.unwrap_or_default();
promote_hidden_inputs(&mut tree_nodes, &cursor_elements);
for (idx, node) in tree_nodes.iter().enumerate() {
let role = node.role.as_str();
let mut should_ref = if INTERACTIVE_ROLES.contains(&role) {
@@ -346,7 +377,6 @@ pub async fn take_snapshot(
let duplicates = tracker.get_duplicates();
let mut tree_nodes = tree_nodes;
for (idx, nth) in &nodes_with_refs {
let node = &tree_nodes[*idx];
let key = format!("{}:{}", node.role, node.name);
@@ -383,6 +413,75 @@ pub async fn take_snapshot(
ref_map.set_next_ref_num(next_ref);
if options.urls {
let link_nodes: Vec<(usize, i64)> = tree_nodes
.iter()
.enumerate()
.filter(|(_, n)| n.role == "link" && n.has_ref && n.backend_node_id.is_some())
.filter_map(|(i, n)| n.backend_node_id.map(|bid| (i, bid)))
.collect();
if !link_nodes.is_empty() {
// CDP has no batch resolve API, so we parallelize individual calls.
// Phase 1: resolve all backend node IDs to JS object IDs in parallel.
let resolve_futs = link_nodes.iter().map(|&(idx, bid)| async move {
let resolved = client
.send_command(
"DOM.resolveNode",
Some(serde_json::json!({ "backendNodeId": bid })),
Some(session_id),
)
.await;
let obj_id = resolved.ok().and_then(|r| {
r.get("object")
.and_then(|o| o.get("objectId"))
.and_then(|v| v.as_str())
.map(|s| s.to_string())
});
(idx, obj_id)
});
let resolved: Vec<(usize, Option<String>)> =
futures_util::future::join_all(resolve_futs).await;
// Phase 2: fetch hrefs for all resolved objects in parallel.
let href_futs: Vec<_> = resolved
.iter()
.filter_map(|(idx, obj_id)| {
let oid = obj_id.as_ref()?;
Some(async move {
let result = client
.send_command(
"Runtime.callFunctionOn",
Some(serde_json::json!({
"objectId": oid,
"functionDeclaration": "function() { return this.href || ''; }",
"returnByValue": true,
})),
Some(session_id),
)
.await;
let href = result.ok().and_then(|r| {
r.get("result")
.and_then(|r| r.get("value"))
.and_then(|v| v.as_str())
.filter(|s| !s.is_empty())
.map(|s| s.to_string())
});
(*idx, href)
})
})
.collect();
let hrefs: Vec<(usize, Option<String>)> =
futures_util::future::join_all(href_futs).await;
for (idx, href) in hrefs {
if let Some(url) = href {
tree_nodes[idx].url = Some(url);
}
}
}
}
let mut output = String::new();
for &root_idx in &effective_roots {
render_tree(&tree_nodes, root_idx, 0, &mut output, options);
@@ -567,6 +666,23 @@ async fn find_cursor_interactive_elements(
var rect = el.getBoundingClientRect();
if (rect.width === 0 || rect.height === 0) continue;
// Detect hidden radio/checkbox inputs inside this element (common pattern:
// <label> wrapping a display:none <input type="radio"> styled as a card).
// Note: we only check display/visibility/hidden, NOT opacity:0 or sr-only,
// because those inputs remain in Chrome's AX tree and already appear as
// role="radio" without promotion.
var hiddenInputType = null;
var hiddenInputChecked = null;
var hiddenInput = el.querySelector('input[type="radio"], input[type="checkbox"]');
if (hiddenInput) {
var hiddenInputStyle = getComputedStyle(hiddenInput);
var isInputHidden = hiddenInputStyle.display === 'none' || hiddenInputStyle.visibility === 'hidden' || hiddenInput.hidden;
if (isInputHidden) {
hiddenInputType = hiddenInput.type;
hiddenInputChecked = hiddenInput.indeterminate ? 'mixed' : String(hiddenInput.checked);
}
}
el.setAttribute('data-__ab-ci', String(results.length));
results.push({
text: text,
@@ -574,7 +690,9 @@ async fn find_cursor_interactive_elements(
hasOnClick: hasOnClick,
hasCursorPointer: hasCursorPointer,
hasTabIndex: hasTabIndex,
isEditable: isEditable
isEditable: isEditable,
hiddenInputType: hiddenInputType,
hiddenInputChecked: hiddenInputChecked
});
}
return results;
@@ -747,6 +865,15 @@ async fn find_cursor_interactive_elements(
.trim()
.to_string();
let hidden_input_kind = elem
.get("hiddenInputType")
.and_then(|v| v.as_str())
.and_then(HiddenInputKind::parse);
let hidden_input_checked = elem
.get("hiddenInputChecked")
.and_then(|v| v.as_str())
.map(|s| s.to_string());
if let Some(bid) = backend_node_id {
map.insert(
bid,
@@ -754,6 +881,8 @@ async fn find_cursor_interactive_elements(
kind: kind.to_string(),
hints,
text,
hidden_input_kind,
hidden_input_checked,
},
);
}
@@ -762,6 +891,38 @@ async fn find_cursor_interactive_elements(
Ok(map)
}
/// Promote LabelText/generic nodes that wrap a hidden radio/checkbox input.
/// When a `<label>` contains a `display:none` `<input type="radio">`, Chrome excludes
/// the input from the AX tree entirely, leaving only the label with role="LabelText"
/// and an empty name. We detect these via cursor-interactive scanning and promote
/// the label to the correct input role so consumers see role="radio" in data.refs.
fn promote_hidden_inputs(
tree_nodes: &mut [TreeNode],
cursor_elements: &HashMap<i64, CursorElementInfo>,
) {
for node in tree_nodes.iter_mut() {
if !matches!(node.role.as_str(), "LabelText" | "generic") {
continue;
}
let cursor_info = match node
.backend_node_id
.and_then(|bid| cursor_elements.get(&bid))
{
Some(info) => info,
None => continue,
};
if let Some(input_kind) = cursor_info.hidden_input_kind {
node.role = input_kind.as_role().to_string();
if node.name.is_empty() && !cursor_info.text.is_empty() {
node.name = cursor_info.text.clone();
}
if let Some(ref checked) = cursor_info.hidden_input_checked {
node.checked = Some(checked.clone());
}
}
}
}
fn build_tree(nodes: &[AXNode]) -> (Vec<TreeNode>, Vec<usize>) {
let mut tree_nodes: Vec<TreeNode> = Vec::with_capacity(nodes.len());
let mut id_to_idx: HashMap<String, usize> = HashMap::new();
@@ -797,6 +958,7 @@ fn build_tree(nodes: &[AXNode]) -> (Vec<TreeNode>, Vec<usize>) {
ref_id: None,
depth: 0,
cursor_info: None,
url: None,
});
id_to_idx.insert(node.node_id.clone(), i);
}
@@ -993,6 +1155,10 @@ fn render_tree(
attrs.push(format!("ref={}", ref_id));
}
if let Some(ref url) = node.url {
attrs.push(format!("url={}", url));
}
if !attrs.is_empty() {
line.push_str(&format!(" [{}]", attrs.join(", ")));
}
@@ -1334,4 +1500,87 @@ mod tests {
assert_eq!(session, parent_session);
assert_eq!(params, serde_json::json!({}));
}
// -----------------------------------------------------------------------
// promote_hidden_inputs
// -----------------------------------------------------------------------
fn make_node(role: &str, name: &str, backend_node_id: Option<i64>) -> TreeNode {
let mut node = TreeNode::empty();
node.role = role.to_string();
node.name = name.to_string();
node.backend_node_id = backend_node_id;
node
}
fn make_cursor_info(
hidden_kind: Option<HiddenInputKind>,
hidden_checked: Option<&str>,
text: &str,
) -> CursorElementInfo {
CursorElementInfo {
kind: "clickable".to_string(),
hints: vec!["cursor:pointer".to_string()],
text: text.to_string(),
hidden_input_kind: hidden_kind,
hidden_input_checked: hidden_checked.map(|s| s.to_string()),
}
}
#[test]
fn test_promote_label_with_hidden_radio() {
let mut nodes = vec![
make_node("LabelText", "", Some(1)),
make_node("LabelText", "", Some(2)),
make_node("button", "Submit", Some(3)),
];
let mut cursor_elements = HashMap::new();
cursor_elements.insert(
1,
make_cursor_info(Some(HiddenInputKind::Radio), Some("false"), "Option A"),
);
cursor_elements.insert(
2,
make_cursor_info(Some(HiddenInputKind::Radio), Some("true"), "Option B"),
);
promote_hidden_inputs(&mut nodes, &cursor_elements);
assert_eq!(nodes[0].role, "radio");
assert_eq!(nodes[0].name, "Option A");
assert_eq!(nodes[0].checked, Some("false".to_string()));
assert_eq!(nodes[1].role, "radio");
assert_eq!(nodes[1].name, "Option B");
assert_eq!(nodes[1].checked, Some("true".to_string()));
// button should be untouched
assert_eq!(nodes[2].role, "button");
}
#[test]
fn test_promote_preserves_existing_name() {
// If AX tree already has a name, don't overwrite with textContent
let mut nodes = vec![make_node("LabelText", "AX Name", Some(1))];
let mut cursor_elements = HashMap::new();
cursor_elements.insert(
1,
make_cursor_info(Some(HiddenInputKind::Radio), Some("false"), "Text Content"),
);
promote_hidden_inputs(&mut nodes, &cursor_elements);
assert_eq!(nodes[0].role, "radio");
assert_eq!(nodes[0].name, "AX Name"); // preserved, not overwritten
}
#[test]
fn test_promote_skips_without_hidden_input() {
// Cursor-interactive label WITHOUT a hidden input should not be promoted
let mut nodes = vec![make_node("LabelText", "", Some(1))];
let mut cursor_elements = HashMap::new();
cursor_elements.insert(1, make_cursor_info(None, None, "Click me"));
promote_hidden_inputs(&mut nodes, &cursor_elements);
assert_eq!(nodes[0].role, "LabelText"); // unchanged
}
}
+10 -3
View File
@@ -714,14 +714,21 @@ pub fn dispatch_state_command(cmd: &Value) -> Option<Result<Value, String>> {
}
}
pub fn get_sessions_dir() -> PathBuf {
/// Return the agent-browser state root (`~/.agent-browser`, falling back to
/// `<tempdir>/agent-browser` when the home directory can't be resolved).
/// This is the parent of `sessions/`, auth storage, and the encryption key.
pub fn get_state_dir() -> PathBuf {
if let Some(home) = dirs::home_dir() {
home.join(".agent-browser").join("sessions")
home.join(".agent-browser")
} else {
std::env::temp_dir().join("agent-browser").join("sessions")
std::env::temp_dir().join("agent-browser")
}
}
pub fn get_sessions_dir() -> PathBuf {
get_state_dir().join("sessions")
}
#[cfg(test)]
mod tests {
use super::*;
+237
View File
@@ -0,0 +1,237 @@
//! Stealth anti-detection module.
//!
//! Injects browser-level patches to evade bot detection (creepjs, sannysoft,
//! Cloudflare Turnstile, etc.) by normalizing fingerprint signals that betray
//! headless or automated Chrome instances.
use serde_json::json;
use super::cdp::client::CdpClient;
/// Full stealth JS payload compiled at build time (for --launch mode).
const STEALTH_SCRIPTS_RAW: &str = include_str!("stealth_scripts.js");
/// Minimal stealth script for CDP-attach mode (connecting to user's real Chrome).
/// Only removes navigator.webdriver — the browser's own fingerprint is already real.
/// Minimal stealth script for CDP-attach mode.
/// Emulation.setAutomationOverride handles navigator.webdriver at the native
/// level, so no JS patching is needed in CdpAttach mode. An empty script
/// avoids creating any detectable lie-props artifacts.
const MINIMAL_STEALTH_SCRIPT: &str = "";
/// Chrome launch arguments that reduce automation fingerprint surface.
pub const STEALTH_CHROMIUM_ARGS: &[&str] = &[
"--disable-blink-features=AutomationControlled",
"--use-gl=angle",
"--use-angle=default",
];
/// Connection mode determines which stealth patches to apply.
#[derive(Clone, Copy, PartialEq)]
pub enum StealthMode {
/// Connected to user's real Chrome — minimal patches only (webdriver removal).
/// The browser already has a real fingerprint; heavy patches would create detectable lies.
CdpAttach,
/// Launched a new Chrome instance — apply full stealth patches.
FullLaunch,
}
/// Build the stealth JS payload for the given mode and locale.
pub fn build_stealth_script(mode: StealthMode, locale: Option<&str>) -> String {
if mode == StealthMode::CdpAttach {
return MINIMAL_STEALTH_SCRIPT.to_string();
}
// Full launch mode: inject all patches
let locale = locale.unwrap_or("en-US");
let base_lang = locale.split('-').next().unwrap_or(locale);
let languages: Vec<&str> = if base_lang == locale {
vec![locale]
} else {
vec![locale, base_lang]
};
let config_line = format!(
r#"const __abStealth = {{ locale: "{}", languages: {}, allowWebGLContextFallback: false }};"#,
locale,
serde_json::to_string(&languages).unwrap_or_else(|_| r#"["en-US","en"]"#.to_string()),
);
if let Some(rest) = STEALTH_SCRIPTS_RAW.strip_prefix(
r#"const __abStealth = { locale: "en-US", languages: ["en-US", "en"], allowWebGLContextFallback: false };"#,
) {
format!("{}{}", config_line, rest)
} else {
format!("{}\n{}", config_line, STEALTH_SCRIPTS_RAW)
}
}
/// Apply stealth patches to a browser session.
///
/// In `CdpAttach` mode (user's real Chrome): only removes `navigator.webdriver`.
/// In `FullLaunch` mode (new Chrome): injects all 32 patches + UA override.
pub async fn apply_stealth(
client: &CdpClient,
session_id: &str,
mode: StealthMode,
locale: Option<&str>,
) -> Result<(), String> {
// First: disable the automation flag at the CDP protocol level.
// This tells Chrome to natively set navigator.webdriver = false,
// which is undetectable by lie-detection systems like CreepJS.
// Falls back gracefully on older Chrome versions that don't support this.
let _ = client
.send_command(
"Emulation.setAutomationOverride",
Some(json!({ "enabled": false })),
Some(session_id),
)
.await;
let script = build_stealth_script(mode, locale);
// Inject stealth scripts to run before page JS
client
.send_command(
"Page.addScriptToEvaluateOnNewDocument",
Some(json!({ "source": script })),
Some(session_id),
)
.await?;
// In full launch mode, also override UA to remove HeadlessChrome marker
if mode == StealthMode::FullLaunch {
let ua = get_browser_user_agent(client, session_id).await;
if let Some(ua) = ua {
let cleaned = ua.replace("HeadlessChrome", "Chrome");
if cleaned != ua {
client
.send_command(
"Emulation.setUserAgentOverride",
Some(json!({
"userAgent": cleaned,
"acceptLanguage": locale.unwrap_or("en-US"),
"platform": platform_string(),
"userAgentMetadata": build_ua_metadata(&cleaned, locale),
})),
Some(session_id),
)
.await?;
}
}
}
Ok(())
}
/// Get the browser's User-Agent string via CDP.
async fn get_browser_user_agent(client: &CdpClient, session_id: &str) -> Option<String> {
let result = client
.send_command(
"Runtime.evaluate",
Some(json!({ "expression": "navigator.userAgent", "returnByValue": true })),
Some(session_id),
)
.await
.ok()?;
result
.get("result")
.and_then(|r| r.get("value"))
.and_then(|v| v.as_str())
.map(String::from)
}
/// Also run stealth script on the current page (for already-loaded pages after CDP attach).
pub async fn apply_stealth_to_current_page(
client: &CdpClient,
session_id: &str,
mode: StealthMode,
locale: Option<&str>,
) -> Result<(), String> {
let script = build_stealth_script(mode, locale);
client
.send_command(
"Runtime.evaluate",
Some(json!({
"expression": script,
"returnByValue": true,
})),
Some(session_id),
)
.await?;
Ok(())
}
/// Strip sourceURL comments from CDP expressions to avoid leaking
/// automation-framework identifiers in stack traces.
pub fn strip_source_url_labels(input: &str) -> String {
// Remove //# sourceURL=... and //@ sourceURL=...
let re_line = regex_lite::Regex::new(r"(?i)\n?\s*//[@#]\s*sourceURL=[^\n\r]*").unwrap();
let output = re_line.replace_all(input, "");
// Remove /*# sourceURL=...*/ block comments
let re_block =
regex_lite::Regex::new(r"(?is)\n?\s*/\*[@#]\s*sourceURL=[\s\S]*?\*/").unwrap();
re_block.replace_all(&output, "").to_string()
}
fn platform_string() -> &'static str {
if cfg!(target_os = "macos") {
"macOS"
} else if cfg!(target_os = "windows") {
"Win32"
} else {
"Linux"
}
}
fn platform_hint() -> &'static str {
if cfg!(target_os = "macos") {
"macOS"
} else if cfg!(target_os = "windows") {
"Windows"
} else {
"Linux"
}
}
fn platform_version_hint() -> &'static str {
if cfg!(target_os = "macos") {
"14.0.0"
} else if cfg!(target_os = "windows") {
"10.0.0"
} else {
"6.5.0"
}
}
fn build_ua_metadata(ua: &str, locale: Option<&str>) -> serde_json::Value {
// Extract Chrome version from UA string
let chrome_version = ua
.split("Chrome/")
.nth(1)
.and_then(|s| s.split_whitespace().next())
.unwrap_or("130.0.0.0");
let major = chrome_version.split('.').next().unwrap_or("130");
let _lang = locale.unwrap_or("en-US");
json!({
"brands": [
{ "brand": "Chromium", "version": major },
{ "brand": "Google Chrome", "version": major },
{ "brand": "Not?A_Brand", "version": "99" },
],
"fullVersionList": [
{ "brand": "Chromium", "version": chrome_version },
{ "brand": "Google Chrome", "version": chrome_version },
{ "brand": "Not?A_Brand", "version": "99.0.0.0" },
],
"fullVersion": chrome_version,
"platform": platform_hint(),
"platformVersion": platform_version_hint(),
"architecture": if cfg!(target_arch = "aarch64") { "arm" } else { "x86" },
"model": "",
"mobile": false,
"bitness": "64",
"wow64": false,
})
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+325
View File
@@ -0,0 +1,325 @@
use serde_json::{json, Value};
use std::sync::Arc;
use tokio::sync::{broadcast, watch, Mutex, RwLock};
use crate::native::cdp::client::CdpClient;
use crate::native::network;
use super::timestamp_ms;
/// Background task that subscribes to CDP events and broadcasts screencast frames in real-time.
/// Also handles auto-start/stop of screencast based on WebSocket client count.
#[allow(clippy::too_many_arguments)]
pub(super) async fn cdp_event_loop(
frame_tx: broadcast::Sender<String>,
client_slot: Arc<RwLock<Option<Arc<CdpClient>>>>,
client_notify: Arc<tokio::sync::Notify>,
screencasting: Arc<Mutex<bool>>,
client_count: Arc<Mutex<usize>>,
cdp_session_id: Arc<RwLock<Option<String>>>,
viewport_width: Arc<Mutex<u32>>,
viewport_height: Arc<Mutex<u32>>,
last_frame: Arc<RwLock<Option<String>>>,
last_tabs: Arc<RwLock<Vec<Value>>>,
last_engine: Arc<RwLock<String>>,
recording: Arc<Mutex<bool>>,
mut shutdown_rx: watch::Receiver<bool>,
) {
loop {
tokio::select! {
changed = shutdown_rx.changed() => {
if changed.is_err() || *shutdown_rx.borrow() {
let session_id = cdp_session_id.read().await.clone();
if *screencasting.lock().await {
if let Some(ref client) = *client_slot.read().await {
let _ = client
.send_command_no_params("Page.stopScreencast", session_id.as_deref())
.await;
}
let mut sc = screencasting.lock().await;
*sc = false;
}
return;
}
}
_ = client_notify.notified() => {}
}
let count = *client_count.lock().await;
let guard = client_slot.read().await;
if count > 0 {
if let Some(ref client) = *guard {
let mut event_rx = client.subscribe();
let client_arc = Arc::clone(client);
drop(guard);
let session_id = cdp_session_id.read().await.clone();
let vw = *viewport_width.lock().await;
let vh = *viewport_height.lock().await;
let eng = last_engine.read().await.clone();
let supports_screencast = eng == "chrome";
if supports_screencast {
let _ = client_arc
.send_command(
"Page.startScreencast",
Some(json!({
"format": "jpeg",
"quality": 80,
"maxWidth": vw,
"maxHeight": vh,
"everyNthFrame": 1,
})),
session_id.as_deref(),
)
.await;
}
{
let mut sc = screencasting.lock().await;
*sc = supports_screencast;
}
let rec = *recording.lock().await;
let status = json!({
"type": "status",
"connected": true,
"screencasting": supports_screencast,
"viewportWidth": vw,
"viewportHeight": vh,
"engine": eng,
"recording": rec,
});
let _ = frame_tx.send(status.to_string());
loop {
tokio::select! {
changed = shutdown_rx.changed() => {
if changed.is_err() || *shutdown_rx.borrow() {
if supports_screencast {
let session_id = cdp_session_id.read().await.clone();
let _ = client_arc
.send_command_no_params("Page.stopScreencast", session_id.as_deref())
.await;
}
let mut sc = screencasting.lock().await;
*sc = false;
return;
}
}
event = event_rx.recv() => {
match event {
Ok(evt) => {
if evt.method == "Page.frameNavigated" {
if let Some(frame) = evt.params.get("frame") {
let is_main = frame
.get("parentId")
.and_then(|v| v.as_str())
.is_none_or(|s| s.is_empty());
if is_main {
if let Some(url) = frame.get("url").and_then(|v| v.as_str()) {
{
let mut tabs = last_tabs.write().await;
for tab in tabs.iter_mut() {
if tab.get("active").and_then(|v| v.as_bool()).unwrap_or(false) {
tab.as_object_mut().map(|o| o.insert("url".to_string(), json!(url)));
}
}
}
let msg = json!({
"type": "url",
"url": url,
"timestamp": timestamp_ms(),
});
let _ = frame_tx.send(msg.to_string());
}
}
}
} else if evt.method == "Page.screencastFrame" {
if let Some(sid) = evt.params.get("sessionId").and_then(|v| v.as_i64()) {
let _ = client_arc.send_command(
"Page.screencastFrameAck",
Some(json!({ "sessionId": sid })),
evt.session_id.as_deref(),
).await;
}
if let Some(data) = evt.params.get("data").and_then(|v| v.as_str()) {
let meta = evt.params.get("metadata");
let msg = json!({
"type": "frame",
"data": data,
"metadata": {
"offsetTop": meta.and_then(|m| m.get("offsetTop")).and_then(|v| v.as_f64()).unwrap_or(0.0),
"pageScaleFactor": meta.and_then(|m| m.get("pageScaleFactor")).and_then(|v| v.as_f64()).unwrap_or(1.0),
"deviceWidth": vw,
"deviceHeight": vh,
"scrollOffsetX": meta.and_then(|m| m.get("scrollOffsetX")).and_then(|v| v.as_f64()).unwrap_or(0.0),
"scrollOffsetY": meta.and_then(|m| m.get("scrollOffsetY")).and_then(|v| v.as_f64()).unwrap_or(0.0),
"timestamp": meta.and_then(|m| m.get("timestamp")).and_then(|v| v.as_u64()).unwrap_or(0),
}
});
let msg_str = msg.to_string();
{
let mut lf = last_frame.write().await;
*lf = Some(msg_str.clone());
}
let _ = frame_tx.send(msg_str);
}
} else if evt.method == "Runtime.consoleAPICalled" {
let level = evt.params.get("type")
.and_then(|v| v.as_str())
.unwrap_or("log");
let raw_args = evt.params.get("args")
.and_then(|v| v.as_array())
.cloned()
.unwrap_or_default();
let text = network::format_console_args(&raw_args);
if !text.is_empty() {
let mut msg = json!({
"type": "console",
"level": level,
"text": text,
"timestamp": timestamp_ms(),
});
if !raw_args.is_empty() {
msg.as_object_mut().unwrap().insert(
"args".to_string(),
Value::Array(raw_args),
);
}
let _ = frame_tx.send(msg.to_string());
}
} else if evt.method == "Runtime.exceptionThrown" {
let text = evt.params.get("exceptionDetails")
.and_then(|d| {
d.get("exception")
.and_then(|e| e.get("description").and_then(|v| v.as_str()))
.or_else(|| d.get("text").and_then(|v| v.as_str()))
})
.unwrap_or("Unknown error");
let line = evt.params.get("exceptionDetails")
.and_then(|d| d.get("lineNumber").and_then(|v| v.as_i64()));
let column = evt.params.get("exceptionDetails")
.and_then(|d| d.get("columnNumber").and_then(|v| v.as_i64()));
let msg = json!({
"type": "page_error",
"text": text,
"line": line,
"column": column,
"timestamp": timestamp_ms(),
});
let _ = frame_tx.send(msg.to_string());
}
}
Err(broadcast::error::RecvError::Lagged(_)) => continue,
Err(broadcast::error::RecvError::Closed) => break,
}
}
_ = client_notify.notified() => {
let count = *client_count.lock().await;
let new_session_id = cdp_session_id.read().await.clone();
if count == 0 {
if supports_screencast {
let _ = client_arc
.send_command_no_params("Page.stopScreencast", session_id.as_deref())
.await;
}
let mut sc = screencasting.lock().await;
*sc = false;
break;
}
let client_changed = {
let guard = client_slot.read().await;
let same = guard
.as_ref()
.is_some_and(|c| Arc::ptr_eq(c, &client_arc));
!same
};
let session_changed = new_session_id != session_id;
let new_vw = *viewport_width.lock().await;
let new_vh = *viewport_height.lock().await;
let viewport_changed = new_vw != vw || new_vh != vh;
if client_changed || session_changed || viewport_changed {
if supports_screencast {
let _ = client_arc
.send_command_no_params("Page.stopScreencast", session_id.as_deref())
.await;
}
let mut sc = screencasting.lock().await;
*sc = false;
client_notify.notify_one();
break;
}
}
}
}
} else {
drop(guard);
}
} else {
let was_screencasting = *screencasting.lock().await;
if was_screencasting {
if let Some(ref client) = *guard {
let session_id = cdp_session_id.read().await.clone();
let _ = client
.send_command_no_params("Page.stopScreencast", session_id.as_deref())
.await;
}
let mut sc = screencasting.lock().await;
*sc = false;
}
drop(guard);
}
}
}
pub async fn start_screencast(
client: &CdpClient,
session_id: &str,
format: &str,
quality: i32,
max_width: i32,
max_height: i32,
) -> Result<(), String> {
client
.send_command(
"Page.startScreencast",
Some(json!({
"format": format,
"quality": quality,
"maxWidth": max_width,
"maxHeight": max_height,
"everyNthFrame": 1,
})),
Some(session_id),
)
.await?;
Ok(())
}
pub async fn stop_screencast(client: &CdpClient, session_id: &str) -> Result<(), String> {
client
.send_command_no_params("Page.stopScreencast", Some(session_id))
.await?;
Ok(())
}
pub async fn ack_screencast_frame(
client: &CdpClient,
session_id: &str,
screencast_session_id: i64,
) -> Result<(), String> {
client
.send_command(
"Page.screencastFrameAck",
Some(json!({ "sessionId": screencast_session_id })),
Some(session_id),
)
.await?;
Ok(())
}
+970
View File
@@ -0,0 +1,970 @@
use std::sync::OnceLock;
use serde_json::{json, Value};
use tokio::io::AsyncWriteExt;
use super::http::cors_headers_for_origin;
pub(crate) const DEFAULT_AI_GATEWAY_URL: &str = "https://ai-gateway.vercel.sh";
static HTTP_CLIENT: OnceLock<reqwest::Client> = OnceLock::new();
pub(crate) fn http_client() -> &'static reqwest::Client {
HTTP_CLIENT.get_or_init(reqwest::Client::new)
}
pub(crate) fn is_chat_enabled() -> bool {
std::env::var("AI_GATEWAY_API_KEY").is_ok()
}
pub(super) fn chat_status_json() -> String {
let enabled = is_chat_enabled();
let mut obj = json!({ "enabled": enabled });
if enabled {
if let Ok(model) = std::env::var("AI_GATEWAY_MODEL") {
obj["model"] = Value::String(model);
}
}
obj.to_string()
}
pub(super) async fn handle_models_request(
stream: &mut tokio::net::TcpStream,
origin: Option<&str>,
) {
let cors = cors_headers_for_origin(origin);
let gateway_url = std::env::var("AI_GATEWAY_URL")
.unwrap_or_else(|_| DEFAULT_AI_GATEWAY_URL.to_string())
.trim_end_matches('/')
.to_string();
let api_key = match std::env::var("AI_GATEWAY_API_KEY") {
Ok(k) => k,
Err(_) => {
let body = r#"{"data":[]}"#;
let resp = format!(
"HTTP/1.1 200 OK\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n{cors}\r\n",
body.len()
);
let _ = stream.write_all(resp.as_bytes()).await;
let _ = stream.write_all(body.as_bytes()).await;
return;
}
};
let url = format!("{}/v1/models", gateway_url);
let client = http_client();
let result = client
.get(&url)
.header("Authorization", format!("Bearer {}", api_key))
.send()
.await;
let body = match result {
Ok(r) if r.status().is_success() => r
.text()
.await
.unwrap_or_else(|_| r#"{"data":[]}"#.to_string()),
_ => r#"{"data":[]}"#.to_string(),
};
let resp = format!(
"HTTP/1.1 200 OK\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n{cors}\r\n",
body.len()
);
let _ = stream.write_all(resp.as_bytes()).await;
let _ = stream.write_all(body.as_bytes()).await;
}
const SKILL_NAMES: &[&str] = &["agent-browser", "slack", "electron", "dogfood", "agentcore"];
/// Locate the `skills/` directory by walking up from the executable.
/// Works for npm installs (binary in `bin/`, skills at `../skills/`) and
/// dev builds (binary deep in `cli/target/`, skills at repo root).
fn find_skills_dir() -> Option<std::path::PathBuf> {
let exe = std::env::current_exe().ok()?;
let real = exe.canonicalize().unwrap_or(exe);
let mut dir = real.parent();
while let Some(d) = dir {
let candidate = d.join("skills");
if candidate.join("agent-browser").join("SKILL.md").exists() {
return Some(candidate);
}
dir = d.parent();
}
None
}
fn load_skills() -> Vec<(String, String)> {
let Some(skills_dir) = find_skills_dir() else {
return Vec::new();
};
SKILL_NAMES
.iter()
.filter_map(|name| {
let path = skills_dir.join(name).join("SKILL.md");
let content = std::fs::read_to_string(&path).ok()?;
Some((name.to_string(), content))
})
.collect()
}
fn strip_frontmatter(s: &str) -> &str {
if !s.starts_with("---") {
return s;
}
if let Some(end) = s[3..].find("---") {
let after = &s[3 + end + 3..];
after.trim_start_matches(['\n', '\r'])
} else {
s
}
}
pub(crate) fn get_system_prompt() -> &'static str {
static PROMPT: OnceLock<String> = OnceLock::new();
PROMPT.get_or_init(|| {
let skills = load_skills();
let mut sections = String::new();
for (name, content) in &skills {
let body = strip_frontmatter(content);
sections.push_str(&format!("\n\n<skill name=\"{}\">\n{}\n</skill>", name, body.trim()));
}
format!(
r#"You are an AI assistant that controls a browser through agent-browser. You have an active browser session, but you can also create new sessions.
RULES:
- You MUST use the agent_browser tool for every browser action. NEVER claim you performed an action without calling the tool.
- If the user asks you to do something, call the tool first, then describe the result.
- If a request is outside your capabilities (e.g. system operations), say so honestly. Do not improvise or pretend.
- One tool call per command. Do not chain with `&&` or `;`.
- Do not add `--json`.
- Do not run non-agent-browser programs.
- Keep responses concise.
- For screenshots, omit the path argument so they save to the default location (which will be displayed inline). Screenshots from tool calls are ALREADY shown to the user. Do NOT re-display them with markdown image syntax in your text response. Never use `![...]()` to reference screenshots.
- To create a new session: add `--session <name>` to any command (e.g. `agent-browser --session my-session open https://example.com`). If the session does not exist, it will be created automatically.
- To use a different browser engine: add `--engine <engine>` (e.g. `agent-browser --session lp-session --engine lightpanda open https://example.com`). Supported engines: chrome (default), lightpanda.
The following skill references describe agent-browser capabilities in detail. Use them when deciding which commands to run and how to approach tasks.
{sections}"#,
)
})
}
pub(crate) const CHAT_TOOLS: &str = r#"[{"type":"function","function":{"name":"agent_browser","description":"Execute an agent-browser command. Runs against the active session by default. Add --session <name> to target or create a different session, and --engine <engine> to choose a browser engine.","parameters":{"type":"object","properties":{"command":{"type":"string","description":"The command to execute, e.g. 'agent-browser open https://google.com' or 'agent-browser --session new-session open https://example.com' or 'agent-browser snapshot -i' or 'agent-browser click @e3'"}},"required":["command"]}}}]"#;
pub(crate) const COMPACT_THRESHOLD_CHARS: usize = 200_000;
pub(crate) const KEEP_RECENT_MESSAGES: usize = 6;
pub(crate) fn estimate_chars(messages: &[Value]) -> usize {
messages
.iter()
.map(|m| {
let content_len = m
.get("content")
.map(|c| {
if let Some(s) = c.as_str() {
s.len()
} else {
c.to_string().len()
}
})
.unwrap_or(0);
let tc_len = m
.get("tool_calls")
.map(|t| t.to_string().len())
.unwrap_or(0);
content_len + tc_len
})
.sum()
}
pub(crate) fn find_safe_split(messages: &[Value], keep_recent: usize) -> usize {
if messages.len() <= keep_recent + 1 {
return 1;
}
let desired = messages.len() - keep_recent;
for i in (1..=desired).rev() {
if messages[i].get("role").and_then(|r| r.as_str()) == Some("user") {
return i;
}
}
desired.max(1)
}
fn build_summary_text(messages: &[Value]) -> String {
let mut text = String::new();
for msg in messages {
let role = msg
.get("role")
.and_then(|r| r.as_str())
.unwrap_or("unknown");
if let Some(content) = msg.get("content").and_then(|c| c.as_str()) {
if !content.is_empty() {
let truncated = if content.len() > 2000 {
format!("{}...[truncated]", &content[..2000])
} else {
content.to_string()
};
text.push_str(&format!("[{}] {}\n\n", role, truncated));
}
}
if let Some(tcs) = msg.get("tool_calls").and_then(|t| t.as_array()) {
for tc in tcs {
let name = tc
.get("function")
.and_then(|f| f.get("name"))
.and_then(|n| n.as_str())
.unwrap_or("");
let args = tc
.get("function")
.and_then(|f| f.get("arguments"))
.and_then(|a| a.as_str())
.unwrap_or("");
text.push_str(&format!("[assistant tool:{}] {}\n", name, args));
}
}
}
text
}
pub(crate) async fn summarize_for_compaction(
client: &reqwest::Client,
url: &str,
api_key: &str,
model: &str,
messages: &[Value],
) -> Option<String> {
let conversation = build_summary_text(messages);
if conversation.is_empty() {
return None;
}
let body = json!({
"model": model,
"messages": [
{
"role": "system",
"content": "Summarize this browser automation conversation concisely. Preserve: URLs visited, actions performed, current page state, errors encountered, and user goals. Output only the summary."
},
{
"role": "user",
"content": conversation
}
],
"max_tokens": 1024,
"stream": false,
});
let resp = client
.post(url)
.header("Authorization", format!("Bearer {}", api_key))
.header("Content-Type", "application/json")
.body(body.to_string())
.send()
.await
.ok()?;
if !resp.status().is_success() {
return None;
}
let result: Value = resp.json().await.ok()?;
result
.get("choices")
.and_then(|c| c.get(0))
.and_then(|c| c.get("message"))
.and_then(|m| m.get("content"))
.and_then(|c| c.as_str())
.map(|s| s.to_string())
}
const SCREENSHOT_MAX_WIDTH: u32 = 1024;
const SCREENSHOT_JPEG_QUALITY: u8 = 40;
fn compress_image_to_jpeg(raw_bytes: &[u8]) -> Option<Vec<u8>> {
let img = image::load_from_memory(raw_bytes).ok()?;
let img = if img.width() > SCREENSHOT_MAX_WIDTH {
img.resize(
SCREENSHOT_MAX_WIDTH,
u32::MAX,
image::imageops::FilterType::Triangle,
)
} else {
img
};
let mut buf = std::io::Cursor::new(Vec::new());
let encoder =
image::codecs::jpeg::JpegEncoder::new_with_quality(&mut buf, SCREENSHOT_JPEG_QUALITY);
img.write_with_encoder(encoder).ok()?;
Some(buf.into_inner())
}
fn has_image_extension(s: &str) -> bool {
let lower = s.to_lowercase();
lower.ends_with(".png") || lower.ends_with(".jpg") || lower.ends_with(".jpeg")
}
fn extract_image_path(text: &str) -> Option<String> {
for line in text.lines() {
let trimmed = line.trim();
// Whole line is a path (handles paths with spaces)
if has_image_extension(trimmed) && std::path::Path::new(trimmed).exists() {
return Some(trimmed.to_string());
}
for suffix in [".png", ".jpg", ".jpeg"] {
if let Some(pos) = trimmed.to_lowercase().rfind(suffix) {
let end = pos + suffix.len();
let candidate = &trimmed[..end];
let start = candidate
.rfind(|c: char| c.is_whitespace())
.map(|i| i + 1)
.unwrap_or(0);
let path = &candidate[start..];
if !path.is_empty() && std::path::Path::new(path).exists() {
return Some(path.to_string());
}
}
}
}
None
}
fn enrich_tool_output(result: &str) -> String {
let Some(path) = extract_image_path(result) else {
return result.to_string();
};
let Ok(raw_bytes) = std::fs::read(&path) else {
return result.to_string();
};
let (jpeg_bytes, mime) = match compress_image_to_jpeg(&raw_bytes) {
Some(compressed) => (compressed, "image/jpeg"),
None => {
let lower = path.to_lowercase();
(
raw_bytes,
if lower.ends_with(".png") {
"image/png"
} else {
"image/jpeg"
},
)
}
};
let b64 = base64::Engine::encode(&base64::engine::general_purpose::STANDARD, &jpeg_bytes);
let data_url = format!("data:{};base64,{}", mime, b64);
json!({
"text": result,
"image": data_url
})
.to_string()
}
const ALLOWED_COMMANDS: &[&str] = &[
"open",
"goto",
"navigate",
"back",
"forward",
"reload",
"click",
"dblclick",
"fill",
"type",
"hover",
"focus",
"check",
"uncheck",
"select",
"drag",
"upload",
"download",
"press",
"key",
"keydown",
"keyup",
"keyboard",
"scroll",
"scrollintoview",
"scrollinto",
"wait",
"screenshot",
"pdf",
"snapshot",
"eval",
"close",
"quit",
"exit",
"inspect",
"auth",
"confirm",
"deny",
"connect",
"cookies",
"storage",
"window",
"frame",
"dialog",
"trace",
"profiler",
"record",
"har",
"network",
"title",
"url",
"console",
"errors",
"highlight",
"state",
"emulate",
"video",
"tap",
"swipe",
"device",
"batch",
"diff",
"find",
"role",
"text",
"label",
"placeholder",
"alt",
"testid",
"first",
"last",
"nth",
"mouse",
"touchscreen",
"attribute",
"property",
"set",
"get",
"is",
"stream",
"tab",
"clipboard",
"session",
];
const ALLOWED_GLOBAL_FLAGS: &[&str] = &["--session", "--engine"];
pub(crate) async fn execute_chat_tool(session: &str, command: &str) -> String {
let exe = match std::env::current_exe() {
Ok(p) => p,
Err(e) => return format!("Failed to resolve executable: {}", e),
};
let single = command.split("&&").next().unwrap_or(command);
let single = single.split(';').next().unwrap_or(single).trim();
let stripped = single.strip_prefix("agent-browser ").unwrap_or(single);
let words = crate::commands::shell_words_split(stripped);
let mut global_flags: Vec<String> = Vec::new();
let mut cmd_words: Vec<String> = Vec::new();
let mut has_session_flag = false;
let mut i = 0;
while i < words.len() {
if ALLOWED_GLOBAL_FLAGS.contains(&words[i].as_str()) {
if words[i] == "--session" {
has_session_flag = true;
}
global_flags.push(words[i].clone());
if i + 1 < words.len() {
global_flags.push(words[i + 1].clone());
i += 2;
} else {
i += 1;
}
} else {
cmd_words.push(words[i].clone());
i += 1;
}
}
let first_cmd = cmd_words.first().map(|s| s.as_str()).unwrap_or("");
if !ALLOWED_COMMANDS.contains(&first_cmd) {
return format!(
"Blocked: '{}' is not a valid agent-browser command.",
first_cmd
);
}
let mut args: Vec<String> = Vec::new();
if !has_session_flag {
args.push("--session".into());
args.push(session.into());
}
args.extend(global_flags);
args.extend(cmd_words);
let mut cmd = tokio::process::Command::new(&exe);
cmd.args(&args)
.env_remove("AGENT_BROWSER_DASHBOARD")
.env_remove("AGENT_BROWSER_DASHBOARD_PORT")
.env_remove("AGENT_BROWSER_STREAM_PORT");
match cmd.output().await {
Ok(output) => {
let stdout = String::from_utf8_lossy(&output.stdout).trim().to_string();
let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string();
if stdout.is_empty() && !stderr.is_empty() {
stderr
} else if stdout.is_empty() {
"Command completed with no output.".to_string()
} else {
stdout
}
}
Err(e) => format!("Failed to execute command: {}", e),
}
}
async fn stream_gateway_response(
stream: &mut tokio::net::TcpStream,
gw_response: reqwest::Response,
) -> Vec<(String, String, String)> {
use futures_util::StreamExt as _;
let mut text_part_id = uuid::Uuid::new_v4().to_string();
let mut text_started = false;
let mut tool_calls: Vec<(String, String, String)> = Vec::new();
let mut tool_call_args: std::collections::HashMap<usize, (String, String, String)> =
std::collections::HashMap::new();
let mut byte_stream = gw_response.bytes_stream();
let mut buffer = String::new();
while let Some(chunk_result) = byte_stream.next().await {
let chunk = match chunk_result {
Ok(c) => c,
Err(_) => break,
};
buffer.push_str(&String::from_utf8_lossy(&chunk));
while let Some(newline_pos) = buffer.find('\n') {
let line = buffer[..newline_pos].trim_end_matches('\r').to_string();
buffer = buffer[newline_pos + 1..].to_string();
if line.is_empty() {
continue;
}
let Some(data) = line.strip_prefix("data: ") else {
continue;
};
if data == "[DONE]" {
if text_started {
let ev = format!("data: {}\n\n", json!({"type":"text-end","id":text_part_id}));
let _ = stream.write_all(ev.as_bytes()).await;
}
let mut indices: Vec<usize> = tool_call_args.keys().copied().collect();
indices.sort();
for idx in indices {
if let Some(tc) = tool_call_args.remove(&idx) {
tool_calls.push(tc);
}
}
return tool_calls;
}
let Ok(sse_json) = serde_json::from_str::<Value>(data) else {
continue;
};
let delta = sse_json
.get("choices")
.and_then(|c| c.get(0))
.and_then(|c| c.get("delta"));
let Some(delta) = delta else { continue };
if let Some(text) = delta.get("content").and_then(|c| c.as_str()) {
if !text.is_empty() {
if !text_started {
let ev = format!(
"data: {}\n\n",
json!({"type":"text-start","id":text_part_id})
);
if stream.write_all(ev.as_bytes()).await.is_err() {
return tool_calls;
}
text_started = true;
}
let ev = format!(
"data: {}\n\n",
json!({"type":"text-delta","id":text_part_id,"delta":text})
);
if stream.write_all(ev.as_bytes()).await.is_err() {
return tool_calls;
}
}
}
if let Some(tcs) = delta.get("tool_calls").and_then(|t| t.as_array()) {
if text_started {
let ev = format!("data: {}\n\n", json!({"type":"text-end","id":text_part_id}));
let _ = stream.write_all(ev.as_bytes()).await;
text_started = false;
text_part_id = uuid::Uuid::new_v4().to_string();
}
for tc in tcs {
let idx = tc.get("index").and_then(|i| i.as_u64()).unwrap_or(0) as usize;
if let std::collections::hash_map::Entry::Vacant(e) = tool_call_args.entry(idx)
{
let id = tc
.get("id")
.and_then(|i| i.as_str())
.unwrap_or("")
.to_string();
let name = tc
.get("function")
.and_then(|f| f.get("name"))
.and_then(|n| n.as_str())
.unwrap_or("")
.to_string();
let ev = format!(
"data: {}\n\n",
json!({"type":"tool-input-start","toolCallId":id,"toolName":name})
);
let _ = stream.write_all(ev.as_bytes()).await;
e.insert((id, name, String::new()));
}
if let Some(arg_delta) = tc
.get("function")
.and_then(|f| f.get("arguments"))
.and_then(|a| a.as_str())
{
let entry = tool_call_args.get_mut(&idx).unwrap();
entry.2.push_str(arg_delta);
let ev = format!(
"data: {}\n\n",
json!({"type":"tool-input-delta","toolCallId":entry.0,"inputTextDelta":arg_delta})
);
let _ = stream.write_all(ev.as_bytes()).await;
}
}
}
}
}
if text_started {
let ev = format!("data: {}\n\n", json!({"type":"text-end","id":text_part_id}));
let _ = stream.write_all(ev.as_bytes()).await;
}
let mut indices: Vec<usize> = tool_call_args.keys().copied().collect();
indices.sort();
for idx in indices {
if let Some(tc) = tool_call_args.remove(&idx) {
tool_calls.push(tc);
}
}
tool_calls
}
pub(super) async fn handle_chat_request(
stream: &mut tokio::net::TcpStream,
body: &str,
origin: Option<&str>,
) {
let cors = cors_headers_for_origin(origin);
let gateway_url = std::env::var("AI_GATEWAY_URL")
.unwrap_or_else(|_| DEFAULT_AI_GATEWAY_URL.to_string())
.trim_end_matches('/')
.to_string();
let api_key = match std::env::var("AI_GATEWAY_API_KEY") {
Ok(k) => k,
Err(_) => {
let err = r#"{"error":"AI_GATEWAY_API_KEY not set. Set the AI_GATEWAY_API_KEY environment variable to enable AI chat."}"#;
let resp = format!(
"HTTP/1.1 500 Internal Server Error\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n{cors}\r\n",
err.len()
);
let _ = stream.write_all(resp.as_bytes()).await;
let _ = stream.write_all(err.as_bytes()).await;
return;
}
};
let default_model = std::env::var("AI_GATEWAY_MODEL")
.unwrap_or_else(|_| "anthropic/claude-sonnet-4.6".to_string());
let parsed: Value = match serde_json::from_str(body) {
Ok(v) => v,
Err(e) => {
let err = format!(r#"{{"error":"Invalid JSON: {}"}}"#, e);
let resp = format!(
"HTTP/1.1 400 Bad Request\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n{cors}\r\n",
err.len()
);
let _ = stream.write_all(resp.as_bytes()).await;
let _ = stream.write_all(err.as_bytes()).await;
return;
}
};
let messages = parsed.get("messages").cloned().unwrap_or(json!([]));
let model = parsed
.get("model")
.and_then(|v| v.as_str())
.unwrap_or(&default_model)
.to_string();
let session = parsed
.get("session")
.and_then(|v| v.as_str())
.unwrap_or("default")
.to_string();
let mut openai_messages: Vec<Value> =
vec![json!({"role": "system", "content": get_system_prompt()})];
let mut frontend_boundaries: Vec<usize> = Vec::new();
let frontend_arr = messages.as_array();
let frontend_count = frontend_arr.map(|a| a.len()).unwrap_or(0);
if let Some(arr) = frontend_arr {
for msg in arr {
frontend_boundaries.push(openai_messages.len());
let Some(role) = msg.get("role").and_then(|r| r.as_str()) else {
continue;
};
if let Some(parts) = msg.get("parts").and_then(|p| p.as_array()) {
let mut content_parts: Vec<Value> = Vec::new();
for part in parts {
match part.get("type").and_then(|t| t.as_str()) {
Some("text") => {
if let Some(text) = part.get("text").and_then(|t| t.as_str()) {
if !text.is_empty() {
content_parts.push(json!({"type": "text", "text": text}));
}
}
}
Some("file") => {
if let (Some(url), Some(media_type)) = (
part.get("url").and_then(|u| u.as_str()),
part.get("mediaType").and_then(|m| m.as_str()),
) {
if media_type.starts_with("image/") {
content_parts.push(json!({
"type": "image_url",
"image_url": { "url": url }
}));
}
}
}
_ => {}
}
}
if !content_parts.is_empty() {
let content = if content_parts.len() == 1
&& content_parts[0].get("type").and_then(|t| t.as_str()) == Some("text")
{
content_parts[0]["text"].clone()
} else {
json!(content_parts)
};
openai_messages.push(json!({"role": role, "content": content}));
}
} else if let Some(content) = msg.get("content").and_then(|c| c.as_str()) {
openai_messages.push(json!({"role": role, "content": content}));
}
}
}
let tools: Value = serde_json::from_str(CHAT_TOOLS).unwrap();
let url = format!("{}/v1/chat/completions", gateway_url);
let client = http_client();
let total_chars = estimate_chars(&openai_messages);
let mut compaction_summary: Option<String> = None;
let mut compaction_failed = false;
let mut keep_last_n: usize = frontend_count;
if total_chars > COMPACT_THRESHOLD_CHARS && openai_messages.len() > KEEP_RECENT_MESSAGES + 2 {
let split = find_safe_split(&openai_messages, KEEP_RECENT_MESSAGES);
let to_summarize = &openai_messages[1..split];
if let Some(summary) =
summarize_for_compaction(client, &url, &api_key, &model, to_summarize).await
{
let summary_msg = json!({
"role": "system",
"content": format!("[Conversation summary]\n{}", summary)
});
let recent = openai_messages[split..].to_vec();
openai_messages = vec![openai_messages[0].clone(), summary_msg];
openai_messages.extend(recent);
let kept_frontend = frontend_boundaries
.iter()
.filter(|&&boundary| boundary >= split)
.count();
keep_last_n = kept_frontend;
compaction_summary = Some(summary);
} else {
compaction_failed = true;
}
}
let headers = format!(
"HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\nConnection: keep-alive\r\nx-vercel-ai-ui-message-stream: v1\r\n{cors}\r\n"
);
if stream.write_all(headers.as_bytes()).await.is_err() {
return;
}
let message_id = uuid::Uuid::new_v4().to_string();
let start_ev = format!(
"data: {}\n\n",
json!({"type":"start","messageId":message_id})
);
if stream.write_all(start_ev.as_bytes()).await.is_err() {
return;
}
if let Some(ref summary) = compaction_summary {
let ev = format!(
"data: {}\n\n",
json!({
"type": "message-metadata",
"messageMetadata": {
"compacted": true,
"summary": summary,
"keepLastN": keep_last_n
}
})
);
let _ = stream.write_all(ev.as_bytes()).await;
} else if compaction_failed {
let ev = format!(
"data: {}\n\n",
json!({
"type": "message-metadata",
"messageMetadata": {
"compacted": false,
"warning": "Conversation is large but compaction failed. Responses may be degraded."
}
})
);
let _ = stream.write_all(ev.as_bytes()).await;
}
let total_deadline = tokio::time::Instant::now() + std::time::Duration::from_secs(300);
const TOOL_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(60);
for _step in 0..50 {
if tokio::time::Instant::now() >= total_deadline {
let ev = format!(
"data: {}\n\n",
json!({"type":"error","errorText":"Chat session timed out (5 minute limit)."})
);
let _ = stream.write_all(ev.as_bytes()).await;
break;
}
let step_ev = "data: {\"type\":\"start-step\"}\n\n";
if stream.write_all(step_ev.as_bytes()).await.is_err() {
return;
}
let gateway_body = json!({
"model": model,
"messages": openai_messages,
"tools": tools,
"stream": true,
});
let gw_response = match client
.post(&url)
.header("Authorization", format!("Bearer {}", api_key))
.header("Content-Type", "application/json")
.body(gateway_body.to_string())
.send()
.await
{
Ok(r) => r,
Err(e) => {
let ev = format!(
"data: {}\n\n",
json!({"type":"error","errorText":format!("Gateway request failed: {}", e)})
);
let _ = stream.write_all(ev.as_bytes()).await;
break;
}
};
if !gw_response.status().is_success() {
let body_text = gw_response.text().await.unwrap_or_default();
let ev = format!(
"data: {}\n\n",
json!({"type":"error","errorText":body_text})
);
let _ = stream.write_all(ev.as_bytes()).await;
break;
}
let tool_calls = stream_gateway_response(stream, gw_response).await;
if tool_calls.is_empty() {
let finish_step_ev = "data: {\"type\":\"finish-step\"}\n\n";
let _ = stream.write_all(finish_step_ev.as_bytes()).await;
break;
}
let tc_values: Vec<Value> = tool_calls.iter().map(|(id, name, args)| {
json!({"id": id, "type": "function", "function": {"name": name, "arguments": args}})
}).collect();
openai_messages.push(json!({"role": "assistant", "tool_calls": tc_values}));
for (tc_id, tc_name, tc_args) in &tool_calls {
let input: Value = serde_json::from_str(tc_args).unwrap_or(json!({}));
let command = input.get("command").and_then(|c| c.as_str()).unwrap_or("");
let ev = format!(
"data: {}\n\n",
json!({
"type": "tool-input-available",
"toolCallId": tc_id,
"toolName": tc_name,
"input": input
})
);
let _ = stream.write_all(ev.as_bytes()).await;
let result = match tokio::time::timeout(
TOOL_TIMEOUT,
execute_chat_tool(&session, command),
)
.await
{
Ok(r) => r,
Err(_) => "Tool execution timed out after 60 seconds.".to_string(),
};
let frontend_output = enrich_tool_output(&result);
let ev = format!(
"data: {}\n\n",
json!({
"type": "tool-output-available",
"toolCallId": tc_id,
"output": frontend_output
})
);
let _ = stream.write_all(ev.as_bytes()).await;
openai_messages.push(json!({
"role": "tool",
"tool_call_id": tc_id,
"content": result
}));
}
let finish_step_ev = "data: {\"type\":\"finish-step\"}\n\n";
let _ = stream.write_all(finish_step_ev.as_bytes()).await;
}
let finish_ev = "data: {\"type\":\"finish\"}\n\n";
let _ = stream.write_all(finish_ev.as_bytes()).await;
let done_ev = "data: [DONE]\n\n";
let _ = stream.write_all(done_ev.as_bytes()).await;
}
+960
View File
@@ -0,0 +1,960 @@
use futures_util::{SinkExt, StreamExt};
use serde_json::{json, Value};
use tokio::io::{AsyncReadExt, AsyncWriteExt};
use tokio::net::TcpListener;
use tokio_tungstenite::tungstenite::Message;
use crate::connection::get_socket_dir;
use super::chat::{chat_status_json, handle_chat_request, handle_models_request};
use super::discovery::discover_sessions;
use super::http::{serve_embedded_file, CORS_HEADERS};
/// Dashboard same-origin proxy endpoints for session metadata and streams.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum SessionProxyEndpoint {
Tabs,
Status,
Stream,
}
#[derive(Debug, Clone, PartialEq, Eq)]
struct DashboardProxyError {
status: &'static str,
message: String,
}
impl DashboardProxyError {
fn not_found(message: impl Into<String>) -> Self {
Self {
status: "404 Not Found",
message: message.into(),
}
}
fn bad_gateway(message: impl Into<String>) -> Self {
Self {
status: "502 Bad Gateway",
message: message.into(),
}
}
}
const PROXY_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(30);
const PROXY_MAX_RESPONSE_SIZE: u64 = 16 * 1024 * 1024;
fn build_json_error_body(error: &str) -> String {
let escaped = serde_json::to_string(error).unwrap_or_else(|_| format!("\"{}\"", error));
format!(r#"{{"success":false,"error":{escaped}}}"#)
}
async fn write_http_response_inner(
stream: &mut tokio::net::TcpStream,
status: &str,
content_type: &str,
body: &[u8],
include_cors: bool,
) {
let cors_headers = if include_cors { CORS_HEADERS } else { "" };
let response = format!(
"HTTP/1.1 {status}\r\nContent-Type: {content_type}\r\nContent-Length: {}\r\nConnection: close\r\n{cors_headers}\r\n",
body.len()
);
let _ = stream.write_all(response.as_bytes()).await;
let _ = stream.write_all(body).await;
}
async fn write_http_response(
stream: &mut tokio::net::TcpStream,
status: &str,
content_type: &str,
body: &[u8],
) {
write_http_response_inner(stream, status, content_type, body, true).await;
}
async fn write_http_response_no_cors(
stream: &mut tokio::net::TcpStream,
status: &str,
content_type: &str,
body: &[u8],
) {
write_http_response_inner(stream, status, content_type, body, false).await;
}
async fn write_json_error_response_no_cors(
stream: &mut tokio::net::TcpStream,
status: &'static str,
error: &str,
) {
let body = build_json_error_body(error);
write_http_response_no_cors(
stream,
status,
"application/json; charset=utf-8",
body.as_bytes(),
)
.await;
}
fn parse_request_method_and_path(request: &str) -> (&str, &str) {
let first_line = request.lines().next().unwrap_or("");
let method = first_line.split_whitespace().next().unwrap_or("GET");
let path = first_line.split_whitespace().nth(1).unwrap_or("/");
(method, path)
}
fn is_websocket_upgrade(request: &str) -> bool {
request.lines().any(|line| {
if let Some((name, value)) = line.split_once(':') {
name.trim().eq_ignore_ascii_case("upgrade")
&& value.trim().eq_ignore_ascii_case("websocket")
} else {
false
}
})
}
fn request_header_value<'a>(request: &'a str, name: &str) -> Option<&'a str> {
request.lines().find_map(|line| {
let (header_name, value) = line.split_once(':')?;
if header_name.trim().eq_ignore_ascii_case(name) {
Some(value.trim())
} else {
None
}
})
}
fn normalize_origin_authority(origin: &str) -> Option<String> {
let url = url::Url::parse(origin).ok()?;
let host = url.host_str()?.to_ascii_lowercase();
let host = if host.contains(':') {
format!("[{host}]")
} else {
host
};
Some(match url.port() {
Some(port) => format!("{host}:{port}"),
None => host,
})
}
fn normalize_host_authority(host: &str) -> String {
let host = host.trim().to_ascii_lowercase();
if let Some(bracket_end) = host.rfind(']') {
if bracket_end == host.len() - 1 {
return host;
}
if host.as_bytes().get(bracket_end + 1) == Some(&b':') {
let port = &host[bracket_end + 2..];
if port == "80" || port == "443" {
return host[..=bracket_end].to_string();
}
}
return host;
}
if let Some((name, port)) = host.rsplit_once(':') {
if !name.contains(':') && (port == "80" || port == "443") {
return name.to_string();
}
}
host
}
fn header_matches_host(request: &str, header_name: &str) -> Option<bool> {
let authority =
request_header_value(request, header_name).and_then(normalize_origin_authority)?;
let host = request_header_value(request, "host").map(normalize_host_authority)?;
Some(authority == host)
}
/// Validates that a proxied WebSocket request either has no Origin header or
/// presents an Origin whose authority matches the request Host header.
fn is_same_origin_ws_request(request: &str) -> bool {
match header_matches_host(request, "origin") {
Some(matches) => matches,
None => request_header_value(request, "origin").is_none(),
}
}
/// Validates that an HTTP session-proxy request came from a same-origin page.
///
/// For GET requests we require either a same-origin `Origin` or a same-origin
/// `Referer` so browsers cannot hit the proxy routes via side-channel tags or
/// arbitrary cross-origin fetches.
fn is_same_origin_http_request(request: &str) -> bool {
matches!(header_matches_host(request, "origin"), Some(true))
|| matches!(header_matches_host(request, "referer"), Some(true))
}
/// Parse a dashboard route of the form `/api/session/<port>/<endpoint>`.
fn parse_session_proxy_route(path: &str) -> Result<(u16, SessionProxyEndpoint), &'static str> {
if !path.starts_with("/api/session/") {
return Err("Invalid session proxy route.");
}
let mut parts = path.split('/');
if parts.next() != Some("") || parts.next() != Some("api") || parts.next() != Some("session") {
return Err("Invalid session proxy route.");
}
let port_str = parts.next().ok_or("Missing session proxy port.")?;
if port_str.is_empty() {
return Err("Missing session proxy port.");
}
let endpoint = match parts.next().ok_or("Missing session proxy endpoint.")? {
"tabs" => SessionProxyEndpoint::Tabs,
"status" => SessionProxyEndpoint::Status,
"stream" => SessionProxyEndpoint::Stream,
_ => return Err("Unknown session proxy endpoint."),
};
if parts.next().is_some() {
return Err("Unexpected path segments in session proxy route.");
}
let port = port_str
.parse::<u16>()
.map_err(|_| "Session proxy port must be a valid TCP port.")?;
if port == 0 {
return Err("Session proxy port must be a valid TCP port.");
}
Ok((port, endpoint))
}
fn sessions_json_has_active_port(sessions_json: &str, port: u16) -> Result<bool, String> {
let sessions: Vec<Value> = serde_json::from_str(sessions_json)
.map_err(|e| format!("Failed to parse active sessions: {e}"))?;
Ok(sessions.iter().any(|session| {
session
.get("port")
.and_then(|value| value.as_u64())
.map(|value| value == u64::from(port))
.unwrap_or(false)
}))
}
fn require_active_session_port(port: u16) -> Result<(), DashboardProxyError> {
let sessions_json = discover_sessions();
let is_active = sessions_json_has_active_port(&sessions_json, port)
.map_err(DashboardProxyError::bad_gateway)?;
if is_active {
Ok(())
} else {
Err(DashboardProxyError::not_found(format!(
"No active session is listening on port {port}."
)))
}
}
fn split_http_response(response: &[u8]) -> Result<(&[u8], &[u8]), String> {
if let Some(header_end) = response.windows(4).position(|window| window == b"\r\n\r\n") {
let body_start = header_end + 4;
return Ok((&response[..header_end], &response[body_start..]));
}
if let Some(header_end) = response.windows(2).position(|window| window == b"\n\n") {
let body_start = header_end + 2;
return Ok((&response[..header_end], &response[body_start..]));
}
Err("Upstream response was missing an HTTP header terminator.".to_string())
}
fn parse_upstream_http_response(response: &[u8]) -> Result<(String, String, Vec<u8>), String> {
let (header_bytes, body) = split_http_response(response)?;
let header_str = std::str::from_utf8(header_bytes)
.map_err(|e| format!("Upstream response headers were not valid UTF-8: {e}"))?;
let mut lines = header_str.lines();
let status_line = lines
.next()
.ok_or_else(|| "Upstream response was missing a status line.".to_string())?;
let status = status_line
.split_once(' ')
.map(|(_, status)| status.trim().to_string())
.filter(|status| !status.is_empty())
.ok_or_else(|| "Upstream response status line was malformed.".to_string())?;
let content_type = lines
.find_map(|line| {
let (name, value) = line.split_once(':')?;
if name.trim().eq_ignore_ascii_case("content-type") {
Some(value.trim().to_string())
} else {
None
}
})
.unwrap_or_else(|| "application/json; charset=utf-8".to_string());
Ok((status, content_type, body.to_vec()))
}
/// Proxy dashboard-origin HTTP requests for session tabs or status to the loopback session server.
async fn proxy_session_http_route(
port: u16,
endpoint: SessionProxyEndpoint,
) -> Result<(String, String, Vec<u8>), DashboardProxyError> {
debug_assert!(matches!(
endpoint,
SessionProxyEndpoint::Tabs | SessionProxyEndpoint::Status
));
require_active_session_port(port)?;
let upstream_path = match endpoint {
SessionProxyEndpoint::Tabs => "/api/tabs",
SessionProxyEndpoint::Status => "/api/status",
SessionProxyEndpoint::Stream => unreachable!("stream routes use the WebSocket proxy"),
};
let request = format!(
"GET {upstream_path} HTTP/1.1\r\nHost: 127.0.0.1:{port}\r\nConnection: close\r\n\r\n"
);
tokio::time::timeout(PROXY_TIMEOUT, async {
let mut upstream = tokio::net::TcpStream::connect(("127.0.0.1", port))
.await
.map_err(|e| {
DashboardProxyError::bad_gateway(format!(
"Failed to connect to session {port}: {e}"
))
})?;
upstream.write_all(request.as_bytes()).await.map_err(|e| {
DashboardProxyError::bad_gateway(format!(
"Failed to proxy request to session {port}: {e}"
))
})?;
let mut response = Vec::new();
(&mut upstream)
.take(PROXY_MAX_RESPONSE_SIZE + 1)
.read_to_end(&mut response)
.await
.map_err(|e| {
DashboardProxyError::bad_gateway(format!(
"Failed to read session {port} response: {e}"
))
})?;
if response.len() as u64 > PROXY_MAX_RESPONSE_SIZE {
return Err(DashboardProxyError::bad_gateway(format!(
"Session {port} response exceeded {PROXY_MAX_RESPONSE_SIZE} bytes."
)));
}
parse_upstream_http_response(&response).map_err(DashboardProxyError::bad_gateway)
})
.await
.map_err(|_| {
DashboardProxyError::bad_gateway(format!(
"Session {port} proxy request timed out after {}s.",
PROXY_TIMEOUT.as_secs()
))
})?
}
/// Bridge a dashboard-origin WebSocket upgrade to the loopback session stream.
async fn proxy_session_stream(mut stream: tokio::net::TcpStream, port: u16) {
let upstream_url = format!("ws://127.0.0.1:{port}");
let (upstream_ws, _) = match tokio_tungstenite::connect_async(&upstream_url).await {
Ok(ws) => ws,
Err(error) => {
write_json_error_response_no_cors(
&mut stream,
"502 Bad Gateway",
&format!("Failed to connect to session {port}: {error}"),
)
.await;
return;
}
};
let client_ws = match tokio_tungstenite::accept_async(stream).await {
Ok(ws) => ws,
Err(_) => return,
};
let (mut client_tx, mut client_rx) = client_ws.split();
let (mut upstream_tx, mut upstream_rx) = upstream_ws.split();
loop {
tokio::select! {
message = client_rx.next() => {
match message {
Some(Ok(message)) => {
let is_close = matches!(message, Message::Close(_));
if upstream_tx.send(message).await.is_err() {
break;
}
if is_close {
break;
}
}
Some(Err(_)) | None => {
let _ = upstream_tx.send(Message::Close(None)).await;
break;
}
}
}
message = upstream_rx.next() => {
match message {
Some(Ok(message)) => {
let is_close = matches!(message, Message::Close(_));
if client_tx.send(message).await.is_err() {
break;
}
if is_close {
break;
}
}
Some(Err(_)) | None => {
let _ = client_tx.send(Message::Close(None)).await;
break;
}
}
}
}
}
}
pub async fn run_dashboard_server(port: u16) {
let addr = format!("127.0.0.1:{}", port);
let listener = match TcpListener::bind(&addr).await {
Ok(l) => l,
Err(e) => {
eprintln!("Failed to bind dashboard server on {}: {}", addr, e);
return;
}
};
loop {
let Ok((stream, _addr)) = listener.accept().await else {
break;
};
tokio::spawn(async move {
handle_dashboard_connection(stream).await;
});
}
}
async fn handle_dashboard_connection(mut stream: tokio::net::TcpStream) {
let mut buf = vec![0u8; 8192];
let peeked_len = match stream.peek(&mut buf).await {
Ok(n) if n > 0 => n,
_ => return,
};
let peeked_request = String::from_utf8_lossy(&buf[..peeked_len]);
let (peeked_method, peeked_path) = parse_request_method_and_path(&peeked_request);
if peeked_path.starts_with("/api/session/") {
let (port, endpoint) = match parse_session_proxy_route(peeked_path) {
Ok(route) => route,
Err(error) => {
write_json_error_response_no_cors(&mut stream, "400 Bad Request", error).await;
return;
}
};
match endpoint {
SessionProxyEndpoint::Stream => {
if peeked_method != "GET" {
write_json_error_response_no_cors(
&mut stream,
"400 Bad Request",
"Session stream proxy only supports GET WebSocket upgrades.",
)
.await;
return;
}
if !is_websocket_upgrade(&peeked_request) {
write_json_error_response_no_cors(
&mut stream,
"400 Bad Request",
"Session stream proxy requires a WebSocket upgrade request.",
)
.await;
return;
}
if !is_same_origin_ws_request(&peeked_request) {
write_json_error_response_no_cors(
&mut stream,
"403 Forbidden",
"Origin does not match Host header.",
)
.await;
return;
}
if let Err(error) = require_active_session_port(port) {
write_json_error_response_no_cors(&mut stream, error.status, &error.message)
.await;
return;
}
proxy_session_stream(stream, port).await;
return;
}
SessionProxyEndpoint::Tabs | SessionProxyEndpoint::Status => {
if peeked_method != "GET" {
write_json_error_response_no_cors(
&mut stream,
"400 Bad Request",
"Session proxy routes only support GET requests.",
)
.await;
return;
}
}
}
}
let n = match stream.read(&mut buf).await {
Ok(n) if n > 0 => n,
_ => return,
};
let request = String::from_utf8_lossy(&buf[..n]).to_string();
let (method, path) = parse_request_method_and_path(&request);
let origin = request_header_value(&request, "origin").map(|value| value.to_string());
if method == "OPTIONS" {
let response = format!(
"HTTP/1.1 204 No Content\r\n{CORS_HEADERS}Access-Control-Max-Age: 86400\r\nContent-Length: 0\r\nConnection: close\r\n\r\n"
);
let _ = stream.write_all(response.as_bytes()).await;
return;
}
if method == "POST" && path == "/api/chat" {
let body_str = read_post_body(&mut stream, &buf, n).await;
handle_chat_request(&mut stream, &body_str, origin.as_deref()).await;
return;
}
if method == "GET" && path == "/api/models" {
handle_models_request(&mut stream, origin.as_deref()).await;
return;
}
if method == "POST" && (path == "/api/sessions" || path == "/api/exec" || path == "/api/kill") {
let body_str = read_post_body(&mut stream, &buf, n).await;
let result = if path == "/api/exec" {
exec_cli(&body_str).await
} else if path == "/api/kill" {
kill_session(&body_str).await
} else {
spawn_session(&body_str).await
};
let (status, resp_body) = match result {
Ok(msg) => ("200 OK", msg),
Err(e) => ("400 Bad Request", build_json_error_body(&e)),
};
write_http_response(
&mut stream,
status,
"application/json; charset=utf-8",
resp_body.as_bytes(),
)
.await;
return;
}
if path.starts_with("/api/session/") {
let (port, endpoint) = match parse_session_proxy_route(path) {
Ok(route) => route,
Err(error) => {
write_json_error_response_no_cors(&mut stream, "400 Bad Request", error).await;
return;
}
};
match endpoint {
SessionProxyEndpoint::Tabs | SessionProxyEndpoint::Status => {
if !is_same_origin_http_request(&request) {
write_json_error_response_no_cors(
&mut stream,
"403 Forbidden",
"Origin or Referer does not match Host header.",
)
.await;
return;
}
match proxy_session_http_route(port, endpoint).await {
Ok((status, content_type, body)) => {
write_http_response_no_cors(&mut stream, &status, &content_type, &body)
.await;
}
Err(error) => {
write_json_error_response_no_cors(
&mut stream,
error.status,
&error.message,
)
.await;
}
}
return;
}
SessionProxyEndpoint::Stream => {
write_json_error_response_no_cors(
&mut stream,
"400 Bad Request",
"Session stream proxy requires a WebSocket upgrade request.",
)
.await;
return;
}
}
}
let (status, content_type, body): (&str, &str, Vec<u8>) = if path == "/api/sessions" {
(
"200 OK",
"application/json; charset=utf-8",
discover_sessions().into_bytes(),
)
} else if path == "/api/chat/status" {
(
"200 OK",
"application/json; charset=utf-8",
chat_status_json().into_bytes(),
)
} else {
serve_embedded_file(path)
};
write_http_response(&mut stream, status, content_type, &body).await;
}
async fn read_post_body(stream: &mut tokio::net::TcpStream, initial: &[u8], n: usize) -> String {
let header_end = initial[..n]
.windows(4)
.position(|w| w == b"\r\n\r\n")
.map(|p| p + 4)
.or_else(|| {
initial[..n]
.windows(2)
.position(|w| w == b"\n\n")
.map(|p| p + 2)
});
let Some(header_end) = header_end else {
return String::new();
};
let header_str = String::from_utf8_lossy(&initial[..header_end]);
let content_length: usize = header_str
.lines()
.find_map(|l| {
if l.len() > 16 && l[..16].eq_ignore_ascii_case("content-length: ") {
l[16..].trim().parse::<usize>().ok()
} else {
let lower = l.to_lowercase();
lower
.strip_prefix("content-length:")
.and_then(|v| v.trim().parse::<usize>().ok())
}
})
.unwrap_or(0);
if content_length == 0 {
return String::new();
}
let read_body = &initial[header_end..n];
let already_read = read_body.len().min(content_length);
let mut body = Vec::with_capacity(content_length);
body.extend_from_slice(&read_body[..already_read]);
let remaining = content_length - already_read;
if remaining > 0 {
let mut rest = vec![0u8; remaining];
if stream.read_exact(&mut rest).await.is_ok() {
body.extend_from_slice(&rest);
}
}
String::from_utf8(body).unwrap_or_default()
}
async fn exec_cli(body: &str) -> Result<String, String> {
let parsed: Value = serde_json::from_str(body).map_err(|e| format!("Invalid JSON: {}", e))?;
let args: Vec<String> = parsed
.get("args")
.and_then(|v| v.as_array())
.ok_or("Missing \"args\" array")?
.iter()
.filter_map(|v| v.as_str().map(|s| s.to_string()))
.collect();
if args.is_empty() {
return Err("Empty args array".to_string());
}
let exe = std::env::current_exe().map_err(|e| format!("Cannot resolve executable: {}", e))?;
let mut cmd = tokio::process::Command::new(&exe);
cmd.args(&args)
.arg("--json")
.env_remove("AGENT_BROWSER_DASHBOARD")
.env_remove("AGENT_BROWSER_DASHBOARD_PORT")
.env_remove("AGENT_BROWSER_STREAM_PORT");
let output = cmd
.output()
.await
.map_err(|e| format!("Failed to execute: {}", e))?;
let stdout = String::from_utf8_lossy(&output.stdout).trim().to_string();
let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string();
Ok(json!({
"success": output.status.success(),
"exit_code": output.status.code(),
"stdout": stdout,
"stderr": stderr,
})
.to_string())
}
async fn kill_session(body: &str) -> Result<String, String> {
let parsed: Value = serde_json::from_str(body).map_err(|e| format!("Invalid JSON: {}", e))?;
let session = parsed
.get("session")
.and_then(|v| v.as_str())
.ok_or("Missing \"session\" field")?;
if session.is_empty() || session.len() > 64 {
return Err("Session name must be 1-64 characters".to_string());
}
let dir = get_socket_dir();
let pid_path = dir.join(format!("{}.pid", session));
let pid_str = std::fs::read_to_string(&pid_path)
.map_err(|_| format!("No PID file for session '{}'", session))?;
let pid: u32 = pid_str
.trim()
.parse()
.map_err(|_| format!("Invalid PID in file: {}", pid_str.trim()))?;
#[cfg(unix)]
{
// SAFETY: The PID came from the daemon-managed pidfile and is only used
// to send standard termination signals to that process.
unsafe {
libc::kill(pid as i32, libc::SIGTERM);
}
tokio::time::sleep(std::time::Duration::from_millis(500)).await;
// SAFETY: A signal value of 0 performs an existence check on the same pid.
if unsafe { libc::kill(pid as i32, 0) } == 0 {
// SAFETY: The process still exists after SIGTERM, so escalate to SIGKILL.
unsafe {
libc::kill(pid as i32, libc::SIGKILL);
}
}
}
for ext in &["pid", "sock", "stream", "engine", "extensions"] {
let _ = std::fs::remove_file(dir.join(format!("{}.{}", session, ext)));
}
Ok(json!({ "success": true, "killed_pid": pid }).to_string())
}
pub(super) async fn spawn_session(body: &str) -> Result<String, String> {
let parsed: Value = serde_json::from_str(body).map_err(|e| format!("Invalid JSON: {}", e))?;
let session = parsed
.get("session")
.and_then(|v| v.as_str())
.ok_or("Missing \"session\" field")?;
if session.is_empty() || session.len() > 64 {
return Err("Session name must be 1-64 characters".to_string());
}
let exe = std::env::current_exe().map_err(|e| format!("Cannot resolve executable: {}", e))?;
let mut cmd = tokio::process::Command::new(&exe);
cmd.arg("open")
.arg("about:blank")
.arg("--session")
.arg(session);
cmd.stdout(std::process::Stdio::null());
cmd.stderr(std::process::Stdio::null());
let status = cmd
.status()
.await
.map_err(|e| format!("Failed to spawn session: {}", e))?;
if status.success() {
Ok(format!(
r#"{{"success":true,"session":{}}}"#,
serde_json::to_string(session).unwrap_or_default()
))
} else {
Err(format!("Session process exited with {}", status))
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_same_origin_ws_request_matching() {
let req = "GET /api/session/9222/stream HTTP/1.1\r\nHost: localhost:4848\r\nOrigin: http://localhost:4848\r\nUpgrade: websocket\r\n\r\n";
assert!(is_same_origin_ws_request(req));
}
#[test]
fn test_same_origin_ws_request_proxied() {
let req = "GET /api/session/9222/stream HTTP/1.1\r\nHost: dashboard.agent-browser.localhost\r\nOrigin: https://dashboard.agent-browser.localhost\r\nUpgrade: websocket\r\n\r\n";
assert!(is_same_origin_ws_request(req));
}
#[test]
fn test_normalize_origin_authority_https_without_port() {
assert_eq!(
normalize_origin_authority("https://dashboard.agent-browser.localhost"),
Some("dashboard.agent-browser.localhost".to_string())
);
}
#[test]
fn test_same_origin_ws_request_default_https_port() {
let req = "GET /api/session/9222/stream HTTP/1.1\r\nHost: dashboard.agent-browser.localhost:443\r\nOrigin: https://dashboard.agent-browser.localhost\r\nUpgrade: websocket\r\n\r\n";
assert!(is_same_origin_ws_request(req));
}
#[test]
fn test_same_origin_http_request_matching_origin() {
let req = "GET /api/session/9222/tabs HTTP/1.1\r\nHost: localhost:4848\r\nOrigin: http://localhost:4848\r\n\r\n";
assert!(is_same_origin_http_request(req));
}
#[test]
fn test_same_origin_http_request_matching_referer() {
let req = "GET /api/session/9222/tabs HTTP/1.1\r\nHost: dashboard.agent-browser.localhost:443\r\nReferer: https://dashboard.agent-browser.localhost/sessions\r\n\r\n";
assert!(is_same_origin_http_request(req));
}
#[test]
fn test_same_origin_http_request_rejects_missing_origin_and_referer() {
let req = "GET /api/session/9222/tabs HTTP/1.1\r\nHost: localhost:4848\r\n\r\n";
assert!(!is_same_origin_http_request(req));
}
#[test]
fn test_same_origin_http_request_rejects_cross_origin_referer() {
let req = "GET /api/session/9222/tabs HTTP/1.1\r\nHost: localhost:4848\r\nReferer: https://evil.com/path\r\n\r\n";
assert!(!is_same_origin_http_request(req));
}
#[test]
fn test_same_origin_ws_request_coder() {
let req = "GET /api/session/9222/stream HTTP/1.1\r\nHost: workspace.coder.com\r\nOrigin: https://workspace.coder.com\r\nUpgrade: websocket\r\n\r\n";
assert!(is_same_origin_ws_request(req));
}
#[test]
fn test_cross_origin_ws_request_rejected() {
let req = "GET /api/session/9222/stream HTTP/1.1\r\nHost: localhost:4848\r\nOrigin: https://evil.com\r\nUpgrade: websocket\r\n\r\n";
assert!(!is_same_origin_ws_request(req));
}
#[test]
fn test_no_origin_header_allowed() {
let req = "GET /api/session/9222/stream HTTP/1.1\r\nHost: localhost:4848\r\nUpgrade: websocket\r\n\r\n";
assert!(is_same_origin_ws_request(req));
}
#[test]
fn test_parse_session_proxy_route_valid() {
assert_eq!(
parse_session_proxy_route("/api/session/9222/tabs"),
Ok((9222, SessionProxyEndpoint::Tabs))
);
assert_eq!(
parse_session_proxy_route("/api/session/1337/status"),
Ok((1337, SessionProxyEndpoint::Status))
);
assert_eq!(
parse_session_proxy_route("/api/session/65535/stream"),
Ok((65535, SessionProxyEndpoint::Stream))
);
}
#[test]
fn test_parse_session_proxy_route_invalid() {
assert!(parse_session_proxy_route("/api/session/0/tabs").is_err());
assert!(parse_session_proxy_route("/api/session/not-a-port/tabs").is_err());
assert!(parse_session_proxy_route("/api/session/70000/tabs").is_err());
assert!(parse_session_proxy_route("/api/session/9222").is_err());
assert!(parse_session_proxy_route("/api/session/9222/unknown").is_err());
assert!(parse_session_proxy_route("/api/session/9222/tabs/extra").is_err());
}
#[test]
fn test_parse_session_proxy_route_path_traversal() {
assert!(parse_session_proxy_route("/api/session/9222/tabs/..").is_err());
assert!(parse_session_proxy_route("/api/session/9222/tabs/../status").is_err());
assert!(parse_session_proxy_route("/api/session/9222/../../etc/passwd").is_err());
assert!(parse_session_proxy_route("/api/session/../session/9222/tabs").is_err());
}
#[test]
fn test_parse_session_proxy_route_double_slashes() {
assert!(parse_session_proxy_route("/api/session//9222/tabs").is_err());
assert!(parse_session_proxy_route("/api//session/9222/tabs").is_err());
assert!(parse_session_proxy_route("//api/session/9222/tabs").is_err());
}
#[test]
fn test_parse_session_proxy_route_trailing_slash() {
assert!(parse_session_proxy_route("/api/session/9222/tabs/").is_err());
assert!(parse_session_proxy_route("/api/session/9222/status/").is_err());
assert!(parse_session_proxy_route("/api/session/9222/stream/").is_err());
}
#[test]
fn test_parse_session_proxy_route_encoded_paths() {
assert!(parse_session_proxy_route("/api/session/9222/tabs%20extra").is_err());
assert!(parse_session_proxy_route("/api/session/%39%32%32%32/tabs").is_err());
}
#[test]
fn test_sessions_json_has_active_port() {
let sessions_json = r#"[
{"session":"alpha","port":9222,"engine":"chrome"},
{"session":"beta","port":9333,"engine":"chrome"}
]"#;
assert_eq!(sessions_json_has_active_port(sessions_json, 9222), Ok(true));
assert_eq!(
sessions_json_has_active_port(sessions_json, 9444),
Ok(false)
);
}
#[test]
fn test_sessions_json_has_active_port_invalid_json() {
assert!(sessions_json_has_active_port("{", 9222).is_err());
}
#[test]
fn test_parse_upstream_http_response() {
let response = b"HTTP/1.1 200 OK\r\nContent-Type: application/json; charset=utf-8\r\nConnection: close\r\n\r\n{\"ok\":true}";
let parsed = parse_upstream_http_response(response).expect("response should parse");
assert_eq!(parsed.0, "200 OK");
assert_eq!(parsed.1, "application/json; charset=utf-8");
assert_eq!(parsed.2, b"{\"ok\":true}".to_vec());
}
}
+118
View File
@@ -0,0 +1,118 @@
use serde_json::{json, Value};
use std::path::Path;
use crate::connection::get_socket_dir;
pub(super) fn discover_sessions() -> String {
let dir = get_socket_dir();
let mut sessions = Vec::new();
if let Ok(entries) = std::fs::read_dir(&dir) {
for entry in entries.flatten() {
let name = entry.file_name();
let name_str = name.to_string_lossy();
if let Some(session) = name_str.strip_suffix(".stream") {
if let Ok(port_str) = std::fs::read_to_string(entry.path()) {
if let Ok(port) = port_str.trim().parse::<u16>() {
let pid_path = dir.join(format!("{}.pid", session));
if is_process_alive(&pid_path) {
let engine_path = dir.join(format!("{}.engine", session));
let engine = std::fs::read_to_string(&engine_path)
.ok()
.filter(|s| !s.trim().is_empty())
.unwrap_or_else(|| "chrome".to_string());
let provider_path = dir.join(format!("{}.provider", session));
let provider = std::fs::read_to_string(&provider_path)
.ok()
.filter(|s| !s.trim().is_empty());
let extensions = read_extensions_metadata(&dir, session);
let mut entry = json!({
"session": session,
"port": port,
"engine": engine.trim(),
});
if let Some(ref p) = provider {
entry["provider"] = json!(p.trim());
}
if !extensions.is_empty() {
entry["extensions"] = json!(extensions);
}
sessions.push(entry);
} else {
let _ = std::fs::remove_file(entry.path());
}
}
}
}
}
}
serde_json::to_string(&sessions).unwrap_or_else(|_| "[]".to_string())
}
fn read_extensions_metadata(dir: &std::path::Path, session: &str) -> Vec<Value> {
let ext_path = dir.join(format!("{}.extensions", session));
let ext_str = match std::fs::read_to_string(&ext_path) {
Ok(s) => s,
Err(_) => return Vec::new(),
};
ext_str
.split(',')
.map(|p| p.trim())
.filter(|p| !p.is_empty())
.filter_map(|path| {
let manifest_path = std::path::Path::new(path).join("manifest.json");
let manifest_str = std::fs::read_to_string(&manifest_path).ok()?;
let manifest: Value = serde_json::from_str(&manifest_str).ok()?;
let name = manifest
.get("name")
.and_then(|v| v.as_str())
.unwrap_or("Unknown")
.to_string();
let version = manifest
.get("version")
.and_then(|v| v.as_str())
.unwrap_or("")
.to_string();
let description = manifest
.get("description")
.and_then(|v| v.as_str())
.map(|s| s.to_string());
let mut ext = json!({
"name": name,
"version": version,
"path": path,
});
if let Some(desc) = description {
ext["description"] = json!(desc);
}
Some(ext)
})
.collect()
}
fn is_process_alive(pid_path: &Path) -> bool {
let pid_str = match std::fs::read_to_string(pid_path) {
Ok(s) => s,
Err(_) => return false,
};
let pid: u32 = match pid_str.trim().parse() {
Ok(p) => p,
Err(_) => return false,
};
#[cfg(unix)]
{
unsafe { libc::kill(pid as i32, 0) == 0 }
}
#[cfg(not(unix))]
{
let _ = pid;
true
}
}
+715
View File
@@ -0,0 +1,715 @@
use rust_embed::Embed;
use serde_json::{json, Value};
use std::sync::Arc;
use tokio::io::{AsyncReadExt, AsyncWriteExt};
use tokio::sync::RwLock;
use crate::connection::get_socket_dir;
#[cfg(windows)]
use crate::connection::resolve_port;
use super::chat::{chat_status_json, handle_chat_request, handle_models_request};
use super::dashboard::spawn_session;
use super::discovery::discover_sessions;
#[derive(Embed)]
#[folder = "../packages/dashboard/out/"]
struct DashboardAssets;
pub(super) const CORS_HEADERS: &str = "Access-Control-Allow-Origin: *\r\nAccess-Control-Allow-Methods: GET, POST, OPTIONS\r\nAccess-Control-Allow-Headers: Content-Type\r\n";
/// Build CORS headers that reflect the request origin only when it passes
/// `is_allowed_origin`. Used for sensitive endpoints (chat, models) so the
/// API key is not accessible from arbitrary web pages.
pub(super) fn cors_headers_for_origin(origin: Option<&str>) -> String {
let allowed_origin = match origin {
Some(o) if super::is_allowed_origin(Some(o)) => o,
_ => "http://localhost",
};
format!(
"Access-Control-Allow-Origin: {}\r\nAccess-Control-Allow-Methods: GET, POST, OPTIONS\r\nAccess-Control-Allow-Headers: Content-Type\r\n",
allowed_origin
)
}
fn request_headers(request: &str) -> &str {
request
.find("\r\n\r\n")
.or_else(|| request.find("\n\n"))
.map(|header_end| &request[..header_end])
.unwrap_or(request)
}
fn request_header_value<'a>(request: &'a str, name: &str) -> Option<&'a str> {
request_headers(request).lines().find_map(|line| {
let (header_name, value) = line.split_once(':')?;
if header_name.trim().eq_ignore_ascii_case(name) {
Some(value.trim())
} else {
None
}
})
}
fn parse_origin(peeked: &[u8]) -> Option<String> {
let header_str = std::str::from_utf8(peeked).ok()?;
request_header_value(header_str, "origin").map(ToString::to_string)
}
fn normalize_origin_authority(origin: &str) -> Option<String> {
let url = url::Url::parse(origin).ok()?;
let host = url.host_str()?.to_ascii_lowercase();
let host = if host.contains(':') {
format!("[{host}]")
} else {
host
};
let default_port = (url.scheme() == "http" && url.port() == Some(80))
|| (url.scheme() == "https" && url.port() == Some(443));
Some(match url.port() {
Some(port) if !default_port => format!("{host}:{port}"),
_ => host,
})
}
fn normalize_host_authority(host: &str) -> String {
let host = host.trim().to_ascii_lowercase();
if let Some(bracket_end) = host.rfind(']') {
if bracket_end == host.len() - 1 {
return host;
}
if host.as_bytes().get(bracket_end + 1) == Some(&b':') {
let port = &host[bracket_end + 2..];
if port == "80" || port == "443" {
return host[..=bracket_end].to_string();
}
}
return host;
}
if let Some((name, port)) = host.rsplit_once(':') {
if !name.contains(':') && (port == "80" || port == "443") {
return name.to_string();
}
}
host
}
fn authority_host(authority: &str) -> &str {
if let Some(stripped) = authority.strip_prefix('[') {
if let Some(bracket_end) = stripped.find(']') {
return &authority[..=bracket_end + 1];
}
}
if let Some((host, _port)) = authority.rsplit_once(':') {
if !host.contains(':') {
return host;
}
}
authority
}
fn is_loopback_authority(authority: &str) -> bool {
matches!(
authority_host(authority),
"localhost" | "127.0.0.1" | "::1" | "[::1]"
)
}
fn header_authority_matches_host(request: &str, header_name: &str) -> bool {
let Some(authority) =
request_header_value(request, header_name).and_then(normalize_origin_authority)
else {
return false;
};
let Some(host) = request_header_value(request, "host").map(normalize_host_authority) else {
return false;
};
authority == host && is_loopback_authority(&authority) && is_loopback_authority(&host)
}
/// Protects the command relay by requiring same-origin browser metadata.
fn is_same_origin_command_request(request: &str) -> bool {
if request_header_value(request, "origin").is_some() {
header_authority_matches_host(request, "origin")
} else {
header_authority_matches_host(request, "referer")
}
}
fn command_cors_headers(request: &str) -> String {
match request_header_value(request, "origin") {
Some(origin) if is_same_origin_command_request(request) => format!(
"Access-Control-Allow-Origin: {origin}\r\nAccess-Control-Allow-Methods: POST, OPTIONS\r\nAccess-Control-Allow-Headers: Content-Type\r\nVary: Origin\r\n"
),
_ => String::new(),
}
}
async fn write_json_error_response_no_cors(
stream: &mut tokio::net::TcpStream,
status: &str,
error: &str,
) {
let body = format!(
r#"{{"success":false,"error":{}}}"#,
serde_json::to_string(error).unwrap_or_else(|_| format!("\"{}\"", error))
);
let response = format!(
"HTTP/1.1 {status}\r\nContent-Type: application/json; charset=utf-8\r\nContent-Length: {}\r\nConnection: close\r\n\r\n",
body.len()
);
let _ = stream.write_all(response.as_bytes()).await;
let _ = stream.write_all(body.as_bytes()).await;
}
pub(super) async fn handle_http_request(
mut stream: tokio::net::TcpStream,
peeked: &[u8],
last_tabs: &Arc<RwLock<Vec<Value>>>,
last_engine: &Arc<RwLock<String>>,
session_name: &str,
) {
let peeked_len = peeked.len();
let mut discard = vec![0u8; peeked_len];
let _ = stream.read_exact(&mut discard).await;
let request = String::from_utf8_lossy(peeked);
let first_line = request.lines().next().unwrap_or("");
let method = first_line.split_whitespace().next().unwrap_or("GET");
let path = first_line.split_whitespace().nth(1).unwrap_or("/");
let origin = parse_origin(peeked);
if method == "OPTIONS" {
if path == "/api/command" {
if !is_same_origin_command_request(&request) {
write_json_error_response_no_cors(
&mut stream,
"403 Forbidden",
"Origin or Referer does not match Host header.",
)
.await;
return;
}
let cors_headers = command_cors_headers(&request);
let response = format!(
"HTTP/1.1 204 No Content\r\n{cors_headers}Access-Control-Max-Age: 86400\r\nContent-Length: 0\r\nConnection: close\r\n\r\n"
);
let _ = stream.write_all(response.as_bytes()).await;
return;
}
let response = format!(
"HTTP/1.1 204 No Content\r\n{CORS_HEADERS}Access-Control-Max-Age: 86400\r\nContent-Length: 0\r\nConnection: close\r\n\r\n"
);
let _ = stream.write_all(response.as_bytes()).await;
return;
}
if method == "POST" {
if path == "/api/command" && !is_same_origin_command_request(&request) {
write_json_error_response_no_cors(
&mut stream,
"403 Forbidden",
"Origin or Referer does not match Host header.",
)
.await;
return;
}
let full_body = read_full_body(&mut stream, peeked).await;
if full_body.is_none()
&& (path == "/api/chat" || path == "/api/sessions" || path == "/api/command")
{
let body = r#"{"error":"Request body too large"}"#;
let cors_headers = if path == "/api/command" {
command_cors_headers(&request)
} else {
CORS_HEADERS.to_string()
};
let response = format!(
"HTTP/1.1 413 Payload Too Large\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n{cors_headers}\r\n",
body.len()
);
let _ = stream.write_all(response.as_bytes()).await;
let _ = stream.write_all(body.as_bytes()).await;
return;
}
let body_str = full_body.as_deref().unwrap_or("");
if path == "/api/sessions" {
let result = spawn_session(body_str).await;
let (status, resp_body) = match result {
Ok(msg) => ("200 OK", msg),
Err(e) => (
"400 Bad Request",
format!(
r#"{{"success":false,"error":{}}}"#,
serde_json::to_string(&e).unwrap_or_else(|_| format!("\"{}\"", e))
),
),
};
let response = format!(
"HTTP/1.1 {status}\r\nContent-Type: application/json; charset=utf-8\r\nContent-Length: {}\r\nConnection: close\r\n{CORS_HEADERS}\r\n",
resp_body.len()
);
let _ = stream.write_all(response.as_bytes()).await;
let _ = stream.write_all(resp_body.as_bytes()).await;
return;
}
if path == "/api/command" {
let result = relay_command_to_daemon(session_name, body_str).await;
let (status, resp_body) = match result {
Ok(resp) => ("200 OK", resp),
Err(e) => (
"502 Bad Gateway",
format!(
r#"{{"success":false,"error":{}}}"#,
serde_json::to_string(&e).unwrap_or_else(|_| format!("\"{}\"", e))
),
),
};
let cors_headers = command_cors_headers(&request);
let response = format!(
"HTTP/1.1 {status}\r\nContent-Type: application/json; charset=utf-8\r\nContent-Length: {}\r\nConnection: close\r\n{cors_headers}\r\n",
resp_body.len()
);
let _ = stream.write_all(response.as_bytes()).await;
let _ = stream.write_all(resp_body.as_bytes()).await;
return;
}
if path == "/api/chat" {
handle_chat_request(&mut stream, body_str, origin.as_deref()).await;
return;
}
}
if method == "GET" && path == "/api/models" {
handle_models_request(&mut stream, origin.as_deref()).await;
return;
}
let (status, content_type, body): (&str, &str, Vec<u8>) = if path == "/api/sessions" {
(
"200 OK",
"application/json; charset=utf-8",
discover_sessions().into_bytes(),
)
} else if path == "/api/tabs" {
let tabs = last_tabs.read().await;
(
"200 OK",
"application/json; charset=utf-8",
serde_json::to_string(&*tabs)
.unwrap_or_else(|_| "[]".to_string())
.into_bytes(),
)
} else if path == "/api/status" {
let engine = last_engine.read().await;
(
"200 OK",
"application/json; charset=utf-8",
format!(r#"{{"engine":"{}"}}"#, *engine).into_bytes(),
)
} else if path == "/api/chat/status" {
(
"200 OK",
"application/json; charset=utf-8",
chat_status_json().into_bytes(),
)
} else {
serve_embedded_file(path)
};
let response = format!(
"HTTP/1.1 {}\r\nContent-Type: {}\r\nContent-Length: {}\r\nConnection: close\r\n{CORS_HEADERS}\r\n",
status,
content_type,
body.len()
);
let _ = stream.write_all(response.as_bytes()).await;
let _ = stream.write_all(&body).await;
}
fn find_header_end(buf: &[u8]) -> Option<usize> {
buf.windows(4)
.position(|w| w == b"\r\n\r\n")
.map(|p| p + 4)
.or_else(|| buf.windows(2).position(|w| w == b"\n\n").map(|p| p + 2))
}
fn parse_content_length_bytes(headers: &[u8]) -> Option<usize> {
let header_str = std::str::from_utf8(headers).ok()?;
for line in header_str.lines() {
if line.len() > 16 && line[..16].eq_ignore_ascii_case("content-length: ") {
return line[16..].trim().parse().ok();
}
}
None
}
const MAX_BODY_SIZE: usize = 10 * 1024 * 1024;
async fn read_full_body(stream: &mut tokio::net::TcpStream, peeked: &[u8]) -> Option<String> {
let body_offset = find_header_end(peeked)?;
let content_length = parse_content_length_bytes(&peeked[..body_offset])?;
if content_length == 0 {
return Some(String::new());
}
if content_length > MAX_BODY_SIZE {
return None;
}
let peeked_body = &peeked[body_offset..];
let peeked_body_len = peeked_body.len().min(content_length);
let mut body = Vec::with_capacity(content_length);
body.extend_from_slice(&peeked_body[..peeked_body_len]);
let remaining = content_length - peeked_body_len;
if remaining > 0 {
let mut rest = vec![0u8; remaining];
if stream.read_exact(&mut rest).await.is_err() {
return String::from_utf8(body).ok();
}
body.extend_from_slice(&rest);
}
String::from_utf8(body).ok()
}
pub(super) async fn relay_command_to_daemon(
session_name: &str,
body: &str,
) -> Result<String, String> {
let mut cmd: Value = serde_json::from_str(body).map_err(|e| format!("Invalid JSON: {}", e))?;
if cmd.get("id").is_none() {
let id = format!(
"dash-{}",
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.unwrap_or_default()
.as_millis()
);
cmd["id"] = json!(id);
}
let mut json_str = serde_json::to_string(&cmd).map_err(|e| e.to_string())?;
json_str.push('\n');
#[cfg(unix)]
let stream = {
let socket_path = get_socket_dir().join(format!("{}.sock", session_name));
tokio::net::UnixStream::connect(&socket_path)
.await
.map_err(|e| format!("Failed to connect to daemon: {}", e))?
};
#[cfg(windows)]
let stream = {
let port = resolve_port(session_name);
tokio::net::TcpStream::connect(format!("127.0.0.1:{}", port))
.await
.map_err(|e| format!("Failed to connect to daemon: {}", e))?
};
let (reader, mut writer) = tokio::io::split(stream);
writer
.write_all(json_str.as_bytes())
.await
.map_err(|e| format!("Failed to send command: {}", e))?;
let mut buf_reader = tokio::io::BufReader::new(reader);
let mut response_line = String::new();
tokio::io::AsyncBufReadExt::read_line(&mut buf_reader, &mut response_line)
.await
.map_err(|e| format!("Failed to read response: {}", e))?;
Ok(response_line.trim().to_string())
}
pub(super) fn serve_embedded_file(url_path: &str) -> (&'static str, &'static str, Vec<u8>) {
let clean = url_path.trim_start_matches('/');
let key = if clean.is_empty() {
"index.html"
} else {
clean
};
let file = DashboardAssets::get(key).or_else(|| DashboardAssets::get("index.html"));
match file {
Some(content) => {
let ext = key.rsplit('.').next().unwrap_or("");
let ct = match ext {
"html" => "text/html; charset=utf-8",
"js" => "application/javascript; charset=utf-8",
"css" => "text/css; charset=utf-8",
"json" => "application/json; charset=utf-8",
"svg" => "image/svg+xml",
"png" => "image/png",
"ico" => "image/x-icon",
"woff2" => "font/woff2",
"woff" => "font/woff",
"txt" => "text/plain; charset=utf-8",
_ => "application/octet-stream",
};
("200 OK", ct, content.data.to_vec())
}
None => (
"404 Not Found",
"text/html; charset=utf-8",
b"<html><body><p>404 Not Found</p></body></html>".to_vec(),
),
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::test_utils::EnvGuard;
use std::sync::Arc;
use tokio::io::{AsyncBufReadExt, AsyncReadExt, AsyncWriteExt};
use tokio::net::TcpListener;
use tokio::sync::oneshot;
async fn send_request_to_handler(request: &str, session_name: &str) -> String {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let addr = listener.local_addr().unwrap();
let peeked = request.as_bytes().to_vec();
let last_tabs = Arc::new(RwLock::new(Vec::new()));
let last_engine = Arc::new(RwLock::new("chrome".to_string()));
let session_name = session_name.to_string();
let server = tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
handle_http_request(stream, &peeked, &last_tabs, &last_engine, &session_name).await;
});
let mut client = tokio::net::TcpStream::connect(addr).await.unwrap();
client.write_all(request.as_bytes()).await.unwrap();
client.shutdown().await.unwrap();
let mut response = Vec::new();
client.read_to_end(&mut response).await.unwrap();
server.await.unwrap();
String::from_utf8(response).unwrap()
}
#[cfg(unix)]
async fn spawn_fake_daemon(
socket_dir: &std::path::Path,
session_name: &str,
) -> oneshot::Receiver<String> {
let socket_path = socket_dir.join(format!("{session_name}.sock"));
let _ = std::fs::remove_file(&socket_path);
let listener = tokio::net::UnixListener::bind(&socket_path).unwrap();
let (tx, rx) = oneshot::channel();
tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
let mut reader = tokio::io::BufReader::new(stream);
let mut line = String::new();
reader.read_line(&mut line).await.unwrap();
let mut stream = reader.into_inner();
stream
.write_all(br#"{"success":true,"data":{"ok":true}}"#)
.await
.unwrap();
stream.write_all(b"\n").await.unwrap();
let _ = tx.send(line);
});
rx
}
#[cfg(unix)]
#[tokio::test(flavor = "current_thread")]
async fn cross_origin_command_post_is_rejected_without_relaying_to_daemon() {
let temp_parent = std::path::Path::new(env!("CARGO_MANIFEST_DIR"))
.join("target")
.join("t");
std::fs::create_dir_all(&temp_parent).unwrap();
let socket_dir = tempfile::Builder::new()
.prefix("ab-")
.tempdir_in(temp_parent)
.unwrap();
let guard = EnvGuard::new(&["AGENT_BROWSER_SOCKET_DIR", "XDG_RUNTIME_DIR"]);
guard.set(
"AGENT_BROWSER_SOCKET_DIR",
socket_dir.path().to_str().unwrap(),
);
guard.remove("XDG_RUNTIME_DIR");
let session_name = "x";
let daemon_command = spawn_fake_daemon(socket_dir.path(), session_name).await;
let body = r#"{"action":"tabs"}"#;
let request = format!(
"POST /api/command HTTP/1.1\r\nHost: localhost:7777\r\nOrigin: https://evil.example\r\nContent-Type: application/json\r\nContent-Length: {}\r\n\r\n{}",
body.len(),
body
);
let response = send_request_to_handler(&request, session_name).await;
assert!(
response.starts_with("HTTP/1.1 403 Forbidden"),
"unexpected response: {response}"
);
assert!(
tokio::time::timeout(std::time::Duration::from_millis(50), daemon_command)
.await
.is_err(),
"cross-origin request reached daemon command relay"
);
}
#[tokio::test(flavor = "current_thread")]
async fn cross_origin_command_preflight_is_rejected_without_wildcard_cors() {
let request = concat!(
"OPTIONS /api/command HTTP/1.1\r\n",
"Host: localhost:7777\r\n",
"Origin: https://evil.example\r\n",
"Access-Control-Request-Method: POST\r\n",
"Access-Control-Request-Headers: content-type\r\n",
"\r\n"
);
let response = send_request_to_handler(request, "x").await;
assert!(
response.starts_with("HTTP/1.1 403 Forbidden"),
"unexpected response: {response}"
);
assert!(
!response.contains("Access-Control-Allow-Origin: *"),
"forbidden command preflight exposed wildcard CORS: {response}"
);
}
#[tokio::test(flavor = "current_thread")]
async fn command_post_without_origin_or_referer_is_rejected() {
let body = r#"{"action":"tabs"}"#;
let request = format!(
"POST /api/command HTTP/1.1\r\nHost: localhost:7777\r\nContent-Type: application/json\r\nContent-Length: {}\r\n\r\n{}",
body.len(),
body
);
let response = send_request_to_handler(&request, "x").await;
assert!(
response.starts_with("HTTP/1.1 403 Forbidden"),
"unexpected response: {response}"
);
assert!(
!response.contains("Access-Control-Allow-Origin: *"),
"forbidden command response exposed wildcard CORS: {response}"
);
}
#[tokio::test(flavor = "current_thread")]
async fn command_post_with_dns_rebinding_host_is_rejected() {
let body = r#"{"action":"tabs"}"#;
let request = format!(
"POST /api/command HTTP/1.1\r\nHost: attacker.example:7777\r\nOrigin: http://attacker.example:7777\r\nContent-Type: application/json\r\nContent-Length: {}\r\n\r\n{}",
body.len(),
body
);
let response = send_request_to_handler(&request, "x").await;
assert!(
response.starts_with("HTTP/1.1 403 Forbidden"),
"unexpected response: {response}"
);
assert!(
!response.contains("Access-Control-Allow-Origin: *"),
"forbidden command response exposed wildcard CORS: {response}"
);
}
#[tokio::test(flavor = "current_thread")]
async fn command_post_ignores_header_like_body_lines() {
let body = "Referer: http://localhost:7777\r\n{\"action\":\"tabs\"}";
let request = format!(
"POST /api/command HTTP/1.1\r\nHost: localhost:7777\r\nContent-Type: application/json\r\nContent-Length: {}\r\n\r\n{}",
body.len(),
body
);
let response = send_request_to_handler(&request, "x").await;
assert!(
response.starts_with("HTTP/1.1 403 Forbidden"),
"unexpected response: {response}"
);
assert!(
!response.contains("Access-Control-Allow-Origin: *"),
"forbidden command response exposed wildcard CORS: {response}"
);
}
#[cfg(unix)]
#[tokio::test(flavor = "current_thread")]
async fn same_origin_command_post_relays_without_wildcard_cors() {
let temp_parent = std::path::Path::new(env!("CARGO_MANIFEST_DIR"))
.join("target")
.join("t");
std::fs::create_dir_all(&temp_parent).unwrap();
let socket_dir = tempfile::Builder::new()
.prefix("ab-")
.tempdir_in(temp_parent)
.unwrap();
let guard = EnvGuard::new(&["AGENT_BROWSER_SOCKET_DIR", "XDG_RUNTIME_DIR"]);
guard.set(
"AGENT_BROWSER_SOCKET_DIR",
socket_dir.path().to_str().unwrap(),
);
guard.remove("XDG_RUNTIME_DIR");
let session_name = "x";
let daemon_command = spawn_fake_daemon(socket_dir.path(), session_name).await;
let body = r#"{"action":"tabs"}"#;
let request = format!(
"POST /api/command HTTP/1.1\r\nHost: localhost:7777\r\nOrigin: http://localhost:7777\r\nContent-Type: application/json\r\nContent-Length: {}\r\n\r\n{}",
body.len(),
body
);
let response = send_request_to_handler(&request, session_name).await;
assert!(
response.starts_with("HTTP/1.1 200 OK"),
"unexpected response: {response}"
);
assert!(
response.contains("Access-Control-Allow-Origin: http://localhost:7777"),
"same-origin command response did not reflect origin: {response}"
);
assert!(
!response.contains("Access-Control-Allow-Origin: *"),
"same-origin command response exposed wildcard CORS: {response}"
);
let relayed = tokio::time::timeout(std::time::Duration::from_secs(1), daemon_command)
.await
.unwrap()
.unwrap();
assert!(relayed.contains(r#""action":"tabs""#), "{relayed}");
}
}
+486
View File
@@ -0,0 +1,486 @@
mod cdp_loop;
pub(crate) mod chat;
mod dashboard;
mod discovery;
mod http;
mod websocket;
pub use cdp_loop::{ack_screencast_frame, start_screencast, stop_screencast};
pub use dashboard::run_dashboard_server;
use serde_json::{json, Value};
use std::sync::Arc;
use tokio::net::TcpListener;
use tokio::sync::{broadcast, watch, Mutex, Notify, RwLock};
use super::cdp::client::CdpClient;
/// Frame metadata from CDP Page.screencastFrame events.
#[derive(Debug, Clone)]
pub struct FrameMetadata {
pub offset_top: f64,
pub page_scale_factor: f64,
pub device_width: u32,
pub device_height: u32,
pub scroll_offset_x: f64,
pub scroll_offset_y: f64,
pub timestamp: u64,
}
impl Default for FrameMetadata {
fn default() -> Self {
Self {
offset_top: 0.0,
page_scale_factor: 1.0,
device_width: 1280,
device_height: 720,
scroll_offset_x: 0.0,
scroll_offset_y: 0.0,
timestamp: 0,
}
}
}
pub struct StreamServer {
port: u16,
session_name: String,
frame_tx: broadcast::Sender<String>,
client_count: Arc<Mutex<usize>>,
client_slot: Arc<RwLock<Option<Arc<CdpClient>>>>,
/// The active CDP page session ID (from Target.attachToTarget).
cdp_session_id: Arc<RwLock<Option<String>>>,
client_notify: Arc<Notify>,
screencasting: Arc<Mutex<bool>>,
viewport_width: Arc<Mutex<u32>>,
viewport_height: Arc<Mutex<u32>>,
last_tabs: Arc<RwLock<Vec<Value>>>,
last_engine: Arc<RwLock<String>>,
last_frame: Arc<RwLock<Option<String>>>,
recording: Arc<Mutex<bool>>,
shutdown_tx: watch::Sender<bool>,
accept_task: Mutex<Option<tokio::task::JoinHandle<()>>>,
cdp_task: Mutex<Option<tokio::task::JoinHandle<()>>>,
}
impl StreamServer {
pub async fn start(
preferred_port: u16,
client: Arc<CdpClient>,
session_id: String,
) -> Result<Self, String> {
let client_slot = Arc::new(RwLock::new(Some(client)));
let (server, _) = Self::start_inner(preferred_port, client_slot, session_id, true).await?;
Ok(server)
}
/// Start the stream server without a CDP client.
/// Returns the server and a shared slot to set the client when the browser launches.
/// Input messages are ignored until the client is set.
/// When `allow_port_fallback` is true, binding to an occupied port falls back to an
/// OS-assigned port (used by daemon startup). When false, the error propagates
/// (used by the runtime `stream_enable` command).
pub async fn start_without_client(
preferred_port: u16,
session_id: String,
allow_port_fallback: bool,
) -> Result<(Self, Arc<RwLock<Option<Arc<CdpClient>>>>), String> {
let client_slot = Arc::new(RwLock::new(None::<Arc<CdpClient>>));
Self::start_inner(preferred_port, client_slot, session_id, allow_port_fallback).await
}
/// Notify the background CDP listener that the client has changed (browser launched/closed).
pub fn notify_client_changed(&self) {
self.client_notify.notify_one();
}
/// Update the active CDP page session ID used for screencast commands.
pub async fn set_cdp_session_id(&self, session_id: Option<String>) {
let mut guard = self.cdp_session_id.write().await;
*guard = session_id;
}
/// Check whether the server currently has active screencast running.
pub async fn is_screencasting(&self) -> bool {
*self.screencasting.lock().await
}
/// Update the stored viewport dimensions and restart the active screencast (if any)
/// so frames are captured at the new size.
pub async fn set_viewport(&self, width: u32, height: u32) {
let mut vw = self.viewport_width.lock().await;
let mut vh = self.viewport_height.lock().await;
if *vw == width && *vh == height {
return;
}
*vw = width;
*vh = height;
drop(vw);
drop(vh);
self.client_notify.notify_one();
}
/// Get the current viewport dimensions.
pub async fn viewport(&self) -> (u32, u32) {
let w = *self.viewport_width.lock().await;
let h = *self.viewport_height.lock().await;
(w, h)
}
/// Override the cached screencast state for explicit CLI start/stop commands.
pub async fn set_screencasting(&self, active: bool) {
let mut guard = self.screencasting.lock().await;
*guard = active;
}
/// Update and broadcast the recording state.
pub async fn set_recording(&self, active: bool, engine: &str) {
*self.recording.lock().await = active;
let connected = self.client_slot.read().await.is_some();
let sc = *self.screencasting.lock().await;
let (vw, vh) = self.viewport().await;
self.broadcast_status(connected, sc, vw, vh, engine).await;
}
/// Shut down the accept loop and background CDP listener, releasing the bound port.
pub async fn shutdown(&self) {
let _ = self.shutdown_tx.send(true);
if let Some(task) = self.accept_task.lock().await.take() {
let _ = task.await;
}
if let Some(task) = self.cdp_task.lock().await.take() {
let _ = task.await;
}
}
async fn start_inner(
preferred_port: u16,
client_slot: Arc<RwLock<Option<Arc<CdpClient>>>>,
session_id: String,
allow_port_fallback: bool,
) -> Result<(Self, Arc<RwLock<Option<Arc<CdpClient>>>>), String> {
let addr = format!("127.0.0.1:{}", preferred_port);
let listener = match TcpListener::bind(&addr).await {
Ok(l) => l,
Err(_) if allow_port_fallback && preferred_port != 0 => {
TcpListener::bind("127.0.0.1:0")
.await
.map_err(|e| format!("Failed to bind stream server: {}", e))?
}
Err(e) => return Err(format!("Failed to bind stream server: {}", e)),
};
let actual_addr = listener
.local_addr()
.map_err(|e| format!("Failed to get stream address: {}", e))?;
let port = actual_addr.port();
let (frame_tx, _) = broadcast::channel::<String>(64);
let client_count = Arc::new(Mutex::new(0usize));
let client_notify = Arc::new(Notify::new());
let screencasting = Arc::new(Mutex::new(false));
let cdp_session_id = Arc::new(RwLock::new(None::<String>));
let viewport_width = Arc::new(Mutex::new(1280u32));
let viewport_height = Arc::new(Mutex::new(720u32));
let last_tabs = Arc::new(RwLock::new(Vec::<Value>::new()));
let last_engine = Arc::new(RwLock::new("chrome".to_string()));
let last_frame = Arc::new(RwLock::new(None::<String>));
let recording = Arc::new(Mutex::new(false));
let (shutdown_tx, shutdown_rx) = watch::channel(false);
let frame_tx_clone = frame_tx.clone();
let client_count_clone = client_count.clone();
let client_slot_clone = client_slot.clone();
let notify_clone = client_notify.clone();
let screencasting_clone = screencasting.clone();
let cdp_session_clone = cdp_session_id.clone();
let vw_clone = viewport_width.clone();
let vh_clone = viewport_height.clone();
let last_tabs_clone = last_tabs.clone();
let last_engine_clone = last_engine.clone();
let last_frame_clone = last_frame.clone();
let recording_clone = recording.clone();
let accept_shutdown_rx = shutdown_rx.clone();
let session_name_clone = session_id.clone();
let accept_task = tokio::spawn(async move {
websocket::accept_loop(
listener,
frame_tx_clone,
client_count_clone,
client_slot_clone,
notify_clone,
screencasting_clone,
cdp_session_clone,
vw_clone,
vh_clone,
last_tabs_clone,
last_engine_clone,
last_frame_clone,
recording_clone,
accept_shutdown_rx,
session_name_clone,
)
.await;
});
let frame_tx_bg = frame_tx.clone();
let client_slot_bg = client_slot.clone();
let client_notify_bg = client_notify.clone();
let screencasting_bg = screencasting.clone();
let client_count_bg = client_count.clone();
let cdp_session_bg = cdp_session_id.clone();
let vw_bg = viewport_width.clone();
let vh_bg = viewport_height.clone();
let last_frame_bg = last_frame.clone();
let last_tabs_bg = last_tabs.clone();
let last_engine_bg = last_engine.clone();
let recording_bg = recording.clone();
let cdp_task = tokio::spawn(async move {
cdp_loop::cdp_event_loop(
frame_tx_bg,
client_slot_bg,
client_notify_bg,
screencasting_bg,
client_count_bg,
cdp_session_bg,
vw_bg,
vh_bg,
last_frame_bg,
last_tabs_bg,
last_engine_bg,
recording_bg,
shutdown_rx,
)
.await;
});
Ok((
Self {
port,
session_name: session_id,
frame_tx,
client_count,
client_slot: client_slot.clone(),
cdp_session_id,
client_notify,
screencasting,
viewport_width,
viewport_height,
last_tabs,
last_engine,
last_frame,
recording,
shutdown_tx,
accept_task: Mutex::new(Some(accept_task)),
cdp_task: Mutex::new(Some(cdp_task)),
},
client_slot,
))
}
pub fn port(&self) -> u16 {
self.port
}
/// Broadcast a raw frame string (legacy).
pub fn broadcast_frame(&self, frame_json: &str) {
let s = frame_json.to_string();
if let Ok(mut lf) = self.last_frame.try_write() {
*lf = Some(s.clone());
}
let _ = self.frame_tx.send(s);
}
/// Broadcast a screencast frame with structured metadata.
pub fn broadcast_screencast_frame(&self, base64_data: &str, metadata: &FrameMetadata) {
let msg = json!({
"type": "frame",
"data": base64_data,
"metadata": {
"offsetTop": metadata.offset_top,
"pageScaleFactor": metadata.page_scale_factor,
"deviceWidth": metadata.device_width,
"deviceHeight": metadata.device_height,
"scrollOffsetX": metadata.scroll_offset_x,
"scrollOffsetY": metadata.scroll_offset_y,
"timestamp": metadata.timestamp,
}
});
let s = msg.to_string();
if let Ok(mut lf) = self.last_frame.try_write() {
*lf = Some(s.clone());
}
let _ = self.frame_tx.send(s);
}
/// Broadcast a status message to all connected clients.
pub async fn broadcast_status(
&self,
connected: bool,
screencasting: bool,
viewport_width: u32,
viewport_height: u32,
engine: &str,
) {
{
let mut guard = self.last_engine.write().await;
*guard = engine.to_string();
}
let rec = *self.recording.lock().await;
let msg = json!({
"type": "status",
"connected": connected,
"screencasting": screencasting,
"viewportWidth": viewport_width,
"viewportHeight": viewport_height,
"engine": engine,
"recording": rec,
});
let _ = self.frame_tx.send(msg.to_string());
}
/// Broadcast an error message to all connected clients.
pub fn broadcast_error(&self, message: &str) {
let msg = json!({
"type": "error",
"message": message,
});
let _ = self.frame_tx.send(msg.to_string());
}
/// Broadcast a command event when a command begins executing.
pub fn broadcast_command(&self, action: &str, id: &str, params: &Value) {
let msg = json!({
"type": "command",
"action": action,
"id": id,
"params": params,
"timestamp": timestamp_ms(),
});
let _ = self.frame_tx.send(msg.to_string());
}
/// Broadcast a result event after a command finishes executing.
pub fn broadcast_result(
&self,
id: &str,
action: &str,
success: bool,
data: &Value,
duration_ms: u64,
) {
let msg = json!({
"type": "result",
"id": id,
"action": action,
"success": success,
"data": data,
"duration_ms": duration_ms,
"timestamp": timestamp_ms(),
});
let _ = self.frame_tx.send(msg.to_string());
}
/// Broadcast a console event from the browser.
pub fn broadcast_console(&self, level: &str, text: &str, args: &[Value]) {
let mut msg = json!({
"type": "console",
"level": level,
"text": text,
"timestamp": timestamp_ms(),
});
if !args.is_empty() {
msg.as_object_mut()
.unwrap()
.insert("args".to_string(), Value::Array(args.to_vec()));
}
let _ = self.frame_tx.send(msg.to_string());
}
/// Broadcast a page error (uncaught exception) from the browser.
pub fn broadcast_page_error(&self, text: &str, line: Option<i64>, column: Option<i64>) {
let msg = json!({
"type": "page_error",
"text": text,
"line": line,
"column": column,
"timestamp": timestamp_ms(),
});
let _ = self.frame_tx.send(msg.to_string());
}
/// Broadcast the current tab list so the dashboard can render a tab bar.
/// Also caches the list so newly connected WebSocket clients receive it immediately.
pub async fn broadcast_tabs(&self, tabs: &[Value]) {
{
let mut guard = self.last_tabs.write().await;
*guard = tabs.to_vec();
}
let msg = json!({
"type": "tabs",
"tabs": tabs,
"timestamp": timestamp_ms(),
});
let _ = self.frame_tx.send(msg.to_string());
}
}
pub(crate) fn timestamp_ms() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0)
}
pub fn is_allowed_origin(origin: Option<&str>) -> bool {
match origin {
None => true,
Some(o) => {
if o.starts_with("file://") {
return true;
}
if let Ok(url) = url::Url::parse(o) {
let host = url.host_str().unwrap_or("");
host == "localhost" || host == "127.0.0.1" || host == "::1" || host == "[::1]"
} else {
false
}
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_allowed_origin_none() {
assert!(is_allowed_origin(None));
}
#[test]
fn test_allowed_origin_file() {
assert!(is_allowed_origin(Some("file:///path/to/file")));
}
#[test]
fn test_allowed_origin_localhost() {
assert!(is_allowed_origin(Some("http://localhost:3000")));
assert!(is_allowed_origin(Some("http://127.0.0.1:8080")));
}
#[test]
fn test_disallowed_origin() {
assert!(!is_allowed_origin(Some("http://evil.com")));
}
#[test]
fn test_frame_metadata_default() {
let meta = FrameMetadata::default();
assert_eq!(meta.device_width, 1280);
assert_eq!(meta.device_height, 720);
assert_eq!(meta.page_scale_factor, 1.0);
}
}
+338
View File
@@ -0,0 +1,338 @@
use serde_json::{json, Value};
use std::net::SocketAddr;
use std::sync::Arc;
use futures_util::{SinkExt, StreamExt};
use tokio::net::TcpListener;
use tokio::sync::{broadcast, watch, Mutex, Notify, RwLock};
use tokio_tungstenite::tungstenite::Message;
use crate::native::cdp::client::CdpClient;
use super::http::handle_http_request;
use super::{is_allowed_origin, timestamp_ms};
#[allow(clippy::too_many_arguments)]
pub(super) async fn accept_loop(
listener: TcpListener,
frame_tx: broadcast::Sender<String>,
client_count: Arc<Mutex<usize>>,
client_slot: Arc<RwLock<Option<Arc<CdpClient>>>>,
client_notify: Arc<Notify>,
screencasting: Arc<Mutex<bool>>,
cdp_session_id: Arc<RwLock<Option<String>>>,
viewport_width: Arc<Mutex<u32>>,
viewport_height: Arc<Mutex<u32>>,
last_tabs: Arc<RwLock<Vec<Value>>>,
last_engine: Arc<RwLock<String>>,
last_frame: Arc<RwLock<Option<String>>>,
recording: Arc<Mutex<bool>>,
mut shutdown_rx: watch::Receiver<bool>,
session_name: String,
) {
let session_name: Arc<str> = Arc::from(session_name);
loop {
tokio::select! {
changed = shutdown_rx.changed() => {
if changed.is_err() || *shutdown_rx.borrow() {
break;
}
}
accept_result = listener.accept() => {
let Ok((stream, addr)) = accept_result else {
break;
};
let frame_tx = frame_tx.clone();
let client_count = client_count.clone();
let client_slot = client_slot.clone();
let client_notify = client_notify.clone();
let screencasting = screencasting.clone();
let cdp_session_id = cdp_session_id.clone();
let vw = viewport_width.clone();
let vh = viewport_height.clone();
let lt = last_tabs.clone();
let le = last_engine.clone();
let lf = last_frame.clone();
let rec = recording.clone();
let shutdown_rx = shutdown_rx.clone();
let sn = session_name.clone();
tokio::spawn(async move {
handle_connection(
stream,
addr,
frame_tx,
client_count,
client_slot,
client_notify,
screencasting,
cdp_session_id,
vw,
vh,
lt,
le,
lf,
rec,
shutdown_rx,
sn,
)
.await;
});
}
}
}
}
fn is_websocket_upgrade(request: &str) -> bool {
request.lines().any(|line| {
if let Some((name, value)) = line.split_once(':') {
name.trim().eq_ignore_ascii_case("upgrade")
&& value.trim().eq_ignore_ascii_case("websocket")
} else {
false
}
})
}
/// Peek at the TCP stream to dispatch between WebSocket upgrade and plain HTTP.
#[allow(clippy::too_many_arguments)]
async fn handle_connection(
stream: tokio::net::TcpStream,
addr: SocketAddr,
frame_tx: broadcast::Sender<String>,
client_count: Arc<Mutex<usize>>,
client_slot: Arc<RwLock<Option<Arc<CdpClient>>>>,
client_notify: Arc<Notify>,
screencasting: Arc<Mutex<bool>>,
cdp_session_id: Arc<RwLock<Option<String>>>,
viewport_width: Arc<Mutex<u32>>,
viewport_height: Arc<Mutex<u32>>,
last_tabs: Arc<RwLock<Vec<Value>>>,
last_engine: Arc<RwLock<String>>,
last_frame: Arc<RwLock<Option<String>>>,
recording: Arc<Mutex<bool>>,
shutdown_rx: watch::Receiver<bool>,
session_name: Arc<str>,
) {
let mut buf = [0u8; 4096];
let n = match stream.peek(&mut buf).await {
Ok(n) => n,
Err(_) => return,
};
let request = String::from_utf8_lossy(&buf[..n]);
if is_websocket_upgrade(&request) {
let frame_rx = frame_tx.subscribe();
handle_ws_client(
stream,
addr,
frame_rx,
client_count,
client_slot,
client_notify,
screencasting,
cdp_session_id,
viewport_width,
viewport_height,
last_tabs,
last_engine,
last_frame,
recording,
shutdown_rx,
)
.await;
} else {
handle_http_request(stream, &buf[..n], &last_tabs, &last_engine, &session_name).await;
}
}
#[allow(clippy::result_large_err, clippy::too_many_arguments)]
async fn handle_ws_client(
stream: tokio::net::TcpStream,
_addr: SocketAddr,
mut frame_rx: broadcast::Receiver<String>,
client_count: Arc<Mutex<usize>>,
client_slot: Arc<RwLock<Option<Arc<CdpClient>>>>,
client_notify: Arc<Notify>,
screencasting: Arc<Mutex<bool>>,
cdp_session_id: Arc<RwLock<Option<String>>>,
viewport_width: Arc<Mutex<u32>>,
viewport_height: Arc<Mutex<u32>>,
last_tabs: Arc<RwLock<Vec<Value>>>,
last_engine: Arc<RwLock<String>>,
last_frame: Arc<RwLock<Option<String>>>,
recording: Arc<Mutex<bool>>,
mut shutdown_rx: watch::Receiver<bool>,
) {
let callback =
|req: &tokio_tungstenite::tungstenite::handshake::server::Request,
resp: tokio_tungstenite::tungstenite::handshake::server::Response| {
let origin = req
.headers()
.get("origin")
.and_then(|v| v.to_str().ok())
.map(|s| s.to_string());
if !is_allowed_origin(origin.as_deref()) {
let mut reject =
tokio_tungstenite::tungstenite::handshake::server::ErrorResponse::new(Some(
"Origin not allowed".to_string(),
));
*reject.status_mut() = tokio_tungstenite::tungstenite::http::StatusCode::FORBIDDEN;
return Err(reject);
}
Ok(resp)
};
let ws_stream = match tokio_tungstenite::accept_hdr_async(stream, callback).await {
Ok(ws) => ws,
Err(_) => return,
};
{
let mut count = client_count.lock().await;
*count += 1;
}
let (mut ws_tx, mut ws_rx) = ws_stream.split();
{
let guard = client_slot.read().await;
let connected = guard.is_some();
let sc = *screencasting.lock().await;
let vw = *viewport_width.lock().await;
let vh = *viewport_height.lock().await;
let eng = last_engine.read().await.clone();
let rec = *recording.lock().await;
let status = json!({
"type": "status",
"connected": connected,
"screencasting": sc,
"viewportWidth": vw,
"viewportHeight": vh,
"engine": eng,
"recording": rec,
});
let _ = ws_tx.send(Message::Text(status.to_string())).await;
let tabs = last_tabs.read().await;
if !tabs.is_empty() {
let tabs_msg = json!({
"type": "tabs",
"tabs": *tabs,
"timestamp": timestamp_ms(),
});
let _ = ws_tx.send(Message::Text(tabs_msg.to_string())).await;
}
if let Some(ref cached) = *last_frame.read().await {
let _ = ws_tx.send(Message::Text(cached.clone())).await;
}
}
client_notify.notify_one();
loop {
tokio::select! {
changed = shutdown_rx.changed() => {
if changed.is_err() || *shutdown_rx.borrow() {
let _ = ws_tx.send(Message::Close(None)).await;
break;
}
}
frame = frame_rx.recv() => {
match frame {
Ok(data) => {
if ws_tx.send(Message::Text(data)).await.is_err() {
break;
}
}
Err(broadcast::error::RecvError::Lagged(_)) => {
continue;
}
Err(broadcast::error::RecvError::Closed) => break,
}
}
msg = ws_rx.next() => {
match msg {
Some(Ok(Message::Text(text))) => {
let guard = client_slot.read().await;
if let Some(ref client) = *guard {
let sid = cdp_session_id.read().await;
handle_client_message(&text, client.as_ref(), sid.as_deref()).await;
}
}
Some(Ok(Message::Close(_))) | None => break,
_ => {}
}
}
}
}
{
let mut count = client_count.lock().await;
*count = count.saturating_sub(1);
}
client_notify.notify_one();
}
async fn handle_client_message(msg: &str, client: &CdpClient, session_id: Option<&str>) {
let parsed: Value = match serde_json::from_str(msg) {
Ok(v) => v,
Err(_) => return,
};
let msg_type = parsed.get("type").and_then(|v| v.as_str()).unwrap_or("");
match msg_type {
"input_mouse" => {
let _ = client
.send_command(
"Input.dispatchMouseEvent",
Some(json!({
"type": parsed.get("eventType").and_then(|v| v.as_str()).unwrap_or("mouseMoved"),
"x": parsed.get("x").and_then(|v| v.as_f64()).unwrap_or(0.0),
"y": parsed.get("y").and_then(|v| v.as_f64()).unwrap_or(0.0),
"button": parsed.get("button").and_then(|v| v.as_str()).unwrap_or("none"),
"clickCount": parsed.get("clickCount").and_then(|v| v.as_i64()).unwrap_or(0),
"deltaX": parsed.get("deltaX").and_then(|v| v.as_f64()).unwrap_or(0.0),
"deltaY": parsed.get("deltaY").and_then(|v| v.as_f64()).unwrap_or(0.0),
"modifiers": parsed.get("modifiers").and_then(|v| v.as_i64()).unwrap_or(0),
})),
session_id,
)
.await;
}
"input_keyboard" => {
let _ = client
.send_command(
"Input.dispatchKeyEvent",
Some(json!({
"type": parsed.get("eventType").and_then(|v| v.as_str()).unwrap_or("keyDown"),
"key": parsed.get("key"),
"code": parsed.get("code"),
"text": parsed.get("text"),
"windowsVirtualKeyCode": parsed.get("windowsVirtualKeyCode").and_then(|v| v.as_i64()).unwrap_or(0),
"modifiers": parsed.get("modifiers").and_then(|v| v.as_i64()).unwrap_or(0),
})),
session_id,
)
.await;
}
"input_touch" => {
let _ = client
.send_command(
"Input.dispatchTouchEvent",
Some(json!({
"type": parsed.get("eventType").and_then(|v| v.as_str()).unwrap_or("touchStart"),
"touchPoints": parsed.get("touchPoints").unwrap_or(&json!([])),
"modifiers": parsed.get("modifiers").and_then(|v| v.as_i64()).unwrap_or(0),
})),
session_id,
)
.await;
}
"status" => {}
_ => {}
}
}
@@ -0,0 +1,18 @@
<!DOCTYPE html>
<html>
<head><title>Upload Test</title></head>
<body>
<h1>Upload Test</h1>
<label for="fileInput">Choose file:</label>
<input type="file" id="fileInput" name="fileInput">
<div id="result"></div>
<script>
document.getElementById('fileInput').addEventListener('change', function(e) {
var file = e.target.files[0];
if (file) {
document.getElementById('result').textContent = 'uploaded:' + file.name;
}
});
</script>
</body>
</html>
+347 -38
View File
@@ -296,6 +296,31 @@ pub fn print_response_with_opts(resp: &Response, action: Option<&str>, opts: &Ou
println!("{}", count);
return;
}
// Bounding box (get box)
if action == Some("boundingbox") {
if let Some(obj) = data.as_object() {
let x = obj.get("x").and_then(|v| v.as_f64()).unwrap_or(0.0);
let y = obj.get("y").and_then(|v| v.as_f64()).unwrap_or(0.0);
let w = obj.get("width").and_then(|v| v.as_f64()).unwrap_or(0.0);
let h = obj.get("height").and_then(|v| v.as_f64()).unwrap_or(0.0);
println!("x: {}", x);
println!("y: {}", y);
println!("width: {}", w);
println!("height: {}", h);
}
return;
}
// Computed styles (get styles)
if let Some(styles) = data.get("styles").and_then(|v| v.as_object()) {
for (key, val) in styles {
let display = match val.as_str() {
Some(s) => s.to_string(),
None => val.to_string(),
};
println!("{}: {}", key, display);
}
return;
}
// Boolean results
if let Some(visible) = data.get("visible").and_then(|v| v.as_bool()) {
println!("{}", visible);
@@ -381,7 +406,9 @@ pub fn print_response_with_opts(resp: &Response, action: Option<&str>, opts: &Ou
}
// Tabs
if let Some(tabs) = data.get("tabs").and_then(|v| v.as_array()) {
for (i, tab) in tabs.iter().enumerate() {
for tab in tabs {
let tab_id = tab.get("tabId").and_then(|v| v.as_str()).unwrap_or("?");
let tab_label = tab.get("label").and_then(|v| v.as_str());
let title = tab
.get("title")
.and_then(|v| v.as_str())
@@ -393,10 +420,63 @@ pub fn print_response_with_opts(resp: &Response, action: Option<&str>, opts: &Ou
} else {
" ".to_string()
};
println!("{} [{}] {} - {}", marker, i, title, url);
if let Some(label) = tab_label {
println!("{} [{}] {} {} - {}", marker, tab_id, label, title, url);
} else {
println!("{} [{}] {} - {}", marker, tab_id, title, url);
}
}
return;
}
// Tab switch
if action == Some("tab_switch") {
if let Some(tab_id) = data.get("tabId").and_then(|v| v.as_str()) {
if let Some(url) = data.get("url").and_then(|v| v.as_str()) {
println!(
"{} Switched to tab [{}] ({})",
color::success_indicator(),
tab_id,
url
);
} else {
println!(
"{} Switched to tab [{}]",
color::success_indicator(),
tab_id
);
}
return;
}
}
// New tab/window
if let Some(tab_id) = data.get("tabId").and_then(|v| v.as_str()) {
if let Some(total) = data.get("total").and_then(|v| v.as_i64()) {
let label_noun = match action {
Some("window_new") => "Window opened",
_ => "Tab opened",
};
let tab_label = data.get("label").and_then(|v| v.as_str());
if let Some(lbl) = tab_label {
println!(
"{} {} [{}] {} ({} total)",
color::success_indicator(),
label_noun,
tab_id,
lbl,
total
);
} else {
println!(
"{} {} [{}] ({} total)",
color::success_indicator(),
label_noun,
tab_id,
total
);
}
return;
}
}
// Console logs
if let Some(logs) = data.get("messages").and_then(|v| v.as_array()) {
if opts.content_boundaries {
@@ -537,7 +617,13 @@ pub fn print_response_with_opts(resp: &Response, action: Option<&str>, opts: &Ou
// Closed (browser or tab)
if data.get("closed").is_some() {
let label = match action {
Some("tab_close") => "Tab closed",
Some("tab_close") => {
if let Some(closed_id) = data.get("tabId").and_then(|v| v.as_str()) {
println!("{} Tab [{}] closed", color::success_indicator(), closed_id);
return;
}
"Tab closed"
}
_ => "Browser closed",
};
println!("{} {}", color::success_indicator(), label);
@@ -973,27 +1059,41 @@ pub fn print_command_help(command: &str) -> bool {
// === Navigation ===
"open" | "goto" | "navigate" => {
r##"
agent-browser open - Navigate to a URL
agent-browser open - Launch the browser, optionally navigate
Usage: agent-browser open <url>
Usage: agent-browser open [url]
Navigates the browser to the specified URL. If no protocol is provided,
https:// is automatically prepended.
Without a URL, launches the browser but stays on about:blank. This lets
you stage state (network routes, cookies, init scripts) before the first
real navigation useful for SSR debug, auth setup, and capturing fresh
`react suspense` / `vitals` state without noise from a prior page.
Aliases: goto, navigate
With a URL, launches and navigates. If no protocol is provided, https://
is automatically prepended.
The `goto` and `navigate` aliases still require a URL.
Global Options:
--json Output as JSON
--session <name> Use specific session
--headers <json> Set HTTP headers (scoped to this origin)
--headed Show browser window
--enable react-devtools Inject the React DevTools hook before any page JS
--init-script <path> Register a page init script (repeatable)
Examples:
agent-browser open # Launch, no nav
agent-browser open example.com
agent-browser open https://github.com
agent-browser open localhost:3000
agent-browser open api.example.com --headers '{"Authorization": "Bearer token"}'
# ^ Headers only sent to api.example.com, not other domains
# Pre-navigation setup in one turn:
agent-browser batch \
'["open"]' \
'["network","route","*","--abort","--resource-type","script"]' \
'["navigate","http://localhost:3000/target"]'
"##
}
"back" => {
@@ -1483,6 +1583,8 @@ Usage: agent-browser screenshot [selector] [path]
Captures a screenshot of the current page. If no path is provided,
saves to a temporary directory with a generated filename.
Headless Chromium screenshots hide native scrollbars for consistent image output.
Pass --hide-scrollbars false when launching to keep native scrollbars visible.
Options:
--full, -f Capture full page (not just viewport)
@@ -1544,6 +1646,7 @@ Designed for AI agents to understand page structure.
Options:
-i, --interactive Only include interactive elements
-u, --urls Include href URLs for link elements
-c, --compact Remove empty structural elements
-d, --depth <n> Limit tree depth
-s, --selector <sel> Scope snapshot to CSS selector
@@ -1555,6 +1658,7 @@ Global Options:
Examples:
agent-browser snapshot
agent-browser snapshot -i
agent-browser snapshot -i --urls
agent-browser snapshot --compact --depth 5
agent-browser snapshot -s "#main-content"
"##
@@ -1941,13 +2045,18 @@ agent-browser tab - Manage browser tabs
Usage: agent-browser tab [operation] [args]
Manage browser tabs in the current window.
Manage browser tabs in the current window. Stable tab ids look like `t1`,
`t2`, `t3`. An id is never reused within a session, so scripts can keep
referring to the same tab across commands. Optional user-assigned labels
(e.g. `docs`, `app`) are interchangeable with ids everywhere a tab ref is
accepted.
Operations:
list List all tabs (default)
new [url] Open new tab
close [index] Close tab (current if no index)
<index> Switch to tab by index
list List open tabs with their ids and labels (default)
new [url] Open a new tab
new --label <name> [url] Open a new tab with a label like `docs` or `app`
close [t<N>|label] Close a tab (current if no ref given)
<t<N>|label> Switch to a tab by id or label
Global Options:
--json Output as JSON
@@ -1958,9 +2067,12 @@ Examples:
agent-browser tab list
agent-browser tab new
agent-browser tab new https://example.com
agent-browser tab 2
agent-browser tab new --label docs https://docs.example.com
agent-browser tab t2
agent-browser tab docs
agent-browser tab close
agent-browser tab close 1
agent-browser tab close t1
agent-browser tab close docs
"##
}
@@ -2392,25 +2504,63 @@ Examples:
"##
}
// === Doctor ===
"doctor" => {
r##"
agent-browser doctor - Diagnose and repair your install
Usage: agent-browser doctor [options]
Runs a battery of checks across environment, Chrome install, daemon state,
config files, encryption key, providers, network reachability, and a live
headless browser launch test.
Auto-cleans stale daemon socket/pid/version sidecar files. Destructive
repairs (reinstalling Chrome, purging old state files, generating a missing
encryption key) are gated behind --fix.
Options:
--offline Skip network probes
--quick Skip the live headless launch test
--fix Also run destructive repairs
--json JSON output
Exit codes:
0 All checks pass (warnings OK)
1 At least one check failed
Examples:
agent-browser doctor
agent-browser doctor --offline --quick
agent-browser doctor --fix
agent-browser doctor --json
"##
}
// === Dashboard ===
"dashboard" => {
r##"
agent-browser dashboard - Observability dashboard
Usage: agent-browser dashboard [start|stop|install] [options]
Usage: agent-browser dashboard [start|stop] [options]
Manage the observability dashboard, a local web UI that shows live
browser viewports and command activity feeds for all sessions.
The dashboard is bundled into the binary and requires no separate install.
Subcommands:
start [--port <n>] Start the dashboard server (default port: 4848)
stop Stop the dashboard server
install Download and install the dashboard to ~/.agent-browser/dashboard/
Running 'agent-browser dashboard' with no subcommand is equivalent to 'dashboard start'.
The dashboard runs as a standalone background process, independent of
browser sessions. All sessions automatically stream to the dashboard.
It works from http://localhost:4848 or a proxied/forwarded URL that
reaches the dashboard server, such as https://dashboard.agent-browser.localhost
or a Coder workspace URL. The browser stays on the dashboard origin;
session tabs, status, and stream traffic are proxied internally, so
session ports do not need to be exposed.
Options:
--port <n> Port for the dashboard server (default: 4848)
@@ -2419,7 +2569,6 @@ Global Options:
--json Output as JSON
Examples:
agent-browser dashboard install
agent-browser dashboard start
agent-browser dashboard start --port 8080
agent-browser dashboard stop
@@ -2622,20 +2771,24 @@ Examples:
"batch" => {
r##"
agent-browser batch - Execute multiple commands from stdin
agent-browser batch - Execute multiple commands sequentially
Usage: echo '<json>' | agent-browser batch [options]
Usage: agent-browser batch [options] "<cmd1>" "<cmd2>" ...
echo '<json>' | agent-browser batch [options]
Reads a JSON array of commands from stdin and executes them sequentially.
Each command is an array of strings matching normal CLI arguments.
Results are printed in order, separated by blank lines (or as a JSON array
with --json).
Runs multiple commands in sequence. Commands can be passed as quoted
arguments or piped as JSON via stdin. Results are printed in order,
separated by blank lines (or as a JSON array with --json).
Options:
--bail Stop on first error (default: continue all commands)
--json Output results as a JSON array
Input Format:
Argument Mode:
Each quoted argument is a full command string:
agent-browser batch "open https://example.com" "snapshot -i" "screenshot"
Stdin Mode (JSON):
A JSON array of string arrays. Each inner array is one command:
[
["open", "https://example.com"],
@@ -2646,12 +2799,102 @@ Input Format:
]
Examples:
agent-browser batch "open https://example.com" "screenshot"
agent-browser batch --bail "open https://example.com" "click @e1" "screenshot"
echo '[["open", "https://example.com"], ["snapshot"]]' | agent-browser batch
echo '[["open", "https://example.com"], ["get", "title"]]' | agent-browser batch --json
agent-browser batch --bail < commands.json
"##
}
"profiles" => {
r##"
agent-browser profiles - List available Chrome profiles
Usage: agent-browser profiles
Lists all Chrome profiles found in your Chrome user data directory, showing
the directory name and display name for each profile. Use the directory name
with --profile to launch Chrome with that profile's login state.
Global Options:
--json Output as JSON
Examples:
agent-browser profiles
agent-browser profiles --json
agent-browser --profile Default open https://gmail.com
"##
}
"chat" => {
r##"
agent-browser chat - Natural language browser control via AI
Usage:
agent-browser chat <message> Single-shot: execute instruction and exit
agent-browser chat Interactive REPL (when stdin is a TTY)
echo "instruction" | agent-browser chat Piped input
Sends natural language instructions to an AI model that translates them
into agent-browser commands and executes them against the active session.
Requires AI_GATEWAY_API_KEY to be set.
In interactive mode, type "quit", "exit", or "q" to leave the REPL.
Chat Options:
--model <name> AI model (or AI_GATEWAY_MODEL env, default: anthropic/claude-sonnet-4.6)
-v, --verbose Show tool commands and their raw output
-q, --quiet Show only the AI text response (hide tool calls)
Global Options:
--json Structured JSON output per turn
--session <name> Target session for commands
Examples:
agent-browser chat "open google.com and search for cats"
agent-browser chat "take a screenshot of the current page"
agent-browser -q chat "summarize this page"
agent-browser -v chat "fill in the login form with test@example.com"
agent-browser --model openai/gpt-4o chat "navigate to hacker news"
agent-browser chat
"##
}
"skills" => {
r##"
agent-browser skills - List and retrieve bundled skill content
Usage: agent-browser skills [subcommand] [options]
Subcommands:
list List all available skills (default)
get <name> [name...] Output a skill's full content
get <name> --full Include references and templates
get --all Output every skill
path [name] Print filesystem path to skill directory
Options:
--json Output as JSON
The skills command serves bundled skill content that always matches the
installed CLI version. Agents should use this to get current instructions
rather than relying on cached copies.
Examples:
agent-browser skills
agent-browser skills list
agent-browser skills get core
agent-browser skills get core --full
agent-browser skills get electron --full
agent-browser skills get --all
agent-browser skills path core
agent-browser skills list --json
Environment:
AGENT_BROWSER_SKILLS_DIR Override the skills directory path
"##
}
_ => return false,
};
println!("{}", help.trim());
@@ -2665,6 +2908,20 @@ agent-browser - fast browser automation CLI for AI agents
Usage: agent-browser <command> [args] [options]
Start here (for AI agents):
agent-browser skills get core --full
Skills ship with the CLI (always version-matched) and include workflow
patterns, ref/selector usage, and copy-paste examples. Prefer this over
guessing commands from flag docs alone. Specialized skills cover Electron
apps, Slack, exploratory testing, and cloud browser providers.
skills [list] List available skills
skills get core Core usage guide (overview + common patterns)
skills get core --full Include full command reference and templates
skills get <name> Load a specialized skill (electron, slack, ...)
skills path [name] Print skill directory path
Core Commands:
open <url> Navigate to URL
click <sel> Click element (or @ref)
@@ -2715,13 +2972,14 @@ Browser Settings: agent-browser set <setting> [value]
media [dark|light] [reduced-motion]
Network: agent-browser network <action>
route <url> [--abort|--body <json>]
route <url> [--abort|--body <json>] [--resource-type <csv>]
unroute [url]
requests [--clear] [--filter <pattern>]
har <start|stop> [path]
Storage:
cookies [get|set|clear] Manage cookies (set supports --url, --domain, --path, --httpOnly, --secure, --sameSite, --expires)
Or: cookies set --curl <file> [--domain <host>] (auto-detects JSON/cURL/Cookie-header files)
storage <local|session> Manage web storage
Tabs:
@@ -2748,9 +3006,30 @@ Streaming:
stream disable Stop runtime WebSocket streaming
stream status Show streaming status and active port
React (requires `open --enable react-devtools`):
react tree Full React component tree (depth id parent name columns)
react inspect <id> Inspect one fiber (props, hooks, state, source)
react renders start Start recording re-renders via onCommitFiberRoot
react renders stop [--json] Stop and print render profile
react suspense [--only-dynamic] [--json]
Walk Suspense boundaries + classifier report
--only-dynamic hides the "static" list
Performance:
vitals [url] [--json] Core Web Vitals (LCP/CLS/TTFB/FCP/INP) +
React hydration timing when profiling build detected
SPA:
pushstate <url> SPA client-side nav. Auto-detects window.next.router.push
(triggers RSC fetch on Next.js); falls back to
history.pushState + popstate/navigate events for other frameworks
Init scripts:
removeinitscript <id> Remove a script registered via --init-script or addinitscript
Batch:
batch [--bail] Execute commands from stdin (JSON array of string arrays)
--bail stops on first error (default: continue all)
batch [--bail] ["cmd" ...] Execute multiple commands sequentially (args or stdin)
--bail stops on first error (default: continue all)
Auth Vault:
auth save <name> [opts] Save auth profile (--url, --username, --password/--password-stdin)
@@ -2767,6 +3046,11 @@ Sessions:
session Show current session name
session list List active sessions
Chat (AI):
chat <message> Send a natural language instruction (single-shot)
chat Start interactive chat (REPL mode when stdin is a TTY)
Options: --model <name>, -v/--verbose, -q/--quiet
Dashboard:
dashboard [start] Start the dashboard server (default port: 4848)
dashboard start --port <n> Start on a specific port
@@ -2776,7 +3060,9 @@ Setup:
install Install browser binaries
install --with-deps Also install system dependencies (Linux)
upgrade Upgrade to the latest version
dashboard install Install the observability dashboard
doctor [--fix] Diagnose install; auto-clean stale files
dashboard start Start the observability dashboard
profiles List available Chrome profiles
Snapshot Options:
-i, --interactive Only interactive elements
@@ -2785,20 +3071,26 @@ Snapshot Options:
-s, --selector <sel> Scope to CSS selector
Authentication:
--profile <path> Persist login sessions across restarts (cookies, IndexedDB, cache)
--profile <name|path> Chrome profile name (e.g., Default) to reuse login state,
or a directory path for a persistent custom profile
(or AGENT_BROWSER_PROFILE env)
--session-name <name> Auto-save/restore cookies and localStorage by name
(or AGENT_BROWSER_SESSION_NAME env)
--state <path> Load saved auth state (cookies + storage) from JSON file
(or AGENT_BROWSER_STATE env)
--auto-connect Connect to a running Chrome to reuse its auth state
Tip: agent-browser --auto-connect state save ./auth.json
--auto-connect Connect to a running Chrome (DEFAULT - shares cookies/sessions)
Tip: enable CDP via chrome://inspect/#remote-debugging
--launch, --new Launch a fresh browser instead of connecting to existing
--headers <json> HTTP headers scoped to URL's origin (e.g., Authorization bearer token)
Options:
--session <name> Isolated session (or AGENT_BROWSER_SESSION env)
--executable-path <path> Custom browser executable (or AGENT_BROWSER_EXECUTABLE_PATH)
--extension <path> Load browser extensions (repeatable)
--init-script <path> Register a page init script before the first navigation (repeatable)
(or AGENT_BROWSER_INIT_SCRIPTS env, comma-separated)
--enable <feature> Built-in init scripts: react-devtools (repeatable or comma-separated)
(or AGENT_BROWSER_ENABLE env)
--args <args> Browser launch args, comma or newline separated (or AGENT_BROWSER_ARGS)
e.g., --args "--no-sandbox,--disable-blink-features=AutomationControlled"
--user-agent <ua> Custom User-Agent (or AGENT_BROWSER_USER_AGENT)
@@ -2808,6 +3100,8 @@ Options:
e.g., --proxy-bypass "localhost,*.internal.com"
--ignore-https-errors Ignore HTTPS certificate errors
--allow-file-access Allow file:// URLs to access local files (Chromium only)
--hide-scrollbars <bool> Hide native scrollbars in headless Chromium screenshots (default: true)
Use --hide-scrollbars false to keep scrollbars visible
-p, --provider <name> Browser provider: ios, browserbase, kernel, browseruse, browserless, agentcore
--device <name> iOS device name (e.g., "iPhone 15 Pro")
--json JSON output
@@ -2827,6 +3121,9 @@ Options:
--confirm-interactive Interactive confirmation prompts; auto-denies if stdin is not a TTY (or AGENT_BROWSER_CONFIRM_INTERACTIVE)
--engine <name> Browser engine: chrome (default), lightpanda (or AGENT_BROWSER_ENGINE)
--no-auto-dialog Disable automatic dismissal of alert/beforeunload dialogs (or AGENT_BROWSER_NO_AUTO_DIALOG)
--model <name> AI model for chat (or AI_GATEWAY_MODEL env)
-v, --verbose Show tool commands and their raw output
-q, --quiet Show only AI text responses (hide tool calls)
--config <path> Use a custom config file (or AGENT_BROWSER_CONFIG env)
--debug Debug output
--version, -V Show version
@@ -2844,11 +3141,12 @@ Configuration:
Boolean flags accept an optional true/false value to override config:
--headed (same as --headed true)
--headed false (disables "headed": true from config)
--hide-scrollbars false (keeps native scrollbars visible in headless Chromium screenshots)
Extensions from user and project configs are merged (not replaced).
Example agent-browser.json:
{{"headed": true, "proxy": "http://localhost:8080", "profile": "./browser-data"}}
{{"headed": true, "hideScrollbars": false, "proxy": "http://localhost:8080"}}
Environment:
AGENT_BROWSER_CONFIG Path to config file (or use --config)
@@ -2858,6 +3156,8 @@ Environment:
AGENT_BROWSER_STATE_EXPIRE_DAYS Auto-delete states older than N days (default: 30)
AGENT_BROWSER_EXECUTABLE_PATH Custom browser executable path
AGENT_BROWSER_EXTENSIONS Comma-separated browser extension paths
AGENT_BROWSER_INIT_SCRIPTS Comma-separated paths to page init scripts
AGENT_BROWSER_ENABLE Comma-separated built-in init script features (e.g. react-devtools)
AGENT_BROWSER_HEADED Show browser window (not headless)
AGENT_BROWSER_JSON JSON output
AGENT_BROWSER_ANNOTATE Annotated screenshot with numbered labels and legend
@@ -2866,6 +3166,7 @@ Environment:
AGENT_BROWSER_PROVIDER Browser provider (ios, browserbase, kernel, browseruse, browserless, agentcore)
AGENT_BROWSER_AUTO_CONNECT Auto-discover and connect to running Chrome
AGENT_BROWSER_ALLOW_FILE_ACCESS Allow file:// URLs to access local files
AGENT_BROWSER_HIDE_SCROLLBARS Hide scrollbars in headless Chromium screenshots (default: true)
AGENT_BROWSER_COLOR_SCHEME Color scheme preference (dark, light, no-preference)
AGENT_BROWSER_DOWNLOAD_PATH Default download directory for browser downloads
AGENT_BROWSER_DEFAULT_TIMEOUT Default action timeout in ms (default: 25000)
@@ -2890,6 +3191,9 @@ Environment:
AGENT_BROWSER_SCREENSHOT_DIR Default screenshot output directory
AGENT_BROWSER_SCREENSHOT_QUALITY JPEG quality 0-100
AGENT_BROWSER_SCREENSHOT_FORMAT Screenshot format: png, jpeg
AI_GATEWAY_URL Vercel AI Gateway base URL (default: https://ai-gateway.vercel.sh)
AI_GATEWAY_API_KEY API key for the AI Gateway (enables chat command and dashboard AI chat)
AI_GATEWAY_MODEL Default AI model (default: anthropic/claude-sonnet-4.6, or --model flag)
Install:
npm install -g agent-browser # npm
@@ -2906,21 +3210,26 @@ Examples:
agent-browser get text @e1
agent-browser screenshot --full
agent-browser screenshot --annotate # Labeled screenshot for vision models
agent-browser wait --load networkidle # Wait for slow pages to load
agent-browser wait 2000 # Wait for slow pages to settle
agent-browser --cdp 9222 snapshot # Connect via CDP port
agent-browser --auto-connect snapshot # Auto-discover running Chrome
agent-browser stream enable # Start runtime streaming on an auto-selected port
agent-browser stream status # Inspect runtime streaming state
agent-browser --color-scheme dark open example.com # Dark mode
agent-browser --profile ~/.myapp open example.com # Persistent profile
agent-browser --profile Default open gmail.com # Reuse Chrome login state
agent-browser --profile ~/.myapp open example.com # Persistent custom profile
agent-browser profiles # List available Chrome profiles
agent-browser --session-name myapp open example.com # Auto-save/restore state
agent-browser chat "open google.com and search for cats" # AI chat (single-shot)
agent-browser chat # AI chat (interactive REPL)
agent-browser -q chat "summarize this page" # Quiet mode (text only)
Command Chaining:
Chain commands with && in a single shell call (browser persists via daemon):
agent-browser open example.com && agent-browser wait --load networkidle && agent-browser snapshot -i
agent-browser open example.com && agent-browser snapshot -i
agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "pass" && agent-browser click @e3
agent-browser open example.com && agent-browser wait --load networkidle && agent-browser screenshot page.png
agent-browser open example.com && agent-browser screenshot
iOS Simulator (requires Xcode and Appium):
agent-browser -p ios open example.com # Use default iPhone
+622
View File
@@ -0,0 +1,622 @@
use serde_json::json;
use std::env;
use std::fs;
use std::path::{Path, PathBuf};
use std::process::exit;
use crate::color;
struct SkillInfo {
name: String,
description: String,
dir: PathBuf,
/// When true, the skill is omitted from `skills list` and `skills get --all`
/// but can still be fetched by name via `skills get <name>`. Used for
/// bootstrap stubs that exist for external tooling (e.g. `npx skills add`)
/// but aren't the intended entry point for agents already inside the CLI.
hidden: bool,
}
/// Skill content is split across two directories:
///
/// - `skills/` — discovery stubs (picked up by `npx skills add`). Carry
/// `hidden: true` so they don't show up in `skills list` or `skills get
/// --all` inside the CLI, since they exist only to redirect external
/// agents to `skills get core`.
/// - `skill-data/` — runtime skill content served by the CLI (`core`,
/// `electron`, `slack`, `dogfood`, etc.).
///
/// Both are shipped in the npm package and searched by `discover_skills`.
const SKILL_DIRS: &[&str] = &["skills", "skill-data"];
/// Locate the package root that contains the skill directories.
///
/// Resolution order:
/// 1. AGENT_BROWSER_SKILLS_DIR env var (points directly at a single directory)
/// 2. ../ relative to the executable (npm installs: binary is in bin/)
/// 3. Walk up from the executable to find a project root with skills/
/// (dev builds where binary is in target/debug/ or target/release/)
fn find_package_root() -> Option<PathBuf> {
if let Ok(exe) = env::current_exe() {
let exe = exe.canonicalize().unwrap_or(exe);
if let Some(parent) = exe.parent() {
// npm install layout: bin/agent-browser-* -> ../
let candidate = parent.join("..");
if candidate.join("skills").is_dir() {
return Some(candidate.canonicalize().unwrap_or(candidate));
}
// dev build layout: walk up from target/debug/ or target/release/
let mut dir = parent;
loop {
if dir.join("skills").is_dir() {
return Some(dir.to_path_buf());
}
match dir.parent() {
Some(p) => dir = p,
None => break,
}
}
}
}
None
}
/// Collect all skill directories to search, respecting the env var override.
fn find_skills_dirs() -> Vec<PathBuf> {
// Env var override: single directory, used as-is
if let Ok(dir) = env::var("AGENT_BROWSER_SKILLS_DIR") {
let p = PathBuf::from(dir);
if p.is_dir() {
return vec![p];
}
}
let Some(root) = find_package_root() else {
return vec![];
};
SKILL_DIRS
.iter()
.map(|d| root.join(d))
.filter(|p| p.is_dir())
.collect()
}
/// Parse YAML frontmatter from a SKILL.md file. Returns (name, description, hidden).
fn parse_frontmatter(content: &str) -> Option<(String, String, bool)> {
let content = content.trim_start();
if !content.starts_with("---") {
return None;
}
let after_opening = &content[3..];
let end = after_opening.find("\n---")?;
let frontmatter = &after_opening[..end];
let mut name = None;
let mut description = None;
let mut hidden = false;
let lines: Vec<&str> = frontmatter.lines().collect();
let mut i = 0;
while i < lines.len() {
let line = lines[i];
if let Some(val) = line.strip_prefix("name:") {
name = Some(val.trim().to_string());
} else if let Some(val) = line.strip_prefix("description:") {
let mut desc = val.trim().to_string();
// Consume YAML continuation lines (indented with spaces or tab)
while i + 1 < lines.len()
&& (lines[i + 1].starts_with(" ") || lines[i + 1].starts_with('\t'))
{
i += 1;
desc.push(' ');
desc.push_str(lines[i].trim());
}
description = Some(desc);
} else if let Some(val) = line.strip_prefix("hidden:") {
hidden = matches!(val.trim(), "true" | "yes");
}
i += 1;
}
Some((name?, description.unwrap_or_default(), hidden))
}
/// Discover all skills across the given directories.
fn discover_skills(dirs: &[PathBuf]) -> Vec<SkillInfo> {
let mut skills = Vec::new();
for skills_dir in dirs {
let entries = match fs::read_dir(skills_dir) {
Ok(e) => e,
Err(_) => continue,
};
for entry in entries.flatten() {
let path = entry.path();
if !path.is_dir() {
continue;
}
let skill_md = path.join("SKILL.md");
if !skill_md.exists() {
continue;
}
let content = match fs::read_to_string(&skill_md) {
Ok(c) => c,
Err(_) => continue,
};
if let Some((name, description, hidden)) = parse_frontmatter(&content) {
skills.push(SkillInfo {
name,
description,
dir: path,
hidden,
});
}
}
}
skills.sort_by(|a, b| a.name.cmp(&b.name));
skills
}
fn truncate_description(desc: &str, max_len: usize) -> String {
if desc.len() <= max_len {
return desc.to_string();
}
let boundary = desc
.char_indices()
.take_while(|(i, _)| *i <= max_len)
.last()
.map(|(i, _)| i)
.unwrap_or(max_len);
let end = desc[..boundary].rfind(' ').unwrap_or(boundary);
format!("{}...", &desc[..end])
}
/// Read the full SKILL.md content (including frontmatter).
fn read_skill_full(skill_md: &Path) -> Option<String> {
fs::read_to_string(skill_md).ok()
}
/// Collect all supplementary files (references/, templates/) for a skill.
fn collect_supplementary_files(skill_dir: &Path) -> Vec<(String, String)> {
let mut files = Vec::new();
for subdir_name in &["references", "templates"] {
let subdir = skill_dir.join(subdir_name);
if !subdir.is_dir() {
continue;
}
let mut entries: Vec<_> = match fs::read_dir(&subdir) {
Ok(e) => e.flatten().collect(),
Err(_) => continue,
};
entries.sort_by_key(|e| e.file_name());
for entry in entries {
let path = entry.path();
if path.is_file() {
if let Ok(content) = fs::read_to_string(&path) {
let rel = format!(
"{}/{}",
subdir_name,
path.file_name().unwrap_or_default().to_string_lossy()
);
files.push((rel, content));
}
}
}
}
files
}
fn run_list(skills_dirs: &[PathBuf], json_mode: bool) {
let skills: Vec<SkillInfo> = discover_skills(skills_dirs)
.into_iter()
.filter(|s| !s.hidden)
.collect();
if skills.is_empty() {
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({ "success": true, "data": [] })).unwrap_or_default()
);
} else {
println!("No skills found");
}
return;
}
if json_mode {
let items: Vec<serde_json::Value> = skills
.iter()
.map(|s| {
json!({
"name": s.name,
"description": s.description,
})
})
.collect();
println!(
"{}",
serde_json::to_string(&json!({ "success": true, "data": items })).unwrap_or_default()
);
} else {
let max_name = skills.iter().map(|s| s.name.len()).max().unwrap_or(0);
for s in &skills {
println!(
" {:<width$} {}",
s.name,
truncate_description(&s.description, 70),
width = max_name
);
}
}
}
fn run_get(skills_dirs: &[PathBuf], names: &[String], get_all: bool, full: bool, json_mode: bool) {
let all_skills = discover_skills(skills_dirs);
let targets: Vec<&SkillInfo> = if get_all {
all_skills.iter().filter(|s| !s.hidden).collect()
} else {
let mut targets = Vec::new();
for name in names {
if name.starts_with('-') {
eprintln!(
"{} Unknown flag ignored: {}",
color::warning_indicator(),
name
);
continue;
}
match all_skills.iter().find(|s| s.name == *name) {
Some(s) => targets.push(s),
None => {
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": false,
"error": format!("Skill not found: {}", name),
}))
.unwrap_or_default()
);
} else {
eprintln!("{} Skill not found: {}", color::error_indicator(), name);
}
exit(1);
}
}
}
targets
};
if targets.is_empty() {
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": false,
"error": "No skill name provided. Usage: agent-browser skills get <name>",
}))
.unwrap_or_default()
);
} else {
eprintln!(
"{} No skill name provided. Usage: agent-browser skills get <name>",
color::error_indicator()
);
}
exit(1);
}
if json_mode {
let items: Vec<serde_json::Value> = targets
.iter()
.map(|s| {
let skill_md = s.dir.join("SKILL.md");
let content = read_skill_full(&skill_md).unwrap_or_default();
let mut obj = json!({
"name": s.name,
"content": content,
});
if full {
let supplementary = collect_supplementary_files(&s.dir);
if !supplementary.is_empty() {
let files: Vec<serde_json::Value> = supplementary
.iter()
.map(|(path, content)| json!({ "path": path, "content": content }))
.collect();
obj["files"] = json!(files);
}
}
obj
})
.collect();
println!(
"{}",
serde_json::to_string(&json!({ "success": true, "data": items })).unwrap_or_default()
);
} else {
for (i, s) in targets.iter().enumerate() {
if i > 0 {
println!("\n---\n");
}
let skill_md = s.dir.join("SKILL.md");
if let Some(content) = read_skill_full(&skill_md) {
print!("{}", content);
if !content.ends_with('\n') {
println!();
}
}
if full {
let supplementary = collect_supplementary_files(&s.dir);
for (path, content) in &supplementary {
println!("\n--- {} ---\n", path);
print!("{}", content);
if !content.ends_with('\n') {
println!();
}
}
}
}
}
}
fn run_path(skills_dirs: &[PathBuf], name: Option<&str>, json_mode: bool) {
match name {
Some(name) => {
let all_skills = discover_skills(skills_dirs);
match all_skills.iter().find(|s| s.name == name) {
Some(s) => {
let path = s.dir.to_string_lossy().to_string();
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": true,
"data": { "name": s.name, "path": path },
}))
.unwrap_or_default()
);
} else {
println!("{}", path);
}
}
None => {
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": false,
"error": format!("Skill not found: {}", name),
}))
.unwrap_or_default()
);
} else {
eprintln!("{} Skill not found: {}", color::error_indicator(), name);
}
exit(1);
}
}
}
None => {
let paths: Vec<String> = skills_dirs
.iter()
.map(|d| d.to_string_lossy().to_string())
.collect();
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": true,
"data": { "paths": paths },
}))
.unwrap_or_default()
);
} else {
for p in &paths {
println!("{}", p);
}
}
}
}
}
pub fn run_skills(args: &[String], json_mode: bool) {
let skills_dirs = find_skills_dirs();
if skills_dirs.is_empty() {
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": false,
"error": "Skills directory not found. Set AGENT_BROWSER_SKILLS_DIR or reinstall via npm.",
}))
.unwrap_or_default()
);
} else {
eprintln!(
"{} Skills directory not found. Set AGENT_BROWSER_SKILLS_DIR or reinstall via npm.",
color::error_indicator()
);
}
exit(1);
}
let subcommand = args.get(1).map(|s| s.as_str());
match subcommand {
None | Some("list") => run_list(&skills_dirs, json_mode),
Some("get") => {
let names: Vec<String> = args[2..]
.iter()
.filter(|a| *a != "--full" && *a != "--all")
.cloned()
.collect();
let full = args[2..].iter().any(|a| a == "--full");
let get_all = args[2..].iter().any(|a| a == "--all");
run_get(&skills_dirs, &names, get_all, full, json_mode);
}
Some("path") => {
let name = args.get(2).map(|s| s.as_str());
run_path(&skills_dirs, name, json_mode);
}
Some(unknown) => {
if json_mode {
println!(
"{}",
serde_json::to_string(&json!({
"success": false,
"error": format!("Unknown skills subcommand: {}", unknown),
}))
.unwrap_or_default()
);
} else {
eprintln!(
"{} Unknown skills subcommand: {}",
color::error_indicator(),
unknown
);
}
exit(1);
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::fs;
fn create_test_skill(dir: &Path, name: &str, description: &str) {
let skill_dir = dir.join(name);
fs::create_dir_all(&skill_dir).unwrap();
fs::write(
skill_dir.join("SKILL.md"),
format!(
"---\nname: {}\ndescription: {}\n---\n\n# {}\n\nContent here.\n",
name, description, name
),
)
.unwrap();
}
#[test]
fn test_parse_frontmatter_basic() {
let content = "---\nname: test-skill\ndescription: A test skill.\n---\n\n# Test\n";
let (name, desc, hidden) = parse_frontmatter(content).unwrap();
assert_eq!(name, "test-skill");
assert_eq!(desc, "A test skill.");
assert!(!hidden);
}
#[test]
fn test_parse_frontmatter_multiline_description() {
let content =
"---\nname: test\ndescription: First line\n continued here\n and here\n---\n";
let (name, desc, hidden) = parse_frontmatter(content).unwrap();
assert_eq!(name, "test");
assert_eq!(desc, "First line continued here and here");
assert!(!hidden);
}
#[test]
fn test_parse_frontmatter_hidden_true() {
let content = "---\nname: stub\ndescription: A bootstrap stub.\nhidden: true\n---\n";
let (name, desc, hidden) = parse_frontmatter(content).unwrap();
assert_eq!(name, "stub");
assert_eq!(desc, "A bootstrap stub.");
assert!(hidden);
}
#[test]
fn test_parse_frontmatter_hidden_false() {
let content = "---\nname: visible\ndescription: Visible.\nhidden: false\n---\n";
let (_, _, hidden) = parse_frontmatter(content).unwrap();
assert!(!hidden);
}
#[test]
fn test_parse_frontmatter_no_frontmatter() {
let content = "# Just a heading\n\nNo frontmatter here.\n";
assert!(parse_frontmatter(content).is_none());
}
#[test]
fn test_parse_frontmatter_missing_name() {
let content = "---\ndescription: No name field\n---\n";
assert!(parse_frontmatter(content).is_none());
}
#[test]
fn test_discover_skills_single_dir() {
let tmp = tempfile::tempdir().unwrap();
create_test_skill(tmp.path(), "alpha", "Alpha skill");
create_test_skill(tmp.path(), "beta", "Beta skill");
// Non-skill directory (no SKILL.md)
fs::create_dir_all(tmp.path().join("not-a-skill")).unwrap();
fs::write(tmp.path().join("not-a-skill").join("README.md"), "hi").unwrap();
let dirs = vec![tmp.path().to_path_buf()];
let skills = discover_skills(&dirs);
assert_eq!(skills.len(), 2);
assert_eq!(skills[0].name, "alpha");
assert_eq!(skills[1].name, "beta");
}
#[test]
fn test_discover_skills_multiple_dirs() {
let tmp1 = tempfile::tempdir().unwrap();
let tmp2 = tempfile::tempdir().unwrap();
create_test_skill(tmp1.path(), "alpha", "Alpha skill");
create_test_skill(tmp2.path(), "beta", "Beta skill");
create_test_skill(tmp2.path(), "gamma", "Gamma skill");
let dirs = vec![tmp1.path().to_path_buf(), tmp2.path().to_path_buf()];
let skills = discover_skills(&dirs);
assert_eq!(skills.len(), 3);
assert_eq!(skills[0].name, "alpha");
assert_eq!(skills[1].name, "beta");
assert_eq!(skills[2].name, "gamma");
}
#[test]
fn test_truncate_description() {
assert_eq!(truncate_description("short", 10), "short");
assert_eq!(
truncate_description("this is a longer description that should be truncated", 20),
"this is a longer..."
);
}
#[test]
fn test_truncate_description_multibyte() {
let desc = "Browse \u{00e9}l\u{00e9}ments and \u{65e5}\u{672c}\u{8a9e} pages quickly";
let result = truncate_description(desc, 20);
assert!(result.ends_with("..."));
assert!(result.len() <= 30);
}
#[test]
fn test_collect_supplementary_files() {
let tmp = tempfile::tempdir().unwrap();
let refs_dir = tmp.path().join("references");
fs::create_dir_all(&refs_dir).unwrap();
fs::write(refs_dir.join("auth.md"), "# Auth\n").unwrap();
fs::write(refs_dir.join("commands.md"), "# Commands\n").unwrap();
let templates_dir = tmp.path().join("templates");
fs::create_dir_all(&templates_dir).unwrap();
fs::write(templates_dir.join("example.sh"), "#!/bin/bash\n").unwrap();
let files = collect_supplementary_files(tmp.path());
assert_eq!(files.len(), 3);
assert_eq!(files[0].0, "references/auth.md");
assert_eq!(files[1].0, "references/commands.md");
assert_eq!(files[2].0, "templates/example.sh");
}
}
+141
View File
@@ -0,0 +1,141 @@
//! Integration tests for `agent-browser doctor`.
//!
//! These tests spawn the real CLI binary via `env!("CARGO_BIN_EXE_*")` and
//! verify the doctor command produces sane output. They override
//! `AGENT_BROWSER_SOCKET_DIR` and `HOME` / `USERPROFILE` so the doctor
//! inspects a throwaway directory and never touches the user's real state.
use std::process::Command;
use tempfile::TempDir;
const BIN: &str = env!("CARGO_BIN_EXE_agent-browser");
fn build_doctor_cmd(tmp: &TempDir, args: &[&str]) -> Command {
let socket_dir = tmp.path().join("sockets");
let home = tmp.path().join("home");
std::fs::create_dir_all(&socket_dir).unwrap();
std::fs::create_dir_all(&home).unwrap();
let mut cmd = Command::new(BIN);
cmd.args(args)
.env("AGENT_BROWSER_SOCKET_DIR", &socket_dir)
.env("HOME", &home)
.env("USERPROFILE", &home)
// Keep the launch test's skip-logic deterministic across hosts.
.env_remove("AGENT_BROWSER_PROVIDER")
.env_remove("AGENT_BROWSER_CDP")
// Don't emit color codes into captured stdout.
.env("NO_COLOR", "1");
cmd
}
#[test]
fn doctor_offline_quick_json_emits_valid_payload() {
let tmp = TempDir::new().unwrap();
let output = build_doctor_cmd(&tmp, &["doctor", "--offline", "--quick", "--json"])
.output()
.expect("failed to invoke agent-browser doctor");
let code = output.status.code().unwrap_or(-1);
let stdout = String::from_utf8(output.stdout).expect("stdout should be utf8");
let stderr = String::from_utf8_lossy(&output.stderr).into_owned();
// Exit code 0 (all pass) or 1 (one or more fails) are both valid outcomes;
// the doctor may legitimately report a failure on a host without Chrome.
assert!(
code == 0 || code == 1,
"unexpected exit code {}\nstdout:\n{}\nstderr:\n{}",
code,
stdout,
stderr,
);
let payload: serde_json::Value = serde_json::from_str(&stdout)
.unwrap_or_else(|e| panic!("stdout was not JSON: {}\n---\n{}", e, stdout));
assert!(payload.get("success").is_some(), "missing success field");
assert!(payload.get("summary").is_some(), "missing summary field");
assert!(payload.get("fixed").is_some(), "missing fixed field");
let summary = &payload["summary"];
assert!(summary["pass"].is_number());
assert!(summary["warn"].is_number());
assert!(summary["fail"].is_number());
let checks = payload["checks"]
.as_array()
.expect("checks should be an array");
assert!(!checks.is_empty(), "checks array should not be empty");
// Every check must have a non-empty id / category / status / message.
for c in checks {
assert!(
c["id"].as_str().is_some_and(|s| !s.is_empty()),
"check missing id: {}",
c
);
assert!(
c["category"].as_str().is_some_and(|s| !s.is_empty()),
"check missing category: {}",
c
);
let status = c["status"].as_str().expect("status should be string");
assert!(
["pass", "warn", "fail", "info"].contains(&status),
"unexpected status {:?}",
status
);
assert!(
c["message"].as_str().is_some_and(|s| !s.is_empty()),
"check missing message: {}",
c
);
}
// Check IDs must be unique now that providers / sessions / skipped-launch
// states each carry their own ID suffix.
let mut seen = std::collections::HashSet::new();
for c in checks {
let id = c["id"].as_str().unwrap();
assert!(
seen.insert(id.to_string()),
"duplicate check id in JSON output: {}\nfull payload:\n{}",
id,
stdout
);
}
}
#[test]
fn doctor_help_describes_flags_and_examples() {
let tmp = TempDir::new().unwrap();
let output = build_doctor_cmd(&tmp, &["doctor", "--help"])
.output()
.expect("failed to invoke agent-browser doctor --help");
assert!(
output.status.success(),
"doctor --help should exit 0; got {:?}",
output.status
);
let stdout = String::from_utf8(output.stdout).expect("stdout should be utf8");
for needle in [
"agent-browser doctor",
"--offline",
"--quick",
"--fix",
"--json",
"Exit codes",
] {
assert!(
stdout.contains(needle),
"doctor --help output missing {:?}\n---\n{}",
needle,
stdout
);
}
}
+1 -1
View File
@@ -1,5 +1,5 @@
# Multi-platform Rust cross-compilation image
FROM rust:1.85-bookworm
FROM rust:1.94-bookworm
# Install cross-compilation toolchains
RUN apt-get update && apt-get install -y \
+25 -8
View File
@@ -20,13 +20,19 @@ services:
# Build both targets in parallel
(echo "→ Linux x64" && cargo zigbuild --release --target x86_64-unknown-linux-gnu && cp /build/target/x86_64-unknown-linux-gnu/release/agent-browser /output/agent-browser-linux-x64 && chmod +x /output/agent-browser-linux-x64 && echo "✓ Linux x64 done") &
PID1=$!
PID1=$$!
(echo "→ Linux ARM64" && cargo zigbuild --release --target aarch64-unknown-linux-gnu && cp /build/target/aarch64-unknown-linux-gnu/release/agent-browser /output/agent-browser-linux-arm64 && chmod +x /output/agent-browser-linux-arm64 && echo "✓ Linux ARM64 done") &
PID2=$!
PID2=$$!
# Wait for both to complete
wait $PID1 $PID2
# Wait for both and check exit codes individually — without this
# the outer script exits 0 even if one of the parallel builds
# failed, silently leaving a stale binary in /output from the
# previous release. Caused 0.27.0-fork.5 to ship with a stale
# linux-x64 binary at the first publish attempt until caught
# manually by checking the embedded version string.
wait $$PID1 || { echo "✗ Linux x64 build failed"; exit 1; }
wait $$PID2 || { echo "✗ Linux ARM64 build failed"; exit 1; }
echo ""
echo "✓ Linux platforms built successfully!"
@@ -65,10 +71,21 @@ services:
environment:
- TARGET=${TARGET:-x86_64-unknown-linux-gnu}
- OUTPUT_NAME=${OUTPUT_NAME:-agent-browser-linux-x64}
# NOTE: $$ escapes a literal $ for the in-container shell. A single $ is
# interpolated by docker compose at YAML parse time against the *host*
# environment, which silently drops script-local variables like SRC
# (caused 0.27.0-fork.7 to ship with a stale linux-arm64 binary because
# the cp command resolved to `cp "" "/output/"` after compose ate $SRC
# and $OUTPUT_NAME). $TARGET / $OUTPUT_NAME are set via `environment:`
# below — those are also passed into the container, so $$TARGET and
# $$OUTPUT_NAME read them at script time.
command: |
-c '
cargo zigbuild --release --target $TARGET
cp /build/target/$TARGET/release/agent-browser* /output/$OUTPUT_NAME
chmod +x /output/$OUTPUT_NAME 2>/dev/null || true
echo "✓ Built $OUTPUT_NAME"
set -e
cargo zigbuild --release --target $$TARGET
SRC="/build/target/$$TARGET/release/agent-browser"
if [ -f "$$SRC.exe" ]; then SRC="$$SRC.exe"; fi
cp "$$SRC" "/output/$$OUTPUT_NAME"
chmod +x /output/$$OUTPUT_NAME 2>/dev/null || true
echo "✓ Built $$OUTPUT_NAME"
'
-41
View File
@@ -1,41 +0,0 @@
# See https://help.github.com/articles/ignoring-files/ for more about ignoring files.
# dependencies
/node_modules
/.pnp
.pnp.*
.yarn/*
!.yarn/patches
!.yarn/plugins
!.yarn/releases
!.yarn/versions
# testing
/coverage
# next.js
/.next/
/out/
# production
/build
# misc
.DS_Store
*.pem
# debug
npm-debug.log*
yarn-debug.log*
yarn-error.log*
.pnpm-debug.log*
# env files (can opt-in for committing if needed)
.env*
# vercel
.vercel
# typescript
*.tsbuildinfo
next-env.d.ts
-22
View File
@@ -1,22 +0,0 @@
{
"$schema": "https://ui.shadcn.com/schema.json",
"style": "new-york",
"rsc": true,
"tsx": true,
"tailwind": {
"config": "",
"css": "src/app/globals.css",
"baseColor": "neutral",
"cssVariables": true,
"prefix": ""
},
"iconLibrary": "lucide",
"aliases": {
"components": "@/components",
"utils": "@/lib/utils",
"ui": "@/components/ui",
"lib": "@/lib",
"hooks": "@/hooks"
},
"registries": {}
}
-18
View File
@@ -1,18 +0,0 @@
import { defineConfig, globalIgnores } from "eslint/config";
import nextVitals from "eslint-config-next/core-web-vitals";
import nextTs from "eslint-config-next/typescript";
const eslintConfig = defineConfig([
...nextVitals,
...nextTs,
// Override default ignores of eslint-config-next.
globalIgnores([
// Default ignores of eslint-config-next:
".next/**",
"out/**",
"build/**",
"next-env.d.ts",
]),
]);
export default eslintConfig;
-83
View File
@@ -1,83 +0,0 @@
import type { MDXComponents } from "mdx/types";
import Link from "next/link";
import { CodeBlock } from "@/components/code-block";
function slugify(text: string): string {
return text
.toLowerCase()
.replace(/[^\w\s-]/g, "")
.replace(/\s+/g, "-")
.trim();
}
function extractText(children: React.ReactNode): string {
if (typeof children === "string") return children;
if (typeof children === "number") return String(children);
if (Array.isArray(children)) return children.map(extractText).join("");
if (children && typeof children === "object") {
const obj = children as unknown as Record<string, unknown>;
if ("props" in obj) {
const props = obj.props as { children?: React.ReactNode } | undefined;
return extractText(props?.children);
}
}
return "";
}
export function useMDXComponents(components: MDXComponents): MDXComponents {
return {
...components,
h2: ({ children }: { children?: React.ReactNode }) => {
const id = slugify(extractText(children));
return <h2 id={id}>{children}</h2>;
},
h3: ({ children }: { children?: React.ReactNode }) => {
const id = slugify(extractText(children));
return <h3 id={id}>{children}</h3>;
},
a: ({
href,
children,
}: {
href?: string;
children?: React.ReactNode;
}) => {
if (href?.startsWith("/")) {
return <Link href={href}>{children}</Link>;
}
return (
<a href={href} target="_blank" rel="noopener noreferrer">
{children}
</a>
);
},
code: ({
children,
className,
}: {
children?: React.ReactNode;
className?: string;
}) => {
if (className) {
return <code className={className}>{children}</code>;
}
return <code>{children}</code>;
},
pre: async ({ children }: { children?: React.ReactNode }) => {
const codeElement = children as React.ReactElement<{
className?: string;
children?: string;
}>;
const className = codeElement?.props?.className || "";
const lang = className.replace("language-", "") || "bash";
const code = codeElement?.props?.children || "";
return (
<CodeBlock
code={typeof code === "string" ? code : String(code)}
lang={lang}
/>
);
},
};
}
-11
View File
@@ -1,11 +0,0 @@
import createMDX from "@next/mdx";
/** @type {import('next').NextConfig} */
const nextConfig = {
pageExtensions: ["js", "jsx", "ts", "tsx", "md", "mdx"],
serverExternalPackages: ["just-bash", "bash-tool"],
};
const withMDX = createMDX({});
export default withMDX(nextConfig);
-48
View File
@@ -1,48 +0,0 @@
{
"name": "docs",
"version": "0.1.0",
"private": true,
"scripts": {
"dev": "portless agent-browser next dev",
"build": "next build",
"start": "next start",
"lint": "eslint"
},
"dependencies": {
"@ai-sdk/react": "^3.0.80",
"@mdx-js/loader": "^3.1.1",
"@mdx-js/mdx": "^3.1.1",
"@mdx-js/react": "^3.1.1",
"@next/mdx": "^16.1.6",
"@streamdown/code": "^1.0.2",
"@upstash/ratelimit": "^2.0.8",
"@upstash/redis": "^1.36.2",
"@vercel/analytics": "^1.6.1",
"@vercel/speed-insights": "^1.3.1",
"ai": "^6.0.78",
"bash-tool": "^1.3.14",
"clsx": "^2.1.1",
"geist": "^1.7.0",
"just-bash": "^2.9.6",
"next": "16.1.1",
"next-themes": "^0.4.6",
"radix-ui": "^1.4.3",
"react": "19.2.3",
"react-dom": "19.2.3",
"shiki": "^3.21.0",
"streamdown": "^2.1.0",
"tailwind-merge": "^3.4.0"
},
"devDependencies": {
"@tailwindcss/postcss": "^4",
"@types/mdx": "^2.0.13",
"@types/node": "^20",
"@types/react": "^19",
"@types/react-dom": "^19",
"eslint": "^9",
"eslint-config-next": "16.1.1",
"tailwindcss": "^4",
"tailwindcss-animate": "^1.0.7",
"typescript": "^5"
}
}
-8139
View File
File diff suppressed because it is too large Load Diff
-7
View File
@@ -1,7 +0,0 @@
const config = {
plugins: {
"@tailwindcss/postcss": {},
},
};
export default config;
Binary file not shown.
Binary file not shown.
-119
View File
@@ -1,119 +0,0 @@
import { readFile } from "fs/promises";
import { join } from "path";
import { convertToModelMessages, stepCountIs, streamText } from "ai";
import type { ModelMessage, UIMessage } from "ai";
import { createBashTool } from "bash-tool";
import { headers } from "next/headers";
import { allDocsPages } from "@/lib/docs-navigation";
import { mdxToCleanMarkdown } from "@/lib/mdx-to-markdown";
import { minuteRateLimit, dailyRateLimit } from "@/lib/rate-limit";
export const maxDuration = 60;
const DEFAULT_MODEL = "anthropic/claude-haiku-4.5";
const SYSTEM_PROMPT = `You are a helpful documentation assistant for agent-browser, a browser automation CLI designed for AI agents.
GitHub repository: https://github.com/vercel-labs/agent-browser
Documentation: https://agent-browser.dev
npm package: agent-browser
You have access to the full agent-browser documentation via the bash and readFile tools. The docs are available as markdown files in the /workspace/ directory.
When answering questions:
- Use the bash tool to list files (ls /workspace/) or search for content (grep -r "keyword" /workspace/)
- Use the readFile tool to read specific documentation pages (e.g. readFile with path "/workspace/index.md")
- Do NOT use bash to write, create, modify, or delete files (no tee, cat >, sed -i, echo >, cp, mv, rm, mkdir, touch, etc.) you are read-only
- Always base your answers on the actual documentation content
- Be concise and accurate
- If the docs don't cover a topic, say so honestly
- Do NOT include source references or file paths in your response
- Do NOT use emojis in your responses`;
async function loadDocsFiles(): Promise<Record<string, string>> {
const files: Record<string, string> = {};
const results = await Promise.allSettled(
allDocsPages.map(async (page) => {
const slug = page.href === "/" ? "" : page.href.replace(/^\//, "");
const filePath = slug
? join(process.cwd(), "src", "app", slug, "page.mdx")
: join(process.cwd(), "src", "app", "page.mdx");
const raw = await readFile(filePath, "utf-8");
const md = mdxToCleanMarkdown(raw);
const fileName = slug ? `/${slug}.md` : "/index.md";
return { fileName, md };
}),
);
for (const result of results) {
if (result.status === "fulfilled") {
files[result.value.fileName] = result.value.md;
}
}
return files;
}
function addCacheControl(messages: ModelMessage[]): ModelMessage[] {
if (messages.length === 0) return messages;
return messages.map((message, index) => {
if (index === messages.length - 1) {
return {
...message,
providerOptions: {
...message.providerOptions,
anthropic: { cacheControl: { type: "ephemeral" } },
},
};
}
return message;
});
}
export async function POST(req: Request) {
const headersList = await headers();
const ip = headersList.get("x-forwarded-for")?.split(",")[0] ?? "anonymous";
const [minuteResult, dailyResult] = await Promise.all([
minuteRateLimit.limit(ip),
dailyRateLimit.limit(ip),
]);
if (!minuteResult.success || !dailyResult.success) {
const isMinuteLimit = !minuteResult.success;
return new Response(
JSON.stringify({
error: "Rate limit exceeded",
message: isMinuteLimit
? "Too many requests. Please wait a moment before trying again."
: "Daily limit reached. Please try again tomorrow.",
}),
{
status: 429,
headers: { "Content-Type": "application/json" },
},
);
}
const { messages }: { messages: UIMessage[] } = await req.json();
const docsFiles = await loadDocsFiles();
const {
tools: { bash, readFile },
} = await createBashTool({ files: docsFiles });
const result = streamText({
model: DEFAULT_MODEL,
system: SYSTEM_PROMPT,
messages: await convertToModelMessages(messages),
stopWhen: stepCountIs(5),
tools: { bash, readFile },
prepareStep: ({ messages: stepMessages }) => ({
messages: addCacheControl(stepMessages),
}),
});
return result.toUIMessageStreamResponse();
}
-40
View File
@@ -1,40 +0,0 @@
import { readFile } from "fs/promises";
import { join } from "path";
import { NextRequest, NextResponse } from "next/server";
import { mdxToCleanMarkdown } from "@/lib/mdx-to-markdown";
export async function GET(req: NextRequest) {
const { searchParams } = new URL(req.url);
const docPath = searchParams.get("path");
if (!docPath) {
return NextResponse.json(
{ error: "Missing ?path= parameter" },
{ status: 400 },
);
}
const normalized = docPath
.replace(/^\//, "")
.replace(/\.\./g, "")
.replace(/[^a-zA-Z0-9/_-]/g, "");
const slug = normalized;
const filePath = slug
? join(process.cwd(), "src", "app", ...slug.split("/"), "page.mdx")
: join(process.cwd(), "src", "app", "page.mdx");
try {
const raw = await readFile(filePath, "utf-8");
const markdown = mdxToCleanMarkdown(raw);
return new NextResponse(markdown, {
headers: {
"Content-Type": "text/markdown; charset=utf-8",
"Cache-Control": "public, max-age=3600",
},
});
} catch {
return NextResponse.json({ error: "Page not found" }, { status: 404 });
}
}
-69
View File
@@ -1,69 +0,0 @@
import { NextRequest, NextResponse } from "next/server";
import { getSearchIndex } from "@/lib/search-index";
export async function GET(req: NextRequest) {
const q = req.nextUrl.searchParams.get("q")?.trim().toLowerCase();
if (!q) {
return NextResponse.json({ results: [] });
}
const index = await getSearchIndex();
const terms = q.split(/\s+/).filter(Boolean);
const results = index
.map((entry) => {
const titleLower = entry.title.toLowerCase();
const contentLower = entry.content.toLowerCase();
const titleMatch = terms.every((t) => titleLower.includes(t));
const contentMatch = terms.every((t) => contentLower.includes(t));
if (!titleMatch && !contentMatch) return null;
let snippet = "";
if (contentMatch) {
const firstTermIdx = Math.min(
...terms.map((t) => {
const idx = contentLower.indexOf(t);
return idx === -1 ? Infinity : idx;
}),
);
if (firstTermIdx !== Infinity) {
const start = Math.max(0, firstTermIdx - 40);
const end = Math.min(entry.content.length, firstTermIdx + 120);
snippet =
(start > 0 ? "..." : "") +
entry.content.slice(start, end).replace(/\n/g, " ") +
(end < entry.content.length ? "..." : "");
}
}
return {
title: entry.title,
href: entry.href,
section: entry.section,
snippet,
score: titleMatch ? 2 : 1,
};
})
.filter(
(
r,
): r is {
title: string;
href: string;
section: string;
snippet: string;
score: number;
} => r !== null,
)
.sort((a, b) => b.score - a.score)
.slice(0, 20)
.map(({ score: _, ...rest }) => rest);
return NextResponse.json(
{ results },
{ headers: { "Cache-Control": "public, max-age=60" } },
);
}
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("cdp-mode");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
-120
View File
@@ -1,120 +0,0 @@
# CDP Mode
Connect to an existing browser via Chrome DevTools Protocol:
```bash
# Start Chrome with: google-chrome --remote-debugging-port=9222
# Connect once, then run commands without --cdp
agent-browser connect 9222
agent-browser snapshot
agent-browser tab
agent-browser close
# Or pass --cdp on each command
agent-browser --cdp 9222 snapshot
```
## Remote WebSocket URLs
Connect to remote browser services via WebSocket URL:
```bash
# Connect to remote browser service
agent-browser --cdp "wss://browser-service.com/cdp?token=..." snapshot
# Works with any CDP-compatible service
agent-browser --cdp "ws://localhost:9222/devtools/browser/abc123" open example.com
```
The `--cdp` flag accepts either:
- A port number (e.g., `9222`) for local connections via `http://localhost:{port}`
- A full WebSocket URL (e.g., `wss://...` or `ws://...`) for remote browser services
## Auto-Connect
Use `--auto-connect` to automatically discover and connect to a running Chrome instance without specifying a port:
```bash
# Auto-discover running Chrome with remote debugging
agent-browser --auto-connect open example.com
agent-browser --auto-connect snapshot
# Or via environment variable
AGENT_BROWSER_AUTO_CONNECT=1 agent-browser snapshot
```
Auto-connect discovers Chrome by:
1. Reading Chrome's `DevToolsActivePort` file from the default user data directory
2. Falling back to probing common debugging ports (9222, 9229)
3. If HTTP-based discovery (`/json/version`, `/json/list`) fails, falling back to a direct WebSocket connection
This is useful when:
- Chrome 144+ has remote debugging enabled via `chrome://inspect/#remote-debugging` (which uses a dynamic port)
- You want a zero-configuration connection to your existing browser
- You don't want to track which port Chrome is using
## Color scheme
Use `--color-scheme` to set a persistent preference when connecting via CDP:
```bash
agent-browser --cdp 9222 --color-scheme dark open https://example.com
agent-browser --cdp 9222 snapshot # stays in dark mode
```
Or set it globally via config or environment variable:
```bash
AGENT_BROWSER_COLOR_SCHEME=dark agent-browser --cdp 9222 open https://example.com
```
## Use cases
This enables control of:
- Electron apps
- Chrome/Chromium with remote debugging
- WebView2 applications
- Remote browser services (via WebSocket URL)
- Any browser exposing a CDP endpoint
## Global options
<table>
<thead>
<tr><th>Option</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td><code>--session &lt;name&gt;</code></td><td>Use isolated session</td></tr>
<tr><td><code>--profile &lt;path&gt;</code></td><td>Persistent browser profile directory</td></tr>
<tr><td><code>-p &lt;provider&gt;</code></td><td>Cloud browser provider (<code>browserbase</code>, <code>browseruse</code>, <code>kernel</code>, <code>browserless</code>)</td></tr>
<tr><td><code>--headers &lt;json&gt;</code></td><td>HTTP headers scoped to origin</td></tr>
<tr><td><code>--executable-path</code></td><td>Custom browser executable</td></tr>
<tr><td><code>--args &lt;args&gt;</code></td><td>Browser launch args (comma-separated)</td></tr>
<tr><td><code>--user-agent &lt;ua&gt;</code></td><td>Custom User-Agent string</td></tr>
<tr><td><code>--proxy &lt;url&gt;</code></td><td>Proxy server URL</td></tr>
<tr><td><code>--proxy-bypass &lt;hosts&gt;</code></td><td>Hosts to bypass proxy</td></tr>
<tr><td><code>--json</code></td><td>JSON output for scripts</td></tr>
<tr><td><code>--name, -n</code></td><td>Locator name filter</td></tr>
<tr><td><code>--exact</code></td><td>Exact text match</td></tr>
<tr><td><code>--headed</code></td><td>Show browser window</td></tr>
<tr><td><code>{"--cdp <port|url>"}</code></td><td>CDP connection (port or WebSocket URL)</td></tr>
<tr><td><code>--auto-connect</code></td><td>Auto-discover and connect to running Chrome</td></tr>
<tr><td><code>--color-scheme &lt;scheme&gt;</code></td><td>Persistent color scheme (<code>dark</code>, <code>light</code>, <code>no-preference</code>)</td></tr>
<tr><td><code>--debug</code></td><td>Debug output</td></tr>
</tbody>
</table>
## Cloud providers
Use the `-p` flag to connect to a cloud browser provider instead of launching a local browser:
```bash
agent-browser -p browserbase open https://example.com
```
See the [Providers](/providers/browser-use) section for setup and configuration of each supported provider: [Browser Use](/providers/browser-use), [Browserbase](/providers/browserbase), [Browserless](/providers/browserless), and [Kernel](/providers/kernel).
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("changelog");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
File diff suppressed because it is too large Load Diff
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("commands");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
-435
View File
@@ -1,435 +0,0 @@
# Commands
## Core
```bash
agent-browser open <url> # Navigate (aliases: goto, navigate)
agent-browser click <sel> # Click element (--new-tab to open in new tab)
agent-browser dblclick <sel> # Double-click
agent-browser fill <sel> <text> # Clear and fill
agent-browser type <sel> <text> # Type into element
agent-browser press <key> # Press key (Enter, Tab, Control+a) (alias: key)
agent-browser keyboard type <text> # Type at current focus (no selector needed)
agent-browser keyboard inserttext <text> # Insert text without key events
agent-browser keydown <key> # Hold key down
agent-browser keyup <key> # Release key
agent-browser hover <sel> # Hover element
agent-browser focus <sel> # Focus element
agent-browser select <sel> <val> # Select dropdown option
agent-browser check <sel> # Check checkbox
agent-browser uncheck <sel> # Uncheck checkbox
agent-browser scroll <dir> [px] # Scroll (up/down/left/right, --selector <sel>)
agent-browser scrollintoview <sel> # Scroll element into view
agent-browser drag <src> <dst> # Drag and drop
agent-browser upload <sel> <files> # Upload files
agent-browser screenshot [path] # Screenshot (--full for full page)
agent-browser screenshot --annotate # Annotated screenshot with numbered element labels
agent-browser screenshot --screenshot-dir ./shots # Save to custom directory
agent-browser screenshot --screenshot-format jpeg --screenshot-quality 80
agent-browser pdf <path> # Save page as PDF
agent-browser snapshot # Accessibility tree with refs
agent-browser eval <js> # Run JavaScript
agent-browser connect <port|url> # Connect to browser via CDP
agent-browser stream enable [--port <port>] # Start runtime WebSocket streaming
agent-browser stream status # Show runtime streaming state and bound port
agent-browser stream disable # Stop runtime WebSocket streaming
agent-browser close # Close browser (aliases: quit, exit)
agent-browser close --all # Close all active sessions
```
## Get info
```bash
agent-browser get text <sel> # Get text content
agent-browser get html <sel> # Get innerHTML
agent-browser get value <sel> # Get input value
agent-browser get attr <sel> <attr> # Get attribute
agent-browser get title # Get page title
agent-browser get url # Get current URL
agent-browser get cdp-url # Get CDP WebSocket URL
agent-browser get count <sel> # Count matching elements
agent-browser get box <sel> # Get bounding box
agent-browser get styles <sel> # Get computed styles
```
## Check state
```bash
agent-browser is visible <sel> # Check if visible
agent-browser is enabled <sel> # Check if enabled
agent-browser is checked <sel> # Check if checked
```
## Find elements
Semantic locators with actions (`click`, `fill`, `type`, `hover`, `focus`, `check`, `uncheck`, `text`):
```bash
agent-browser find role <role> <action> [value]
agent-browser find text <text> <action>
agent-browser find label <label> <action> [value]
agent-browser find placeholder <ph> <action> [value]
agent-browser find alt <text> <action>
agent-browser find title <text> <action>
agent-browser find testid <id> <action> [value]
agent-browser find first <sel> <action> [value]
agent-browser find last <sel> <action> [value]
agent-browser find nth <n> <sel> <action> [value]
```
Options:
- `--name <name>` -- filter role by accessible name
- `--exact` -- require exact text match
Examples:
```bash
agent-browser find role button click --name "Submit"
agent-browser find label "Email" fill "test@test.com"
agent-browser find alt "Logo" click
agent-browser find first ".item" click
agent-browser find last ".item" text
agent-browser find nth 2 ".card" hover
```
## Wait
```bash
agent-browser wait <selector> # Wait for element
agent-browser wait <ms> # Wait for time
agent-browser wait --text "Welcome" # Wait for text (substring match)
agent-browser wait --url "**/dash" # Wait for URL pattern
agent-browser wait --load networkidle # Wait for load state
agent-browser wait --fn "condition" # Wait for JS condition
agent-browser wait --download [path] # Wait for download
agent-browser wait --fn "!document.body.innerText.includes('Loading...')" # Wait for text to disappear
agent-browser wait "#spinner" --state hidden # Wait for element to disappear
```
## Downloads
```bash
agent-browser download <sel> <path> # Click element to trigger download
agent-browser wait --download [path] # Wait for any download to complete
```
Use `--download-path <dir>` (or `AGENT_BROWSER_DOWNLOAD_PATH` env) to set a default download directory. Without it, downloads go to a temporary directory that is deleted when the browser closes.
## Mouse
```bash
agent-browser mouse move <x> <y> # Move mouse
agent-browser mouse down [button] # Press button
agent-browser mouse up [button] # Release button
agent-browser mouse wheel <dy> [dx] # Scroll wheel
```
## Clipboard
```bash
agent-browser clipboard read # Read text from clipboard
agent-browser clipboard write "Hello, World!" # Write text to clipboard
agent-browser clipboard copy # Copy current selection (Ctrl+C)
agent-browser clipboard paste # Paste from clipboard (Ctrl+V)
```
## Settings
```bash
agent-browser set viewport <w> <h> [scale] # Set viewport size (scale for retina, e.g. 2)
agent-browser set device <name> # Emulate device ("iPhone 14")
agent-browser set geo <lat> <lng> # Set geolocation
agent-browser set offline [on|off] # Toggle offline mode
agent-browser set headers <json> # Extra HTTP headers
agent-browser set credentials <u> <p> # HTTP basic auth
agent-browser set media [dark|light] # Emulate color scheme (persists for session)
```
Use `--color-scheme` for persistent dark/light mode across all commands:
```bash
agent-browser --color-scheme dark open https://example.com
```
## Cookies & storage
```bash
agent-browser cookies # Get all cookies
agent-browser cookies set <name> <val> # Set cookie
agent-browser cookies clear # Clear cookies
agent-browser storage local # Get all localStorage
agent-browser storage local <key> # Get specific key
agent-browser storage local set <k> <v> # Set value
agent-browser storage local clear # Clear all
agent-browser storage session # Same for sessionStorage
```
## Network
```bash
agent-browser network route <url> # Intercept requests
agent-browser network route <url> --abort # Block requests
agent-browser network route <url> --body <json> # Mock response
agent-browser network unroute [url] # Remove routes
agent-browser network requests # View tracked requests
agent-browser network requests --clear # Clear request log
agent-browser network requests --filter <pat> # Filter by URL pattern
agent-browser network requests --type xhr,fetch # Filter by resource type
agent-browser network requests --method POST # Filter by HTTP method
agent-browser network requests --status 2xx # Filter by status (200, 2xx, 400-499)
agent-browser network request <requestId> # View full request/response detail
agent-browser network har start # Start HAR recording
agent-browser network har stop [output.har] # Stop and save HAR (temp path if omitted)
```
## Tabs & frames
```bash
agent-browser tab # List tabs
agent-browser tab new [url] # New tab
agent-browser tab <n> # Switch to tab
agent-browser tab close [n] # Close tab
agent-browser window new # Open new browser window
agent-browser frame <sel> # Switch to iframe by CSS selector
agent-browser frame @e3 # Switch to iframe by element ref
agent-browser frame main # Back to main frame
```
### Iframe support
Iframes are detected automatically during snapshots. `Iframe` nodes are resolved and their content is
inlined beneath the iframe element in the snapshot output. Refs assigned to elements inside iframes carry
frame context, so `click`, `fill`, and other interactions work without manually switching frames.
```bash
agent-browser snapshot -i
# @e3 [Iframe] "payment-frame"
# @e4 [input] "Card number"
# @e5 [button] "Pay"
# Interact directly using refs — no frame switch needed
agent-browser fill @e4 "4111111111111111"
agent-browser click @e5
# Or switch frame context for scoped snapshots
agent-browser frame @e3
agent-browser snapshot -i # Only elements inside that iframe
agent-browser frame main # Return to main frame
```
The `frame` command accepts element refs (`@e3`), CSS selectors (`"#my-iframe"`), or frame name/URL.
## Dialogs
```bash
agent-browser dialog accept [text] # Accept dialog (with optional prompt text)
agent-browser dialog dismiss # Dismiss dialog
agent-browser dialog status # Check if a dialog is currently open
```
By default, `alert` and `beforeunload` dialogs are automatically accepted so they never block the agent. `confirm` and `prompt` dialogs still require explicit handling. Use `--no-auto-dialog` (or `AGENT_BROWSER_NO_AUTO_DIALOG=1`) to disable automatic handling.
When a JavaScript dialog (`alert`, `confirm`, `prompt`) is pending, all command responses include a `warning` field with the dialog type and message.
## Streaming
```bash
agent-browser stream enable # Start runtime WebSocket streaming on an auto-selected port
agent-browser stream enable --port 9223 # Bind a specific localhost port
agent-browser stream status # Show enabled state, port, browser connection, screencasting
agent-browser stream disable # Stop runtime streaming and remove the .stream metadata file
```
Streaming is enabled automatically for all sessions. Use these commands to check status, re-enable on a specific port, or disable streaming.
## Debug
```bash
agent-browser trace start [path] # Start trace
agent-browser trace stop [path] # Stop and save trace
agent-browser profiler start # Start Chrome DevTools profiling
agent-browser profiler stop [path] # Stop and save profile (.json)
agent-browser record start <path> # Start video recording (WebM)
agent-browser record stop # Stop and save video
agent-browser record restart <path> # Stop current and start new recording
agent-browser console # View console messages
agent-browser console --json # JSON output with raw CDP args
agent-browser console --clear # Clear console log
agent-browser errors # View page errors
agent-browser errors --clear # Clear error log
agent-browser highlight <sel> # Highlight element
agent-browser inspect # Open Chrome DevTools for the active page
```
## Auth vault
```bash
agent-browser auth save <name> [opts] # Save auth profile
agent-browser auth login <name> # Login using saved credentials
agent-browser auth list # List saved profiles (names and URLs only)
agent-browser auth show <name> # Show profile metadata (no passwords)
agent-browser auth delete <name> # Delete a saved profile
```
Save options:
- `--url <url>` -- login page URL (required)
- `--username <user>` -- username (required)
- `--password <pass>` -- password (required unless `--password-stdin`)
- `--password-stdin` -- read password from stdin (recommended to avoid shell history exposure)
- `--username-selector <sel>` -- custom CSS selector for username field
- `--password-selector <sel>` -- custom CSS selector for password field
- `--submit-selector <sel>` -- custom CSS selector for submit button
`auth login` navigates with `load` and then waits for the username/password/submit selectors to appear before interacting. This improves reliability on SPA login pages where fields render after initial page load.
```bash
echo "pass" | agent-browser auth save github --url https://github.com/login --username user --password-stdin
agent-browser auth login github
agent-browser auth list
```
## Confirmation
When `--confirm-actions` is set, certain action categories return a `confirmation_required` response instead of executing immediately. Use `confirm` or `deny` to approve or reject the action.
```bash
agent-browser confirm <confirmation-id> # Approve a pending action
agent-browser deny <confirmation-id> # Deny a pending action
```
Pending confirmations auto-deny after 60 seconds.
```bash
agent-browser --confirm-actions eval,download eval "document.title"
# Returns confirmation_required with ID
agent-browser confirm c_8f3a1234
```
## State management
```bash
agent-browser state save <path> # Save auth state to file
agent-browser state load <path> # Load auth state from file
agent-browser state list # List saved state files
agent-browser state show <file> # Show state summary
agent-browser state rename <old> <new> # Rename state file
agent-browser state clear [name] # Clear states for session name
agent-browser state clear --all # Clear all saved states
agent-browser state clean --older-than <days> # Delete old states
```
## Sessions
```bash
agent-browser session # Show current session name
agent-browser session list # List active sessions
```
## Dashboard
```bash
agent-browser dashboard [start] # Start the dashboard server (default port: 4848)
agent-browser dashboard start --port <n> # Start on a specific port
agent-browser dashboard stop # Stop the dashboard server
agent-browser dashboard install # Install the dashboard files
```
## Navigation
```bash
agent-browser back # Go back
agent-browser forward # Go forward
agent-browser reload # Reload page
```
## Global options
```bash
--session <name> # Isolated browser session
--session-name <name> # Auto-save/restore session state (cookies, localStorage)
--profile <path> # Persistent browser profile directory
--state <path> # Load storage state from JSON file
--headers <json> # HTTP headers scoped to URL's origin
--executable-path <path> # Custom browser executable
--extension <path> # Load browser extension (repeatable)
--args <args> # Browser launch args (comma separated)
--user-agent <ua> # Custom User-Agent string
--proxy <url> # Proxy server URL
--proxy-bypass <hosts> # Hosts to bypass proxy
--ignore-https-errors # Ignore HTTPS certificate errors
--allow-file-access # Allow file:// URLs to access local files (Chromium only)
-p, --provider <name> # Browser provider (ios, browserbase, kernel, browseruse, browserless)
--device <name> # iOS device name (e.g., "iPhone 15 Pro")
--json # JSON output (for scripts)
--annotate # Annotated screenshot with numbered element labels
--screenshot-dir <path> # Default screenshot output directory (or AGENT_BROWSER_SCREENSHOT_DIR)
--screenshot-quality <n> # JPEG quality 0-100 (or AGENT_BROWSER_SCREENSHOT_QUALITY)
--screenshot-format <fmt> # Format: png (default), jpeg (or AGENT_BROWSER_SCREENSHOT_FORMAT)
--headed # Show browser window (not headless)
--cdp <port|url> # Connect via Chrome DevTools Protocol (port or WebSocket URL)
--auto-connect # Auto-discover and connect to running Chrome
--color-scheme <scheme> # Color scheme: dark, light, no-preference
--download-path <path> # Default download directory
--content-boundaries # Wrap page output in boundary markers for LLM safety
--max-output <chars> # Truncate page output to N characters
--allowed-domains <list> # Comma-separated allowed domain patterns
--action-policy <path> # Path to action policy JSON file
--confirm-actions <list> # Action categories requiring confirmation
--confirm-interactive # Interactive confirmation prompts (auto-denies if stdin is not a TTY)
--config <path> # Use a custom config file
--debug # Debug output
```
## Batch execution
Execute multiple commands in a single invocation by piping a JSON array of string arrays to `batch`:
```bash
echo '[
["open", "https://example.com"],
["snapshot", "-i"],
["click", "@e1"],
["screenshot", "result.png"]
]' | agent-browser batch --json
# Stop on first error
agent-browser batch --bail < commands.json
```
<table>
<thead>
<tr><th>Option</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td><code>--bail</code></td><td>Stop on first error (default: continue all commands)</td></tr>
<tr><td><code>--json</code></td><td>Output results as a JSON array</td></tr>
</tbody>
</table>
## Command chaining
Chain commands with `&&` in a single shell invocation. The browser persists via a background daemon, so chaining works naturally and is more efficient than separate calls:
```bash
agent-browser open example.com && agent-browser wait --load networkidle && agent-browser snapshot -i
agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "pass" && agent-browser click @e3
agent-browser open example.com && agent-browser wait --load networkidle && agent-browser screenshot page.png
```
Use `&&` when you don't need to read intermediate output. Run commands separately when you need to parse output first (e.g., snapshot to discover refs, then interact with those refs).
## Local files
Open local files (PDFs, HTML) using `file://` URLs:
```bash
agent-browser --allow-file-access open file:///path/to/document.pdf
agent-browser --allow-file-access open file:///path/to/page.html
agent-browser screenshot output.png
```
The `--allow-file-access` flag enables JavaScript to access other local files. Chromium only.
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("configuration");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
-200
View File
@@ -1,200 +0,0 @@
# Configuration
Create an `agent-browser.json` file to set persistent defaults instead of repeating flags on every command.
## Config File Locations
agent-browser checks two locations, merged in priority order:
<table>
<thead>
<tr><th>Priority</th><th>Location</th><th>Scope</th></tr>
</thead>
<tbody>
<tr><td>1 (lowest)</td><td><code>~/.agent-browser/config.json</code></td><td>User-level defaults</td></tr>
<tr><td>2</td><td><code>./agent-browser.json</code></td><td>Project-level overrides</td></tr>
<tr><td>3</td><td><code>AGENT_BROWSER_*</code> env vars</td><td>Override config values</td></tr>
<tr><td>4 (highest)</td><td>CLI flags</td><td>Override everything</td></tr>
</tbody>
</table>
Project-level values override user-level values. Environment variables override both. CLI flags always win.
Use `--config <path>` or the `AGENT_BROWSER_CONFIG` environment variable to load a specific config file instead of the default locations:
```bash
agent-browser --config ./ci-config.json open example.com
AGENT_BROWSER_CONFIG=./ci-config.json agent-browser open example.com
```
## Example Config
```json
{
"headed": true,
"proxy": "http://localhost:8080",
"profile": "./browser-data",
"userAgent": "my-agent/1.0",
"ignoreHttpsErrors": true
}
```
## All Options
Every CLI flag can be set in the config file using its camelCase equivalent:
<table>
<thead>
<tr><th>Config Key</th><th>CLI Flag</th><th>Type</th></tr>
</thead>
<tbody>
<tr><td><code>headed</code></td><td><code>--headed</code></td><td>boolean</td></tr>
<tr><td><code>json</code></td><td><code>--json</code></td><td>boolean</td></tr>
<tr><td><code>full</code></td><td><code>--full, -f</code></td><td>boolean</td></tr>
<tr><td><code>debug</code></td><td><code>--debug</code></td><td>boolean</td></tr>
<tr><td><code>session</code></td><td><code>--session</code></td><td>string</td></tr>
<tr><td><code>sessionName</code></td><td><code>--session-name</code></td><td>string</td></tr>
<tr><td><code>executablePath</code></td><td><code>--executable-path</code></td><td>string</td></tr>
<tr><td><code>extensions</code></td><td><code>--extension</code></td><td>string[]</td></tr>
<tr><td><code>profile</code></td><td><code>--profile</code></td><td>string</td></tr>
<tr><td><code>state</code></td><td><code>--state</code></td><td>string</td></tr>
<tr><td><code>proxy</code></td><td><code>--proxy</code></td><td>string</td></tr>
<tr><td><code>proxyBypass</code></td><td><code>--proxy-bypass</code></td><td>string</td></tr>
<tr><td><code>args</code></td><td><code>--args</code></td><td>string</td></tr>
<tr><td><code>userAgent</code></td><td><code>--user-agent</code></td><td>string</td></tr>
<tr><td><code>provider</code></td><td><code>-p, --provider</code></td><td>string</td></tr>
<tr><td><code>device</code></td><td><code>--device</code></td><td>string</td></tr>
<tr><td><code>ignoreHttpsErrors</code></td><td><code>--ignore-https-errors</code></td><td>boolean</td></tr>
<tr><td><code>allowFileAccess</code></td><td><code>--allow-file-access</code></td><td>boolean</td></tr>
<tr><td><code>cdp</code></td><td><code>--cdp</code></td><td>string</td></tr>
<tr><td><code>autoConnect</code></td><td><code>--auto-connect</code></td><td>boolean</td></tr>
<tr><td><code>colorScheme</code></td><td><code>--color-scheme</code></td><td>string (<code>dark</code>, <code>light</code>, <code>no-preference</code>)</td></tr>
<tr><td><code>downloadPath</code></td><td><code>--download-path</code></td><td>string</td></tr>
<tr><td><code>contentBoundaries</code></td><td><code>--content-boundaries</code></td><td>boolean</td></tr>
<tr><td><code>maxOutput</code></td><td><code>--max-output</code></td><td>number</td></tr>
<tr><td><code>allowedDomains</code></td><td><code>--allowed-domains</code></td><td>string[]</td></tr>
<tr><td><code>actionPolicy</code></td><td><code>--action-policy</code></td><td>string</td></tr>
<tr><td><code>confirmActions</code></td><td><code>--confirm-actions</code></td><td>string</td></tr>
<tr><td><code>confirmInteractive</code></td><td><code>--confirm-interactive</code></td><td>boolean</td></tr>
<tr><td><code>engine</code></td><td><code>--engine</code></td><td>string (<code>chrome</code>, <code>lightpanda</code>)</td></tr>
<tr><td><code>noAutoDialog</code></td><td><code>--no-auto-dialog</code></td><td>boolean</td></tr>
<tr><td><code>headers</code></td><td><code>--headers</code></td><td>string (JSON)</td></tr>
</tbody>
</table>
## Common Configurations
### Local Development
```json
{
"headed": true,
"profile": "./browser-data"
}
```
### Behind a Proxy
```json
{
"proxy": "http://proxy.corp.example.com:8080",
"proxyBypass": "localhost,*.internal.com",
"ignoreHttpsErrors": true
}
```
### CI / Devcontainer
```json
{
"args": "--no-sandbox,--disable-gpu",
"ignoreHttpsErrors": true
}
```
### iOS Testing
```json
{
"provider": "ios",
"device": "iPhone 16 Pro"
}
```
### AI Agent Security
```json
{
"contentBoundaries": true,
"maxOutput": 50000,
"allowedDomains": ["your-app.com", "*.your-app.com"],
"actionPolicy": "./policy.json"
}
```
## Overriding Boolean Options
Boolean flags accept an optional `true`/`false` value to override config settings:
```bash
agent-browser --headed false open example.com
```
A bare flag is equivalent to passing `true`:
```bash
agent-browser --headed open example.com # same as --headed true
agent-browser --headed true open example.com # explicit
```
This applies to all boolean flags: `--headed`, `--debug`, `--json`, `--ignore-https-errors`, `--allow-file-access`, `--auto-connect`, `--content-boundaries`, `--confirm-interactive`.
## Extensions Merging
Extensions from user-level and project-level configs are **concatenated**, not replaced. For example, if `~/.agent-browser/config.json` specifies `["/ext1"]` and `./agent-browser.json` specifies `["/ext2"]`, the result is `["/ext1", "/ext2"]`.
The `AGENT_BROWSER_EXTENSIONS` environment variable and CLI `--extension` flags follow the standard priority rules (env replaces config, CLI appends).
## Environment Variables
These environment variables configure additional daemon and runtime behavior:
<table>
<thead>
<tr><th>Variable</th><th>Description</th><th>Default</th></tr>
</thead>
<tbody>
<tr><td><code>AGENT_BROWSER_AUTO_CONNECT</code></td><td>Auto-discover and connect to a running Chrome instance.</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_ALLOW_FILE_ACCESS</code></td><td>Allow <code>file://</code> URLs to access local files.</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_COLOR_SCHEME</code></td><td>Color scheme preference (<code>dark</code>, <code>light</code>, <code>no-preference</code>).</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_DOWNLOAD_PATH</code></td><td>Default directory for browser downloads.</td><td>(temp directory)</td></tr>
<tr><td><code>AGENT_BROWSER_DEFAULT_TIMEOUT</code></td><td>Default timeout in ms. Keep below 30000 to avoid IPC timeouts.</td><td><code>25000</code></td></tr>
<tr><td><code>AGENT_BROWSER_SESSION_NAME</code></td><td>Auto-save/load state persistence name.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_STATE_EXPIRE_DAYS</code></td><td>Auto-delete saved session states older than N days.</td><td><code>30</code></td></tr>
<tr><td><code>AGENT_BROWSER_ENCRYPTION_KEY</code></td><td>64-char hex key for AES-256-GCM session encryption.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_EXTENSIONS</code></td><td>Comma-separated browser extension paths. Extensions work in both headed and headless mode.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_HEADED</code></td><td>Show browser window instead of running headless (<code>1</code> to enable).</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_STREAM_PORT</code></td><td>Override the WebSocket streaming port. By default, an OS-assigned port is used. Set this to bind to a specific port (e.g., <code>9223</code>).</td><td>OS-assigned</td></tr>
<tr><td><code>AGENT_BROWSER_IDLE_TIMEOUT_MS</code></td><td>Auto-shutdown the daemon after N ms of inactivity (no commands received). Useful for ephemeral environments.</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_IOS_DEVICE</code></td><td>Default iOS device name for the <code>ios</code> provider.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_IOS_UDID</code></td><td>Default iOS device UDID for the <code>ios</code> provider.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_DEBUG</code></td><td>Enable debug output (<code>1</code> to enable).</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_CONTENT_BOUNDARIES</code></td><td>Wrap page output in boundary markers for LLM safety.</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_MAX_OUTPUT</code></td><td>Max characters for page output (truncates beyond limit).</td><td>(unlimited)</td></tr>
<tr><td><code>AGENT_BROWSER_ALLOWED_DOMAINS</code></td><td>Comma-separated allowed domain patterns (e.g., <code>example.com,*.example.com</code>).</td><td>(unrestricted)</td></tr>
<tr><td><code>AGENT_BROWSER_ACTION_POLICY</code></td><td>Path to action policy JSON file.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_CONFIRM_ACTIONS</code></td><td>Comma-separated action categories requiring confirmation.</td><td>(none)</td></tr>
<tr><td><code>AGENT_BROWSER_CONFIRM_INTERACTIVE</code></td><td>Enable interactive confirmation prompts (auto-denies if stdin is not a TTY).</td><td>(disabled)</td></tr>
<tr><td><code>AGENT_BROWSER_ENGINE</code></td><td>Browser engine to use: <code>chrome</code> (default), <code>lightpanda</code>.</td><td><code>chrome</code></td></tr>
<tr><td><code>AGENT_BROWSER_NO_AUTO_DIALOG</code></td><td>Disable automatic dismissal of <code>alert</code>/<code>beforeunload</code> dialogs.</td><td>(disabled)</td></tr>
</tbody>
</table>
## Error Handling
- **Auto-discovered config files** (`~/.agent-browser/config.json`, `./agent-browser.json`) that are missing are silently ignored.
- **`--config <path>`** with a missing or malformed file exits with an error.
- **Malformed JSON** in auto-discovered files prints a warning to stderr and continues without that file.
- **Unknown keys** are silently ignored for forward compatibility.
> **Tip:** If your project-level `agent-browser.json` contains environment-specific values (paths, proxies), consider adding it to `.gitignore`.
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("dashboard");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
-137
View File
@@ -1,137 +0,0 @@
# Observability Dashboard
Monitor agent-browser sessions in real time with a local web dashboard showing a live browser viewport and command activity feed.
## Install
Download the dashboard once:
```bash
agent-browser dashboard install
```
This downloads the dashboard to `~/.agent-browser/dashboard/` and is served directly by the daemon when streaming is enabled.
## Usage
Start the dashboard server and open any session -- it appears automatically:
```bash
agent-browser dashboard start
agent-browser open example.com
```
Then open `http://localhost:4848` in your browser to see the live dashboard.
All sessions automatically stream to the dashboard. No extra flags are needed.
### Custom stream port
By default each session binds its WebSocket stream server to an OS-assigned port. To use a specific port, set the `AGENT_BROWSER_STREAM_PORT` environment variable:
```bash
AGENT_BROWSER_STREAM_PORT=9223 agent-browser open example.com
```
You can also use the runtime commands to control streaming on a running session:
```bash
agent-browser stream enable --port 9223
agent-browser stream status
agent-browser stream disable
```
## Dashboard features
The dashboard is a single-page web app with three areas:
<table>
<thead>
<tr>
<th>Area</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Live viewport</strong></td>
<td>Real-time JPEG frames from the browser, rendered to a canvas element</td>
</tr>
<tr>
<td><strong>Activity feed</strong></td>
<td>Chronological stream of commands, results, and console messages with expandable details</td>
</tr>
<tr>
<td><strong>Session creation</strong></td>
<td>Create new sessions from the dashboard with local engines (Chrome, Lightpanda) or cloud providers (AgentCore, Browserbase, Browserless, Browser Use, Kernel)</td>
</tr>
<tr>
<td><strong>Status bar</strong></td>
<td>Connection status, viewport dimensions, and WebSocket endpoint</td>
</tr>
</tbody>
</table>
## WebSocket protocol
The dashboard connects to the same WebSocket endpoint used by [Streaming](/streaming), with additional message types for observability:
### Command events
Sent when a command begins executing:
```json
{
"type": "command",
"action": "click",
"id": "r123",
"params": { "selector": "@e5" },
"timestamp": 1711367000000
}
```
### Result events
Sent when a command finishes:
```json
{
"type": "result",
"id": "r123",
"action": "click",
"success": true,
"data": {},
"duration_ms": 45,
"timestamp": 1711367000045
}
```
### Console events
Sent when the browser logs to the console:
```json
{
"type": "console",
"level": "log",
"text": "Page loaded",
"args": [{"type": "string", "value": "Page loaded"}],
"timestamp": 1711367000100
}
```
The `args` array contains the raw CDP `Runtime.consoleAPICalled` arguments for programmatic access. Object arguments include preview data (e.g. `{userId: "abc", count: 42}` instead of `"Object"`).
These are in addition to the existing `frame`, `status`, and `error` message types documented on the [Streaming](/streaming) page.
## Architecture
The dashboard is a Next.js static export (`output: 'export'`) that produces plain HTML, CSS, and JS. It lives at `packages/dashboard/` in the monorepo and is built with:
```bash
pnpm build:dashboard
```
The built files are served by the daemon's stream server on the same port used for WebSocket connections. Plain HTTP requests serve the dashboard, while WebSocket upgrade requests are handled as before.
When the dashboard is not installed, visiting the HTTP endpoint shows instructions to run `agent-browser dashboard install`.
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("diffing");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
-175
View File
@@ -1,175 +0,0 @@
import { DiffDemo } from "@/components/diff-demo"
# Diffing
Compare page states to detect changes -- structurally via accessibility tree snapshots, visually via pixel comparison, or across two different URLs.
<DiffDemo />
## Commands
<table>
<thead>
<tr><th>Command</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td><code>diff snapshot</code></td><td>Compare current snapshot to last snapshot in session</td></tr>
<tr><td><code>diff snapshot --baseline &lt;file&gt;</code></td><td>Compare current snapshot to a saved file</td></tr>
<tr><td><code>diff screenshot --baseline &lt;file&gt;</code></td><td>Visual pixel diff against a baseline image</td></tr>
<tr><td><code>diff url &lt;url1&gt; &lt;url2&gt;</code></td><td>Compare two pages (snapshot + optional screenshot)</td></tr>
</tbody>
</table>
## Snapshot diff
Compares the accessibility tree between two points in time using a line-level text diff.
```bash
# Compare against the last snapshot taken in this session
agent-browser diff snapshot
# Compare against a saved baseline file
agent-browser diff snapshot --baseline before.txt
# Scope to a specific part of the page
agent-browser diff snapshot --selector "#main" --compact
```
Without `--baseline`, the command automatically compares against the most recent snapshot taken in the current session. This is the primary use case for agents verifying that an action had the intended effect.
### Options
<table>
<thead>
<tr><th>Flag</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td><code>-b, --baseline &lt;file&gt;</code></td><td>Path to a saved snapshot file to compare against</td></tr>
<tr><td><code>-s, --selector &lt;sel&gt;</code></td><td>Scope the current snapshot to a CSS selector or @ref</td></tr>
<tr><td><code>-c, --compact</code></td><td>Use compact snapshot format</td></tr>
<tr><td><code>-d, --depth &lt;n&gt;</code></td><td>Limit snapshot tree depth</td></tr>
</tbody>
</table>
### Output
The diff uses `+` for added lines and `-` for removed lines, similar to unified diff format. A summary line shows the count of additions, removals, and unchanged lines.
```
- button "Submit" [ref=e2]
+ button "Submit" [ref=e2] [disabled]
3 additions, 2 removals, 41 unchanged
```
## Screenshot diff
Compares the current page screenshot against a baseline image at the pixel level. Produces a diff image with changed pixels highlighted in red.
```bash
# Basic visual diff
agent-browser diff screenshot --baseline before.png
# Save diff image to a specific path
agent-browser diff screenshot --baseline before.png --output diff.png
# Adjust threshold and scope to element
agent-browser diff screenshot --baseline before.png --threshold 0.2 --selector "#hero"
```
### Options
<table>
<thead>
<tr><th>Flag</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td><code>-b, --baseline &lt;file&gt;</code></td><td>Baseline PNG/JPEG image to compare against (required)</td></tr>
<tr><td><code>-o, --output &lt;file&gt;</code></td><td>Path for the generated diff image (default: temp dir)</td></tr>
<tr><td><code>-t, --threshold &lt;0-1&gt;</code></td><td>Color distance threshold (default: 0.1). Higher = more tolerant</td></tr>
<tr><td><code>-s, --selector &lt;sel&gt;</code></td><td>Scope the current screenshot to an element</td></tr>
<tr><td><code>--full</code></td><td>Take a full-page screenshot</td></tr>
</tbody>
</table>
### Output
Reports the diff image path, number of different pixels, and mismatch percentage. The diff image shows unchanged pixels dimmed with changed pixels in red.
If the baseline and current images have different dimensions, the command reports a dimension mismatch instead of attempting pixel comparison.
## URL diff
Compares two pages by navigating to each in sequence and diffing the results.
```bash
# Compare two URLs (snapshot diff)
agent-browser diff url https://staging.example.com https://prod.example.com
# Include visual comparison
agent-browser diff url https://v1.example.com https://v2.example.com --screenshot
# Full-page screenshot comparison
agent-browser diff url https://v1.example.com https://v2.example.com --screenshot --full
```
The command navigates to the first URL, captures state, then navigates to the second URL and captures again. Snapshot diff is always included. Screenshot diff requires the `--screenshot` flag.
After completion, the browser remains on the second URL.
### Options
<table>
<thead>
<tr><th>Flag</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td><code>--screenshot</code></td><td>Also perform visual screenshot comparison</td></tr>
<tr><td><code>--full</code></td><td>Use full-page screenshots</td></tr>
<tr><td><code>--wait-until &lt;strategy&gt;</code></td><td>Navigation wait strategy: <code>load</code>, <code>domcontentloaded</code>, <code>networkidle</code> (default: <code>load</code>)</td></tr>
<tr><td><code>-s, --selector &lt;sel&gt;</code></td><td>Scope snapshots to a CSS selector or @ref</td></tr>
<tr><td><code>-c, --compact</code></td><td>Use compact snapshot format</td></tr>
<tr><td><code>-d, --depth &lt;n&gt;</code></td><td>Limit snapshot tree depth</td></tr>
</tbody>
</table>
## Use cases
### Verifying agent actions
The most common use case: confirm that an action (click, fill, submit) changed the page as expected.
```bash
agent-browser snapshot -i # Take interactive-only snapshot (baseline)
agent-browser fill @e3 "test@example.com"
agent-browser diff snapshot # Compare current snapshot to the baseline
```
### Monitoring for changes
Periodically compare a page against a saved baseline to detect updates.
```bash
# Save baseline
agent-browser open https://example.com && agent-browser snapshot > baseline.txt
# Later, check for changes
agent-browser open https://example.com && agent-browser diff snapshot --baseline baseline.txt
```
### Visual regression testing
Compare screenshots before and after a deploy to catch unintended visual changes.
```bash
agent-browser open https://staging.example.com && agent-browser screenshot baseline.png
# ... deploy happens ...
agent-browser open https://staging.example.com && agent-browser diff screenshot --baseline baseline.png
```
### Comparing environments
Diff staging against production to verify parity.
```bash
agent-browser diff url https://staging.example.com https://prod.example.com --screenshot
```
-7
View File
@@ -1,7 +0,0 @@
import { pageMetadata } from "@/lib/page-metadata";
export const metadata = pageMetadata("engines/chrome");
export default function Layout({ children }: { children: React.ReactNode }) {
return children;
}
-104
View File
@@ -1,104 +0,0 @@
# Chrome
Chrome (and Chromium) is the default browser engine. agent-browser discovers, launches, and manages the Chrome process automatically via the Chrome DevTools Protocol (CDP).
## Binary Discovery
When no `--executable-path` is provided, agent-browser searches for Chrome in this order:
<table>
<thead>
<tr><th>Platform</th><th>Locations checked</th></tr>
</thead>
<tbody>
<tr>
<td>macOS</td>
<td>
<code>/Applications/Google Chrome.app</code>,
<code>/Applications/Google Chrome Canary.app</code>,
<code>/Applications/Chromium.app</code>,
<code>/Applications/Brave Browser.app</code>,
Puppeteer cache (<code>~/.cache/puppeteer/chrome/</code> or <code>PUPPETEER_CACHE_DIR</code>),
Chrome for Testing cache
</td>
</tr>
<tr>
<td>Linux</td>
<td>
<code>google-chrome</code>,
<code>google-chrome-stable</code>,
<code>chromium-browser</code>,
<code>chromium</code> in PATH,
Puppeteer cache (<code>~/.cache/puppeteer/chrome/</code> or <code>PUPPETEER_CACHE_DIR</code>),
Chrome for Testing cache
</td>
</tr>
<tr>
<td>Windows</td>
<td>
<code>%LOCALAPPDATA%\Google\Chrome\Application\chrome.exe</code>,
<code>C:\Program Files\Google\Chrome\Application\chrome.exe</code>,
<code>C:\Program Files (x86)\...\chrome.exe</code>
</td>
</tr>
</tbody>
</table>
If Chrome is not found, run `agent-browser install` to download Chrome from Chrome for Testing.
## Usage
Chrome is the default engine -- no `--engine` flag is needed:
```bash
agent-browser open example.com
```
To be explicit:
```bash
agent-browser --engine chrome open example.com
```
## Custom Binary
Point to any Chromium-based browser with `--executable-path`:
```bash
agent-browser --executable-path /path/to/chromium open example.com
```
Or via environment variable:
```bash
export AGENT_BROWSER_EXECUTABLE_PATH=/path/to/chromium
agent-browser open example.com
```
## Chrome-Specific Features
These features are available only with Chrome:
<table>
<thead>
<tr><th>Feature</th><th>Flag</th></tr>
</thead>
<tbody>
<tr><td>Browser extensions</td><td><code>--extension &lt;path&gt;</code></td></tr>
<tr><td>Persistent profiles</td><td><code>--profile &lt;path&gt;</code> (sets Chrome's <code>--user-data-dir</code>)</td></tr>
<tr><td>Storage state</td><td><code>--state &lt;path&gt;</code></td></tr>
<tr><td>File URL access</td><td><code>--allow-file-access</code></td></tr>
<tr><td>Headed mode</td><td><code>--headed</code></td></tr>
<tr><td>Custom launch args</td><td><code>--args &lt;args&gt;</code></td></tr>
</tbody>
</table>
## Containers and CI
In Docker, CI runners, or other sandboxed environments, Chrome's user namespace sandbox may need to be disabled:
```bash
agent-browser --args "--no-sandbox" open example.com
```
agent-browser automatically adds `--no-sandbox` when it detects a container environment (Docker, Podman, running as root).

Some files were not shown because too many files have changed in this diff Show More