* fix: use correct Windows virtual-key codes for punctuation in type command
The `type` command was dropping punctuation characters like `.`, `'`, and
`#` because `char_to_key_info()` used raw ASCII codes as the
`windowsVirtualKeyCode` in CDP `Input.dispatchKeyEvent` calls. For
punctuation the ASCII value collides with unrelated VK codes — most
critically '.' (ASCII 46) equals VK_DELETE (0x2E), causing Chrome to
interpret periods as Delete key presses.
Changes:
- Add `punctuation_key_info()` with correct VK_OEM_* codes matching
Playwright's USKeyboardLayout (e.g. Period=190, Slash=191, Semicolon=186)
- Fall back to `Input.insertText` for characters without a US keyboard
mapping (emoji, CJK, etc.), matching Playwright's `keyboard.type()`
- Update e2e test to use `type` instead of `fill` workaround for email
- Add unit tests verifying VK code parity with Playwright's layout
Fixes#833
* style: fix rustfmt formatting for InsertTextParams
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
The v0.20.0 migration from Playwright to the native Rust daemon introduced
two regressions in checkbox/radio handling:
1. `is_element_checked` only read `this.checked`, which is undefined on
non-input elements. Material Design and ARIA controls use wrapper divs
with `role="checkbox"` and `aria-checked`, or hide the native input
off-screen inside a label. The function now mirrors Playwright's
`getChecked()` with follow-label retargeting: native `.checked`,
`aria-checked` for ARIA roles, `label.control` traversal, and nested
input lookup.
2. `check`/`uncheck` accepted the coordinate-based CDP click result
without verifying the state actually changed. When the AX tree's
`backendDOMNodeId` points to a hidden off-screen input (common in
Material Design), `Input.dispatchMouseEvent` hits nothing. The actions
now re-check state after clicking and fall back to a JS `.click()` on
the resolved input — matching Playwright's `_setChecked` verify step.
Adds e2e regression test covering Material Design (hidden input + ripple
overlay), ARIA-only, and native checkbox patterns.
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
When using --auto-connect, discover_and_attach_targets() was selecting
Chrome internal pages (chrome://, chrome-extension://, devtools://) as
the active target. Follow-up commands like `get url` and `snapshot`
would then return data from targets like chrome://omnibox-popup.top-chrome/
instead of the actual application tab.
Add is_internal_chrome_target() filter to exclude internal Chrome targets
from the discovery results. If no user-facing targets remain after
filtering, the existing "create a new tab" fallback handles it.
Fixes#813
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
* fix: restore WebSocket streaming in native daemon
The v0.20.0 Rust rewrite broke WebSocket streaming — connections opened
but received zero messages before closing. Multiple issues contributed:
1. StreamServer was dropped immediately after creation in daemon.rs,
closing the broadcast channel and killing all WS connections.
2. Screencast frames were only processed during command polling
(drain_cdp_events) instead of in real-time, unlike the 0.19.0
TypeScript cdp.on('Page.screencastFrame') callback.
3. Auto-start/stop screencast on WS client connect/disconnect was
missing from the Rust implementation.
4. Screencast CDP commands used the wrong session ID (daemon session
name instead of the CDP page session from Target.attachToTarget).
5. Broadcast channel Lagged errors killed WS connections instead of
being handled gracefully.
The fix adds a background CDP event loop in StreamServer that subscribes
to Chrome events and broadcasts screencast frames in real-time, properly
tracks the CDP page session ID, restores auto-screencast lifecycle, and
keeps the StreamServer alive in DaemonState.
Fixes#820
* fix: use actual CDP session ID for input dispatch in stream WebSocket
Pass the real cdp_session_id (from Target.attachToTarget) through to
handle_ws_client instead of an empty string. Previously, input commands
(mouse, keyboard, touch) were sent with `"sessionId": ""` which Chrome
silently rejects. Now the correct page session ID is read at dispatch
time, and when no session ID is set yet (before browser launch),
the field is omitted entirely via `None` so Chrome uses browser-level
dispatch.
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: snapshot --selector scopes to the matched element subtree
The native Rust daemon accepted the --selector flag but never used it —
the full accessibility tree was always returned regardless of the
selector. This restores the 0.19.0 behaviour where snapshot --selector
returns only the subtree rooted at the matched CSS selector.
The implementation resolves the selector via Runtime.evaluate, fetches
the full DOM subtree with DOM.describeNode(depth: -1) to collect all
descendant backendNodeIds, then filters the AX tree to render only the
nodes whose backendDOMNodeId falls within that set. This correctly
handles elements like <body> that don't map to a direct AX node.
Also fixes handle_snapshot reading "depth" instead of "maxDepth" from
the command JSON, which caused --depth to be silently ignored.
Fixes#822
* style: run cargo fmt on snapshot.rs
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* fix: add appium: vendor prefix to iOS capabilities for Appium v3
Appium v3 enforces the W3C WebDriver spec strictly, requiring
non-standard capabilities to use vendor prefixes. The iOS provider
was sending capabilities like `automationName`, `noReset`, `deviceName`,
`platformVersion`, and `udid` without the required `appium:` prefix,
causing session creation to fail with InvalidArgumentError.
This change prefixes all non-standard capabilities with `appium:` while
leaving standard W3C capabilities (`platformName`, `browserName`)
unprefixed. Backwards-compatible with Appium v2, which accepts both
formats.
Fixes#629
* fix: extract build_ios_capabilities for testable production code path
Addresses review feedback: removes unused `mut manager` warning and
validates the actual capability-building logic instead of reconstructing
JSON inline.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix: replace screenshot polling with screencast-based piped ffmpeg recording
Recording previously used Page.captureScreenshot polling at 10fps,
which was CPU-heavy and produced inconsistent results. Now uses
Page.startScreencast with throttled acks (35ms interval) to receive
frames event-driven from Chrome, and pipes JPEG data directly to
ffmpeg stdin in real-time instead of saving temp files.
- Spawn ffmpeg at recording start with piped stdin (image2pipe)
- Background task receives screencast frames, interpolates gaps by
repeating the last frame based on timestamps, targets 25fps
- Ack throttling controls Chrome's frame push rate
- Fix: current frame was never written after the first one
- Fix: frame count was read before task finished padding
- Remove tokio-util dependency (replaced CancellationToken with oneshot)
- Add tokio "process" feature for async child process stdin pipe
- Extract start/stop_recording_task helpers on DaemonState
- Add tests for restart, ffmpeg codec selection, and stop without task
* fmt
* chore
* fix: switch WebM codec from VP9 to VP8 for correct framerate and browser
compatibility
VP9 realtime encoder ignored input framerate, producing 10fps output
instead of 25fps. This caused inconsistent playback in browsers.
VP8 respects -framerate 25 and has wider browser playback support.
* fmt
* fix: add kill_on_drop to ffmpeg process to prevent zombie on task panic
* fix: switch from screencast to screenshot polling for reliable recording duration
Screencast only pushes frames on visual changes, producing short videos
on static pages. Screenshot polling captures at a fixed 10fps interval
regardless of page activity, guaranteeing duration matches wall-clock time.
ffmpeg piped stdin architecture is preserved — no temp files.
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
Brave Browser is Chromium-based and uses the same DevToolsActivePort
mechanism. Add its user-data-dir paths to get_chrome_user_data_dirs()
and its executable paths to find_chrome() on all three platforms
(macOS, Linux, Windows).
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: gracefully fall back to role/name lookup when backend_node_id is stale
When the DOM changes between snapshot and click (common with SPAs and
dynamic UIs), the stored backend_node_id becomes invalid. Previously,
DOM.getBoxModel and DOM.resolveNode failures propagated as hard errors,
bypassing the role/name fallback path entirely. Now these failures are
caught and the code falls through to a JS-based element lookup.
Also adds resolve_object_id_by_role_name so that resolve_element_object_id
has a fallback for ref-based lookups (previously it had none), and
improves the role matching JS to correctly map implicit ARIA roles
(e.g. <input type="submit"> → "button", <a href> → "link").
Closes#805
* test: add e2e regression test for stale ref click fallback (#805)
Verifies that clicking a ref whose backend_node_id has become stale
(because the DOM was replaced by JavaScript) falls back to role/name
lookup instead of failing with "Could not compute box model".
* fix: use accessibility tree for stale ref fallback instead of JS heuristic
Replace the hand-rolled JS role/name matching (getImplicitRole,
getAccessibleName) with a re-query of Accessibility.getFullAXTree —
the same data source that built the ref map during snapshot. This
guarantees role/name matching is identical to what was stored,
preventing silent wrong-element clicks from name computation
divergence (e.g. aria-labelledby, <label for>, alt text).
Matches v0.19.0 (Playwright) behavior where getByRole always
re-queried the live accessibility tree.
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
Replace all `eprintln!` calls in daemon-context code with
`let _ = writeln!(std::io::stderr(), ...)` so that broken pipe errors
on stderr are silently ignored instead of panicking.
The CLI client spawns the daemon with piped stderr to capture startup
errors, then drops the pipe handle once the daemon is ready. Any
subsequent `eprintln!` in the daemon panics because Rust's `eprintln!`
macro internally unwraps the write result. This caused the reported
"failed printing to stderr: Broken pipe (os error 32)" panic during
Chrome launch on Linux.
Closes#799
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
The CDP event broadcast channel (capacity 256) can overflow on slow CI
runners when Chrome emits many events during navigation. Previously,
RecvError::Lagged was treated the same as RecvError::Closed, causing
spurious "Event stream closed" errors even though Chrome was still
running. Now all 5 event-receiving loops correctly continue on Lagged
instead of breaking.
Chrome uses /dev/shm for shared memory, which is typically limited to
64MB on CI runners and containers. When Chrome exhausts this, it crashes
mid-session with "Event stream closed" errors. Auto-detect CI/container
environments and pass --disable-dev-shm-usage to use /tmp instead.
The CDP WebSocket client had three issues causing snapshot to hang
indefinitely when connected to remote browsers via WSS:
1. Binary WebSocket frames were silently dropped — remote CDP proxies
(Browserless, Browserbase, etc.) may send large responses like
Accessibility.getFullAXTree as Binary frames instead of Text frames.
2. Default tungstenite size limits (16 MiB frame / 64 MiB message)
could be exceeded by large accessibility tree responses, causing the
WebSocket connection to error out and the reader task to die.
3. When the reader task died, pending commands waited for the full
30-second timeout instead of failing immediately.
Fixes#788
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
Chrome occasionally crashes during startup on CI runners before
printing the DevTools URL, causing random e2e test failures across
different tests each run. Retry the launch with a 500ms delay to
handle these transient crashes.
- Restore the `refs` dictionary in `--json` snapshot output, matching the documented API contract
- The `refs` field was silently dropped during the Node.js to Rust rewrite (v0.20), causing consumers parsing `data.refs` for programmatic element interaction to receive no structured ref data
Fixes#785
* fix: use VP9 codec for webm recording output
The recording command hardcoded libx264 (H.264) which is incompatible
with the WebM container format. WebM only supports VP8/VP9/AV1 codecs,
causing ffmpeg to fail when users specify a .webm output file.
Select codec based on output file extension: libvpx-vp9 for .webm,
libx264 for other formats.
Fixes#778
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* refactor: use CRF mode for VP9 webm encoding
Switch from bitrate target (-b:v 2M) to constant quality mode (-crf 30),
which is the standard approach for screen recording (used by Puppeteer
and recommended by ffmpeg VP9 guide). CRF adapts bitrate to scene
complexity for more consistent quality.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: add -b:v 0 for true constant quality VP9 encoding
Without -b:v 0, libvpx-vp9 uses its default bitrate target alongside
-crf, resulting in constrained quality mode instead of true constant
quality mode.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: pad video dimensions to even numbers for h264 compatibility
libx264 requires width and height to be divisible by 2, but CDP
screencast can capture frames with odd dimensions (e.g. 1280x577).
Add pad filter to ensure even dimensions for all codecs.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The comment said "Ignore SIGPIPE" but the code actually resets SIGPIPE
to SIG_DFL (default behavior = process termination), not SIG_IGN (ignore).
Updated the comment to accurately describe what the code does and why.
Co-authored-by: hyunjinee <leehj0110@kakao.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Regenerate pnpm-lock.yaml to match the cleaned-up package.json (only
@changesets/cli remains). Add CI environment detection to
should_disable_sandbox() so Chrome launches with --no-sandbox on GitHub
Actions runners where AppArmor blocks unprivileged user namespaces.
The test was reading AGENT_BROWSER_HEADED without holding ENV_MUTEX,
causing a race with test_launch_options_from_env_headed_flag when
tests run in parallel.
Three issues prevented --engine lightpanda from working with official
Lightpanda release builds:
1. Missing --log_level info: Lightpanda release builds default to
log_level=warn, which suppresses the info-level "server running"
startup message. wait_for_address() blocks forever reading an empty
stderr pipe. Pass --log_level info explicitly.
2. --timeout 0 means instant disconnect: Lightpanda interprets 0 as
"timeout after 0ms", not "no timeout". Use 604800 (1 week, the
documented maximum) instead.
3. extract_address only matched pretty format: Release builds use
logfmt (address=HOST:PORT without spaces), but the parser only
matched the pretty format (address = HOST:PORT with spaces). Handle
both formats.
* fix: narrow "not found" pattern in to_ai_friendly_error to avoid catching
non-element errors
Change `contains("not found")` to `contains("element not found")` so that
connection/state errors like "Browser not found" pass through unchanged
instead of being incorrectly mapped to "Element not found" message.
* remove comment
* fmt
* test: use real project error message in non-element not found test
When user explicitly sets --headed false, the CLI was ignoring this
flag because the launch condition only checked if flags.headed was
true. This meant that --headed false would not trigger a launch
command, and subsequent commands would auto-launch with default
headless=true.
The fix adds a cli_headed flag to track when the user explicitly
sets --headed (regardless of value), and includes this in the
launch condition check.
Fixes#743
This PR fixes CI build failures by addressing code formatting and linting issues that were causing the builds to fail.
**Changes made:**
1. **Rust formatting fixes in `cli/src/commands.rs`:**
- Removed unnecessary multi-line formatting for clipboard operations
- Applied consistent single-line formatting for return statements
- Fixed line length and formatting for the `test_wait_text_with_timeout` test function
2. **TypeScript fixes in `src/actions.ts`:**
- Fixed `waitForFunction` usage in the `handleWait` function by replacing the function parameter approach with a string-based implementation
- Properly escaped the text parameter using `JSON.stringify` to prevent potential injection issues
These changes ensure the code passes linting checks (clippy for Rust, ESLint for TypeScript) and formatting validation (rustfmt, prettier) that are enforced in the CI pipeline.
Fixes#751
* feat: add screenshot output config, clipboard CLI commands, and fix wait --text native path
## Summary
- Add `--screenshot-dir`, `--screenshot-quality`, and `--screenshot-format` CLI flags (with corresponding `AGENT_BROWSER_SCREENSHOT_DIR`, `AGENT_BROWSER_SCREENSHOT_QUALITY`, `AGENT_BROWSER_SCREENSHOT_FORMAT` env vars) so users can configure where and how screenshots are saved without specifying a full path every time
- Add `clipboard read`, `clipboard write <text>`, `clipboard copy`, and `clipboard paste` CLI commands, exposing the existing protocol-level clipboard handlers that were previously only accessible via JSON-RPC
- Fix `wait --text` in native mode: the CLI was emitting `selector: "text=..."` (a Playwright-style locator) which native's `querySelector` can't handle. Now emits a `text` field that correctly hits the native `wait_for_text` polling path
- Add native clipboard `copy` and `paste` support via CDP `Input.dispatchKeyEvent`, and a `write` operation to the Node.js handler
* fix: resolve CI failures in Rust formatting and TypeScript typecheck
Use string-based page.evaluate for clipboard writeText to avoid
referencing `navigator` in Node.js compilation context. Run cargo fmt
to fix formatting in commands.rs and screenshot.rs.
* fix: clipboard write captures full multi-word text
Use rest[1..].join(" ") instead of rest.get(1) so unquoted multi-word
input like `clipboard write hello world` sends the full string rather
than silently dropping everything after the first word.
* improvements
* fixes
* improvements
* improvements
Fix issue where Chrome extensions specified in the `extensions` field of `config.json` were not being loaded when launching the browser.
## Problem
Extensions configured via the `extensions` field in `config.json` were not being passed to the Chrome browser launch command, causing them to be ignored.
## Changes
- Added `!flags.extensions.is_empty()` to the launch trigger condition to ensure browser launch is triggered when extensions are configured
- Added extensions to the launch command JSON payload so they are properly passed to the browser
Fixes#726
* feat: add browserless provider integration to native browser implementation
This PR adds support for the Browserless provider to the native browser implementation, expanding the available remote browser providers from 3 to 4.
## Changes Made
- **Added `connect_browserless()` function**: Implements session creation with Browserless API using environment variables for configuration
- **Updated provider routing**: Added "browserless" case to the main provider switch statement
- **Added session cleanup**: Implemented proper session termination using the stop URL returned by Browserless
- **Updated documentation**: Modified comments and error messages to include Browserless in the supported provider list
- **Environment variable support**: Added support for configurable Browserless settings including API key, URL, browser type, TTL, and stealth mode
## Implementation Details
- Uses standard Browserless session API with POST to create sessions and DELETE to terminate
- Supports both chromium and chrome browser types with validation
- Includes proper error handling for API failures and missing configuration
- Stores the stop URL as session_id for cleanup purposes
- Follows the existing provider pattern for consistency
Fixes#744
* fix: URL-encode API key in browserless session request
Use reqwest's .query() method instead of string-formatting the token
directly into the URL, matching the Node.js implementation's use of
encodeURIComponent. Prevents malformed URLs if the API key contains
special characters.
---------
Co-authored-by: ctate <366502+ctate@users.noreply.github.com>
* feat: Add browserless as a hosted option + boolean env-parsing utility
* Add ensureDomainFilter, sanitizeExistingPage and move parseBooleanParam
* Add docs in relevant places, fix utils, rename of API env var
* Update readme
* Fix env variable name in readme
* Cleanup session stop urls when errors happen
* Fix browserlessStopUrl not being assigned in happy path
The client (connection.rs) and native daemon (native/daemon.rs) used
different get_port_for_session() implementations on Windows:
- Client: i32, .chars(), djb2 — (hash << 5) - hash + c
- Daemon: i64, .bytes(), Java hashCode — hash * 31 + b
For session name "default", client computes port 50838 while the
daemon binds on 51174, causing a 5-second timeout and startup failure.
Fix: align native/daemon.rs to use the identical djb2 algorithm from
connection.rs (i32, chars, djb2), so both sides agree on the port.
Unix is unaffected (uses Unix domain sockets, no port hashing).
Tests: add port hash regression tests to all three implementations
(native/daemon.rs, connection.rs, daemon.ts) to prevent future drift.
Fixes#705