Compare commits

..
Author SHA1 Message Date
leeguooooo c26afbaba6 chore(release): bump to 0.27.0-fork.8 — auto-retry transient occlusion 2026-05-09 12:39:14 +09:00
leeguooooo ffb386e3af feat(click): auto-retry on transient occlusion before erroring
fork.7 caught the X mask-overlay race correctly but reported it to
the user verbatim — every transient overlay (modal backdrop, focus
ring, click-outside mask, sticky banner) became an error the user
had to wrap in their own retry loop. Most of these clear within a
frame or two on their own.

Now `verify_click_target` retries the elementFromPoint probe a few
times (default 3 × 200ms = 600ms total grace period) before failing.
Real-world overlays that blink in for a render cycle clear during
the first retry; persistent overlays still surface as errors with
the same actionable message — just qualified with "still occluded
after N retries / Mms" so the user knows we tried.

Tunable:
  AGENT_BROWSER_OCCLUSION_RETRIES         (default 3, 0 disables)
  AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS  (default 200)

DOM.resolveNode is called once outside the loop — backendNodeId is
stable across renders, only the element under (x, y) changes when
overlays flicker. Each probe is still capped at 500ms so a stuck
Runtime.callFunctionOn can't stall a click for longer than the user
expects.
2026-05-09 12:39:03 +09:00
leeguooooo 947d150561 fix(docker): escape \$ as \$\$ so docker compose doesn't eat shell vars
Real bug behind 0.27.0-fork.5 and fork.7 shipping stale linux binaries.
Docker compose interpolates \${VAR} (and \$VAR) at YAML parse time
against the host shell — including inside `command:` blocks. So:

  PID1=\$!                ← compose sees \$! → host has no `!` var → ""
  wait \$PID1 ...         ← compose sees \$PID1 → "" → becomes `wait `
  SRC="...\$TARGET..."    ← \$TARGET still works (set in `environment:`)
  cp "\$SRC" "..."        ← \$SRC eaten → empty → cp errors silently

Result: the per-PID error check I added in dbf272c never fired
because both lines were `wait` (no args) — which waits for ALL
children and exits with the LAST one's status, not each individually.
A failing arm64 build couldn't fail the script.

Fix: escape every script-local \$ as \$\$. Docker compose translates
\$\$ → literal \$ when materializing the command for the container,
and the in-container shell then expands \$VAR correctly.

Verified by `docker compose config` showing the resolved command
contains \$\$PID1 / \$\$SRC etc (which becomes \$PID1 / \$SRC in the
container's bash).
2026-05-09 11:07:01 +09:00
leeguooooo 06a29251a2 chore(release): bump to 0.27.0-fork.7 — click occlusion guard 2026-05-09 10:49:39 +09:00
leeguooooo 0eacec9b9f fix(click): occlusion check via document.elementFromPoint before dispatch
Closes the "modal silently closes when clicking 'Add post' on a thread"
bug. Verified root cause via instrumented page-side click logger:

  click @e31 (aria-label="Add post" at button (1034, 285))
  → mouse event dispatched to (1045, 296)
  → document.elementFromPoint(1045, 296) returned:
       DIV[testid="mask"], bounds (0,0,1746x934)
  → X interpreted as "click outside modal" → close + nav to /home

The cached coordinates were correct. Between snapshot and click, X
laid a transient full-viewport mask over the modal (their own
"click-outside-to-close" overlay). stealth dispatched the click
without checking what was actually at that pixel — the overlay
intercepted it.

Fix: just before returning (x, y) from resolve_element_center for
ref-based interactions, run a Runtime.callFunctionOn against the
ref's resolved element with `function(x, y) { return this.contains(
document.elementFromPoint(x, y)) || that.contains(this) ? null :
{...occluder details...}; }`. If the element at the point isn't us
(or our descendant — clicking the SVG icon inside a button is fine
— or our ancestor), we fail with a specific message:

  Ref @e31 is occluded by DIV[testid=mask] at the click point.
  A transient overlay (modal backdrop, mask, sticky banner, etc.)
  appeared between snapshot and click. Wait for it to clear or
  re-snapshot, then retry.

So instead of silently submitting an entire thread or nuking the
user's modal, agent gets a parseable error and can wait + retry.

Tight 500ms timeout per CDP call (matching the verify_ref_identity
defensive guard from fork.6) so a stuck DOM.resolveNode can't
re-introduce the multi-minute hang we just fixed. On any timeout
or error in the guard itself, fall through and let the click
proceed — strictly no worse than the unguarded code path.

Disable with AGENT_BROWSER_VERIFY_CLICK_TARGET=0.
2026-05-09 10:49:27 +09:00
leeguooooo 7159012173 chore(release): bump to 0.27.0-fork.6 — defensive-guard timeouts + accurate CDP tip 2026-05-09 10:04:25 +09:00
leeguooooo 1b3d41e579 fix(timeout): cap defensive CDP guards so click can't hang multi-minute
Reported: a single `click @ref` could hang 5+ minutes, with multiple
queued click invocations adding up to 7+ minutes — worst case 30s
timeout × 3 CDP calls × N parallel processes:

  - verify_ref_identity (Accessibility.getPartialAXTree)  →  default 30s
  - resolveNode / getBoxModel                              →  default 30s
  - wait_for_paint_settled (Runtime.evaluate awaitPromise) →  default 30s

The latter two are best-effort defenses added in fork.3-5 to fix SPA
race / DOM-reuse bugs. They should never block a real click for
30s — the unguarded code path was always faster than the guarded
path-that-hangs.

  - verify_ref_identity   capped at 1s   (skips check on timeout)
  - wait_for_paint_settled capped at 500ms (skips wait on timeout)

Both skip-on-timeout intentionally: the worst case is the click
behaves like fork.2 (race-prone but fast), which is strictly better
than the user pkilling stuck processes.

Also rewrites the misleading "Chrome 144+ chrome://inspect tip" in
the auto-connect failure message — the toggle exposes target
discovery only, not the /json/version HTTP API the auto-connect
flow expects (verified by user: lsof shows :9222 listening but
curl /json/version returns 404).
2026-05-09 10:04:14 +09:00
leeguooooo dbf272ced7 fix(docker): catch parallel-build failures + stop using glob in cp
Two latent bugs in the release pipeline that conspired to ship a stale
linux-x64 binary in 0.27.0-fork.5 (only caught by manually grepping
the embedded version string):

1. build-linux ran x64 and arm64 in parallel and used a single
   `wait $PID1 $PID2` to join them. That command waits for both, but
   its exit code is the LAST waited pid only — so if x64 silently
   broke and arm64 succeeded, the outer script exited 0 and shipped
   whatever was already in /output from the previous release. Now we
   wait on each pid individually and exit 1 on either failure.

2. build-single's cp used `agent-browser*` which globs to BOTH the
   binary and its `.d` dependency file. When two sources are passed,
   cp requires the destination to be a directory. We weren't, so cp
   exited non-zero with "Not a directory" and the build script
   shrugged it off because the next line was `chmod ... || true`.
   Now we resolve a single explicit source path.
2026-05-09 04:30:20 +09:00
7 changed files with 250 additions and 22 deletions
+1 -1
View File
@@ -45,7 +45,7 @@ dependencies = [
[[package]]
name = "agent-browser-stealth"
version = "0.27.0-fork.5"
version = "0.27.0-fork.8"
dependencies = [
"aes-gcm",
"async-trait",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "agent-browser-stealth"
version = "0.27.0-fork.5"
version = "0.27.0-fork.8"
edition = "2021"
description = "Fast browser automation CLI for AI agents"
license = "Apache-2.0"
+6 -4
View File
@@ -1607,8 +1607,9 @@ async fn auto_launch(state: &mut DaemonState) -> Result<(), String> {
To let agent-browser work with your existing Chrome (recommended):\n\
{}\n\n\
Or start a standalone browser with: agent-browser --launch open <url>\n\n\
Tip: On Chrome 144+, you can enable CDP without restarting:\n\
Open chrome://inspect/#remote-debugging and toggle it on.",
Note: chrome://inspect/#remote-debugging only enables remote *target discovery* — \
it does NOT expose the standard CDP HTTP API on /json/version. \
A full restart with --remote-debugging-port=<port> is required.",
chrome_relaunch_hint(),
));
}
@@ -2135,8 +2136,9 @@ async fn handle_launch(cmd: &Value, state: &mut DaemonState) -> Result<Value, St
To let agent-browser work with your existing Chrome (recommended):\n\
{}\n\n\
Or start a standalone browser with: agent-browser --launch open <url>\n\n\
Tip: On Chrome 144+, you can enable CDP without restarting:\n\
Open chrome://inspect/#remote-debugging and toggle it on.",
Note: chrome://inspect/#remote-debugging only enables remote *target discovery* — \
it does NOT expose the standard CDP HTTP API on /json/version. \
A full restart with --remote-debugging-port=<port> is required.",
chrome_relaunch_hint(),
));
}
+204 -3
View File
@@ -201,6 +201,31 @@ pub async fn resolve_element_center(
if let Ok(r) = result {
let (x, y) = box_model_center(&r.model);
// Occlusion check: a transient overlay (X.com's "click
// outside to close" mask, modal backdrop, sticky banner,
// etc.) can land on top of our target between snapshot
// and click. Coordinates are correct, but
// `document.elementFromPoint(x, y)` returns the overlay
// — and the click goes to the overlay's handler, not
// ours. Catch it here so the user gets "occluded by
// DIV[testid=mask]" instead of "modal silently closed +
// thread submitted by accident".
//
// Set AGENT_BROWSER_VERIFY_CLICK_TARGET=0 to skip.
if std::env::var("AGENT_BROWSER_VERIFY_CLICK_TARGET").as_deref() != Ok("0") {
if let Err(e) = verify_click_target(
client,
effective_session_id,
backend_node_id,
&ref_id,
x,
y,
)
.await
{
return Err(e);
}
}
return Ok((x, y, effective_session_id.to_string()));
}
// backend_node_id is stale; re-query the accessibility tree below
@@ -396,9 +421,23 @@ async fn verify_ref_identity(
"backendNodeId": backend_node_id,
"fetchRelatives": false,
});
let resp: Result<GetFullAXTreeResult, String> = client
.send_command_typed("Accessibility.getPartialAXTree", &params, Some(session_id))
.await;
// Tight 1s timeout: this is a defensive guard, not a critical path.
// The default 30s CDP timeout was the dominant factor in the
// "click hangs 5+ minutes" report — three CDP calls (verify +
// resolveNode + paint-settle) at 30s each, multiplied by parallel
// click invocations queueing on the daemon, totalled multi-minute
// user-visible hangs. Cap our own helper so a stuck AX query
// doesn't make `click` worse than the no-guard version was.
let resp: Result<GetFullAXTreeResult, String> = match tokio::time::timeout(
std::time::Duration::from_secs(1),
client.send_command_typed("Accessibility.getPartialAXTree", &params, Some(session_id)),
)
.await
{
Ok(r) => r,
// Timeout: skip identity verification rather than block the click.
Err(_) => return Ok(()),
};
let Ok(tree) = resp else {
// Node likely gone; let the box-model call fail and trigger fallback.
return Ok(());
@@ -426,6 +465,168 @@ async fn verify_ref_identity(
))
}
/// At the moment we'd dispatch the click, ask the page itself which element
/// occupies (x, y). If it's not our target (and not a descendant or
/// ancestor), an overlay has appeared between snapshot and click — we'd
/// silently click the overlay otherwise. Returns Err with details about
/// the occluding element so the caller can wait + re-snapshot.
///
/// Implemented as a single Runtime.callFunctionOn: resolve the cached
/// backendNodeId to a remote object, then run a function on it that
/// compares with elementFromPoint. The function returns null when the
/// click is safe and a JSON string with diagnostic info when it isn't.
async fn verify_click_target(
client: &CdpClient,
session_id: &str,
backend_node_id: i64,
ref_id: &str,
x: f64,
y: f64,
) -> Result<(), String> {
use serde::Deserialize;
// Resolve once. backendNodeId is stable across renders; only the
// element under (x, y) is what changes when an overlay flickers.
let resolve_params = DomResolveNodeParams {
backend_node_id: Some(backend_node_id),
node_id: None,
object_group: Some("agent-browser-occlusion".to_string()),
};
let resolve_fut = client.send_command_typed::<_, serde_json::Value>(
"DOM.resolveNode",
&resolve_params,
Some(session_id),
);
let Ok(resolve_resp) =
tokio::time::timeout(std::time::Duration::from_millis(500), resolve_fut).await
else {
return Ok(());
};
let Ok(resolved) = resolve_resp else { return Ok(()) };
let Some(object_id) = resolved
.get("object")
.and_then(|o| o.get("objectId"))
.and_then(|v| v.as_str())
else {
return Ok(());
};
// Auto-retry on transient occlusion. Many real-world overlays
// (modal backdrops, focus rings, click-outside masks) blink in for
// a frame or two during state transitions and clear on their own.
// Without retries the user gets an "occluded" error and has to
// wrap every click in their own retry loop. With retries the
// common case is invisible — only persistent overlays surface.
//
// AGENT_BROWSER_OCCLUSION_RETRIES (default 3, 0 disables)
// AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS (default 200)
let max_retries: u32 = std::env::var("AGENT_BROWSER_OCCLUSION_RETRIES")
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(3);
let retry_delay_ms: u64 = std::env::var("AGENT_BROWSER_OCCLUSION_RETRY_DELAY_MS")
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(200);
#[derive(Deserialize)]
struct Occluder {
tag: Option<String>,
testid: Option<String>,
role: Option<String>,
#[serde(rename = "ariaLabel")]
aria_label: Option<String>,
text: Option<String>,
reason: Option<String>,
}
// function(x, y) { ... } where `this` is the target element.
// Return null → click is safe.
// Return JSON → describes the occluding element.
let function_decl = "function(x, y) { \
const at = document.elementFromPoint(x, y); \
if (!at) return JSON.stringify({reason:'no-element-at-point'}); \
if (at === this || this.contains(at) || at.contains(this)) return null; \
return JSON.stringify({ \
tag: at.tagName, \
testid: (at.dataset && at.dataset.testid) || null, \
role: at.getAttribute('role'), \
ariaLabel: at.getAttribute('aria-label'), \
text: ((at.textContent||'').trim().slice(0, 60)) \
}); \
}";
let mut last_occ: Option<Occluder> = None;
for attempt in 0..=max_retries {
if attempt > 0 {
tokio::time::sleep(std::time::Duration::from_millis(retry_delay_ms)).await;
}
let call_params = serde_json::json!({
"objectId": object_id,
"functionDeclaration": function_decl,
"arguments": [{"value": x}, {"value": y}],
"returnByValue": true,
});
let call_fut = client.send_command_typed::<_, serde_json::Value>(
"Runtime.callFunctionOn",
&call_params,
Some(session_id),
);
let Ok(call_resp) =
tokio::time::timeout(std::time::Duration::from_millis(500), call_fut).await
else {
return Ok(()); // probe itself stalled — fall through to click
};
let Ok(call_result) = call_resp else {
return Ok(());
};
let value = call_result.get("result").and_then(|r| r.get("value"));
let json_str = match value {
Some(serde_json::Value::String(s)) => s.clone(),
// null / undefined → element at point IS our target. Safe.
_ => return Ok(()),
};
let occ: Occluder = match serde_json::from_str(&json_str) {
Ok(v) => v,
Err(_) => return Ok(()),
};
last_occ = Some(occ);
}
// All retries exhausted — overlay is sticky. Build the descriptive error.
let occ = last_occ.expect("loop ran at least once");
if let Some(reason) = occ.reason {
return Err(format!(
"Ref {} cannot be clicked at its computed position: {}. \
The element may have moved off-screen — re-run snapshot.",
ref_id, reason
));
}
let mut desc = occ.tag.unwrap_or_else(|| "unknown".to_string());
if let Some(t) = occ.testid {
desc.push_str(&format!("[testid={}]", t));
}
if let Some(r) = occ.role {
desc.push_str(&format!("[role={}]", r));
}
if let Some(a) = occ.aria_label {
desc.push_str(&format!("[aria-label=\"{}\"]", a));
}
if let Some(t) = occ.text {
if !t.is_empty() {
desc.push_str(&format!(" text=\"{}\"", t));
}
}
let waited_ms = (max_retries as u64) * retry_delay_ms;
Err(format!(
"Ref {} is occluded by {} at the click point (still occluded after \
{} retries / {}ms). A persistent overlay is in the way — \
re-run snapshot, dismiss the overlay, or set \
AGENT_BROWSER_VERIFY_CLICK_TARGET=0 to bypass.",
ref_id, desc, max_retries, waited_ms,
))
}
/// Re-query the accessibility tree to find a node matching role+name+nth,
/// returning its fresh backendDOMNodeId. This uses the same data source
/// (Accessibility.getFullAXTree) that built the ref map during snapshot,
+12 -4
View File
@@ -903,8 +903,15 @@ async fn wait_for_paint_settled(client: &CdpClient, session_id: &str) {
requestAnimationFrame(() => \
requestAnimationFrame(() => \
queueMicrotask(() => resolve(true)))))";
let _ = client
.send_command_typed::<_, Value>(
// Tight 500ms timeout. RAF normally fires at 16ms, two RAFs total ~33ms.
// If the tab is hidden / throttled / page is doing something pathological
// and RAF doesn't fire in 500ms, we'd rather return now than stall the
// user's click. Without this cap, a stuck RAF inherited the default 30s
// CDP timeout and was the main contributor to the "click hangs 5+ min"
// user report.
let _ = tokio::time::timeout(
std::time::Duration::from_millis(500),
client.send_command_typed::<_, Value>(
"Runtime.evaluate",
&EvaluateParams {
expression: script.to_string(),
@@ -912,8 +919,9 @@ async fn wait_for_paint_settled(client: &CdpClient, session_id: &str) {
await_promise: Some(true),
},
Some(session_id),
)
.await;
),
)
.await;
}
async fn dispatch_click(
+25 -8
View File
@@ -20,13 +20,19 @@ services:
# Build both targets in parallel
(echo "→ Linux x64" && cargo zigbuild --release --target x86_64-unknown-linux-gnu && cp /build/target/x86_64-unknown-linux-gnu/release/agent-browser /output/agent-browser-linux-x64 && chmod +x /output/agent-browser-linux-x64 && echo "✓ Linux x64 done") &
PID1=$!
PID1=$$!
(echo "→ Linux ARM64" && cargo zigbuild --release --target aarch64-unknown-linux-gnu && cp /build/target/aarch64-unknown-linux-gnu/release/agent-browser /output/agent-browser-linux-arm64 && chmod +x /output/agent-browser-linux-arm64 && echo "✓ Linux ARM64 done") &
PID2=$!
PID2=$$!
# Wait for both to complete
wait $PID1 $PID2
# Wait for both and check exit codes individually — without this
# the outer script exits 0 even if one of the parallel builds
# failed, silently leaving a stale binary in /output from the
# previous release. Caused 0.27.0-fork.5 to ship with a stale
# linux-x64 binary at the first publish attempt until caught
# manually by checking the embedded version string.
wait $$PID1 || { echo "✗ Linux x64 build failed"; exit 1; }
wait $$PID2 || { echo "✗ Linux ARM64 build failed"; exit 1; }
echo ""
echo "✓ Linux platforms built successfully!"
@@ -65,10 +71,21 @@ services:
environment:
- TARGET=${TARGET:-x86_64-unknown-linux-gnu}
- OUTPUT_NAME=${OUTPUT_NAME:-agent-browser-linux-x64}
# NOTE: $$ escapes a literal $ for the in-container shell. A single $ is
# interpolated by docker compose at YAML parse time against the *host*
# environment, which silently drops script-local variables like SRC
# (caused 0.27.0-fork.7 to ship with a stale linux-arm64 binary because
# the cp command resolved to `cp "" "/output/"` after compose ate $SRC
# and $OUTPUT_NAME). $TARGET / $OUTPUT_NAME are set via `environment:`
# below — those are also passed into the container, so $$TARGET and
# $$OUTPUT_NAME read them at script time.
command: |
-c '
cargo zigbuild --release --target $TARGET
cp /build/target/$TARGET/release/agent-browser* /output/$OUTPUT_NAME
chmod +x /output/$OUTPUT_NAME 2>/dev/null || true
echo "✓ Built $OUTPUT_NAME"
set -e
cargo zigbuild --release --target $$TARGET
SRC="/build/target/$$TARGET/release/agent-browser"
if [ -f "$$SRC.exe" ]; then SRC="$$SRC.exe"; fi
cp "$$SRC" "/output/$$OUTPUT_NAME"
chmod +x /output/$$OUTPUT_NAME 2>/dev/null || true
echo "✓ Built $$OUTPUT_NAME"
'
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "agent-browser-stealth",
"version": "0.27.0-fork.5",
"version": "0.27.0-fork.8",
"description": "Browser automation CLI for AI agents — stealth fork with anti-detection",
"type": "module",
"files": [