feat: rich-editor fill, box centers, screenshot downscale, disabled+docs (#41-#45)
Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled

Dogfooding backlog from this session's embedded-form/editor work.

#41 fill on rich editors: detect CodeMirror 5 / Monaco / ProseMirror /
contenteditable and set via their own API or execCommand('insertText') so
beforeinput/input fire (a raw .value/textContent write no-op'd juejin's
CodeMirror and skipped React composers). Response echoes the `engine` used.
`fill <sel> --file <path>` / `--stdin` set large multiline text without
shell-escaping. `get value` now reads CodeMirror/Monaco/contenteditable too.

#42 screenshot --max-width/--max-height/--scale, plus a default 2000px
longest-edge cap (AGENT_BROWSER_SCREENSHOT_MAX_EDGE; 0 disables) so retina
full-page shots fit an agent's image reader and --scale 0.5 makes screenshot
px line up with click px. Annotated shots are never downscaled.

#43 `box @ref` (already a top-level alias of `get box`) now also returns
centerX/centerY/inViewport in CSS px — feed straight into `click x y` when a
ref-click no-ops (e.g. a button in a cross-origin iframe).

#44 no code change needed — disabled elements already list as
`button "Save" [disabled, ref=eN]`; the reporter's missing button was
DOM-gated on validity. Added a skill note: `find text` can't reach into a
cross-origin iframe — target those by snapshot @ref.

#45 core skill now distinguishes screenshot-to-locate (discouraged) from
screenshot-to-capture a reusable image asset via `screenshot [--clip] <file>`
(encouraged), so agents stop over-reading the prohibition.

#40 (group-scoped relay) stays deferred — needs an ab-connect extension change.

Verified live: fill --file round-trips multiline+CJK+backticks; contenteditable
engine=contenteditable + get value reads it back; box gives centerX/centerY/
inViewport; screenshot of retina example.com → 2000px; disabled button shows
[disabled]. 856 tests pass.
This commit is contained in:
leeguooooo
2026-06-17 17:49:38 +09:00
parent 50b27ac0e0
commit a83d1b1df9
9 changed files with 315 additions and 33 deletions
+1 -1
View File
@@ -290,7 +290,7 @@ checksum = "613afe47fcd5fac7ccf1db93babcb082c5994d996f20b8b159f2ad1658eb5724"
[[package]]
name = "chrome-use"
version = "1.5.20"
version = "1.5.21"
dependencies = [
"aes",
"aes-gcm",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "chrome-use"
version = "1.5.20"
version = "1.5.21"
edition = "2021"
description = "Fast browser automation CLI for AI agents"
license = "Apache-2.0"
+84 -2
View File
@@ -79,6 +79,7 @@ const KNOWN_COMMANDS: &[&str] = &[
"dialog",
"upload",
"site",
"box",
];
/// Levenshtein distance, capped — small inputs only (command names).
@@ -560,9 +561,37 @@ fn parse_command_inner(args: &[String], flags: &Flags) -> Result<Value, ParseErr
"fill" => {
let sel = rest.first().ok_or_else(|| ParseError::MissingArguments {
context: "fill".to_string(),
usage: "fill <selector> <text>",
usage: "fill <selector> <text> | fill <selector> --file <path> | fill <selector> --stdin",
})?;
Ok(json!({ "id": id, "action": "fill", "selector": sel, "value": rest[1..].join(" ") }))
// Large/multiline content without shell-escaping hell (issue #41):
// `fill <sel> --file <path>` reads the value from a UTF-8 file, and
// `fill <sel> --stdin` reads it from stdin — sent verbatim, so backticks,
// quotes, newlines and non-ASCII pass through untouched.
let value = match rest.get(1).copied() {
Some("--file") => {
let path = rest.get(2).ok_or(ParseError::InvalidValue {
message: "fill --file requires a path".to_string(),
usage: "fill <selector> --file <path>",
})?;
std::fs::read_to_string(path).map_err(|e| ParseError::InvalidValue {
message: format!("fill --file: cannot read {path}: {e}"),
usage: "fill <selector> --file <path>",
})?
}
Some("--stdin") => {
use std::io::Read;
let mut buf = String::new();
io::stdin()
.read_to_string(&mut buf)
.map_err(|e| ParseError::InvalidValue {
message: format!("fill --stdin: {e}"),
usage: "fill <selector> --stdin",
})?;
buf
}
_ => rest[1..].join(" "),
};
Ok(json!({ "id": id, "action": "fill", "selector": sel, "value": value }))
}
"type" => {
// `--key-events` (alias `--keys`): send real per-character keystrokes
@@ -1006,11 +1035,55 @@ fn parse_command_inner(args: &[String], flags: &Flags) -> Result<Value, ParseErr
// path: file path (contains / or . or ends with known extension)
let mut full_page = false;
let mut clip: Option<Value> = None;
let mut max_width: Option<u32> = None;
let mut max_height: Option<u32> = None;
let mut scale: Option<f64> = None;
let mut positional: Vec<&str> = Vec::new();
let mut i = 0;
// Parse a numeric value for a downscale flag (issue #42).
let parse_num = |i: &mut usize, flag: &str| -> Result<String, ParseError> {
let v = rest
.get(*i + 1)
.ok_or_else(|| ParseError::MissingArguments {
context: format!("screenshot {flag}"),
usage: "screenshot [--max-width <px>] [--max-height <px>] [--scale <0..1>]",
})?;
*i += 1;
Ok(v.to_string())
};
while i < rest.len() {
match rest[i] {
"--full" | "-f" => full_page = true,
// Downscale the saved image so retina/full-page shots fit an
// agent's image reader and screenshot px line up with click px (#42).
"--max-width" => {
let v = parse_num(&mut i, "--max-width")?;
max_width = Some(v.parse().map_err(|_| ParseError::InvalidValue {
message: format!("--max-width expects a number, got '{v}'"),
usage: "screenshot --max-width <px>",
})?);
}
"--max-height" => {
let v = parse_num(&mut i, "--max-height")?;
max_height = Some(v.parse().map_err(|_| ParseError::InvalidValue {
message: format!("--max-height expects a number, got '{v}'"),
usage: "screenshot --max-height <px>",
})?);
}
"--scale" => {
let v = parse_num(&mut i, "--scale")?;
let s: f64 = v.parse().map_err(|_| ParseError::InvalidValue {
message: format!("--scale expects a number like 0.5, got '{v}'"),
usage: "screenshot --scale <0..1>",
})?;
if s <= 0.0 || s > 1.0 {
return Err(ParseError::InvalidValue {
message: format!("--scale must be in (0, 1], got '{v}'"),
usage: "screenshot --scale <0..1>",
});
}
scale = Some(s);
}
// `--clip x,y,w,h` captures a pixel region (issue #34).
"--clip" => {
let raw = rest
@@ -1073,6 +1146,15 @@ fn parse_command_inner(args: &[String], flags: &Flags) -> Result<Value, ParseErr
if let Some(c) = clip {
cmd["clip"] = c;
}
if let Some(w) = max_width {
cmd["maxWidth"] = json!(w);
}
if let Some(h) = max_height {
cmd["maxHeight"] = json!(h);
}
if let Some(s) = scale {
cmd["scale"] = json!(s);
}
if let Some(ref fmt) = flags.screenshot_format {
cmd["format"] = json!(fmt);
}
+87 -2
View File
@@ -3114,7 +3114,40 @@ async fn handle_screenshot(cmd: &Value, state: &mut DaemonState) -> Result<Value
)
.await?;
// Downscale the saved image so retina/full-page shots fit an agent's image
// reader and screenshot pixels line up with `click x y` CSS px (issue #42).
// Explicit --scale / --max-width / --max-height win; otherwise a default cap
// (2000px longest edge, AGENT_BROWSER_SCREENSHOT_MAX_EDGE overrides, 0 = off)
// applies. Annotated shots are left untouched so ref overlays stay aligned.
let mut resized: Option<(u32, u32)> = None;
if !annotate {
let scale = cmd.get("scale").and_then(|v| v.as_f64());
let max_w = cmd
.get("maxWidth")
.and_then(|v| v.as_u64())
.map(|v| v as u32);
let max_h = cmd
.get("maxHeight")
.and_then(|v| v.as_u64())
.map(|v| v as u32);
let default_edge = if scale.is_none() && max_w.is_none() && max_h.is_none() {
std::env::var("AGENT_BROWSER_SCREENSHOT_MAX_EDGE")
.ok()
.and_then(|s| s.parse::<u32>().ok())
.or(Some(2000))
.filter(|&e| e > 0)
} else {
None
};
resized = downscale_screenshot(&result.path, scale, max_w, max_h, default_edge);
}
let mut response = json!({ "path": absolutize_saved_path(&result.path) });
if let Some((w, h)) = resized {
response["width"] = json!(w);
response["height"] = json!(h);
response["resized"] = json!(true);
}
if !result.annotations.is_empty() {
response["annotations"] = serde_json::to_value(&result.annotations)
.map_err(|e| format!("Failed to serialize annotations: {}", e))?;
@@ -3130,6 +3163,56 @@ async fn handle_screenshot(cmd: &Value, state: &mut DaemonState) -> Result<Value
Ok(response)
}
/// Downscale a saved screenshot in place (issue #42). Resolves the target longest
/// edge from `scale` (fraction of current), explicit `max_w`/`max_h` caps, or a
/// `default_edge` cap — whichever yields the smaller image. Only ever shrinks;
/// no-op (returns None) if the image is already within bounds or can't be read.
/// Returns the new (width, height) when it actually resized.
fn downscale_screenshot(
path: &str,
scale: Option<f64>,
max_w: Option<u32>,
max_h: Option<u32>,
default_edge: Option<u32>,
) -> Option<(u32, u32)> {
let img = image::open(path).ok()?;
let (w, h) = (img.width(), img.height());
if w == 0 || h == 0 {
return None;
}
// Collect candidate scale factors (≤ 1.0); the smallest wins.
let mut factor = 1.0f64;
if let Some(s) = scale {
factor = factor.min(s);
}
if let Some(mw) = max_w {
if w > mw {
factor = factor.min(mw as f64 / w as f64);
}
}
if let Some(mh) = max_h {
if h > mh {
factor = factor.min(mh as f64 / h as f64);
}
}
if let Some(edge) = default_edge {
let longest = w.max(h);
if longest > edge {
factor = factor.min(edge as f64 / longest as f64);
}
}
if factor >= 1.0 {
return None; // already within bounds — never upscale
}
let nw = ((w as f64 * factor).round() as u32).max(1);
let nh = ((h as f64 * factor).round() as u32).max(1);
let resized = img.resize(nw, nh, image::imageops::FilterType::Lanczos3);
resized.save(path).ok()?;
Some((resized.width(), resized.height()))
}
async fn handle_click(cmd: &Value, state: &mut DaemonState) -> Result<Value, String> {
// First-class coordinate click (issue #8.4): click a raw viewport point with
// no element resolution. Parsed from `click <x> <y>` / `click --coords x,y`.
@@ -3288,7 +3371,7 @@ async fn handle_fill(cmd: &Value, state: &mut DaemonState) -> Result<Value, Stri
let mgr = state.browser.as_ref().ok_or("Browser not launched")?;
let session_id = mgr.active_session_id()?.to_string();
interaction::fill(
let engine = interaction::fill(
&mgr.client,
&session_id,
&state.ref_map,
@@ -3297,7 +3380,9 @@ async fn handle_fill(cmd: &Value, state: &mut DaemonState) -> Result<Value, Stri
&state.iframe_sessions,
)
.await?;
Ok(json!({ "filled": selector }))
// Echo the input path used (input/contenteditable/codemirror5/monaco/select)
// so the agent can confirm a rich editor was handled, not silently no-op'd (#41).
Ok(json!({ "filled": selector, "engine": engine }))
}
async fn handle_type(cmd: &Value, state: &mut DaemonState) -> Result<Value, String> {
+30 -4
View File
@@ -1529,9 +1529,27 @@ pub async fn get_element_input_value(
.send_command_typed(
"Runtime.callFunctionOn",
&CallFunctionOnParams {
function_declaration:
"function() { return typeof this.value === 'string' ? this.value : ''; }"
.to_string(),
// Read rich-editor content too (issue #41): CodeMirror 5 / Monaco
// keep their text in a model, not `.value`; contenteditable keeps
// it as innerText. Falls back to `.value` for plain inputs.
function_declaration: r#"function() {
const el = this;
const cm5 = el.closest && el.closest('.CodeMirror');
if (cm5 && cm5.CodeMirror) return cm5.CodeMirror.getValue();
if (window.monaco && monaco.editor) {
try {
const eds = monaco.editor.getEditors ? monaco.editor.getEditors() : [];
const ed = eds.find(e => e.getDomNode && e.getDomNode().contains(el)) || eds[0];
if (ed) return ed.getValue();
const m = monaco.editor.getModels ? monaco.editor.getModels() : [];
if (m[0]) return m[0].getValue();
} catch (e) {}
}
if (typeof el.value === 'string') return el.value;
if (el.isContentEditable) return el.innerText;
return '';
}"#
.to_string(),
object_id: Some(object_id),
arguments: None,
return_by_value: Some(true),
@@ -1609,7 +1627,15 @@ pub async fn get_element_bounding_box(
&CallFunctionOnParams {
function_declaration: r#"function() {
const r = this.getBoundingClientRect();
return { x: r.x, y: r.y, width: r.width, height: r.height };
const inViewport = r.bottom > 0 && r.right > 0
&& r.top < (innerHeight || document.documentElement.clientHeight)
&& r.left < (innerWidth || document.documentElement.clientWidth);
return {
x: r.x, y: r.y, width: r.width, height: r.height,
centerX: Math.round(r.x + r.width / 2),
centerY: Math.round(r.y + r.height / 2),
inViewport,
};
}"#
.to_string(),
object_id: Some(object_id),
+52 -17
View File
@@ -579,7 +579,7 @@ pub async fn fill(
selector_or_ref: &str,
value: &str,
iframe_sessions: &HashMap<String, String>,
) -> Result<(), String> {
) -> Result<String, String> {
let (object_id, effective_session_id) = resolve_element_object_id(
client,
session_id,
@@ -590,14 +590,15 @@ pub async fn fill(
.await?;
// Emulate a real edit so framework-controlled inputs (React/Vue) and
// site-side listeners actually see the change (issue #25): the old path set
// `this.value` directly and used Input.insertText, which left React's
// internal value-tracker out of sync and never fired change/blur — so
// dependent logic (e.g. Mercari's postal-code → 都道府県 autocomplete) never
// ran even though the value was visible. Set the value through the element's
// PROTOTYPE setter (which React's _valueTracker hooks), then dispatch
// input → change → blur/focusout. `type <sel> <text>` remains for sites that
// need per-keystroke events.
// site-side listeners actually see the change (issue #25): set the value
// through the element's PROTOTYPE setter (which React's _valueTracker hooks),
// then dispatch input → changeblur/focusout. Beyond plain inputs, detect
// rich editors and use their own API/events (issue #41): CodeMirror 5 and
// Monaco have a model that `.value`/`textContent` can't touch; ProseMirror /
// contenteditable need `execCommand('insertText')` so beforeinput/input fire
// (a raw `textContent =` corrupts PM's doc and skips React composers).
// Returns the engine used so the caller can report it. `type <sel> <text>`
// remains for sites that need per-keystroke events.
let fill_js = format!(
r#"function() {{
const el = this;
@@ -605,13 +606,43 @@ pub async fn fill(
try {{ el.focus(); }} catch (e) {{}}
const tag = el.tagName;
const fire = (type, ctor) => el.dispatchEvent(new (ctor || Event)(type, {{ bubbles: true }}));
if (tag === 'SELECT') {{
el.value = v; fire('input'); fire('change'); return true;
// CodeMirror 5: a hidden <textarea> inside .CodeMirror with a live instance.
const cm5 = el.closest && el.closest('.CodeMirror');
if (cm5 && cm5.CodeMirror) {{ cm5.CodeMirror.setValue(v); return 'codemirror5'; }}
// Monaco: global `monaco`; prefer the editor whose DOM contains el.
if (window.monaco && monaco.editor) {{
try {{
const eds = monaco.editor.getEditors ? monaco.editor.getEditors() : [];
const ed = eds.find(e => e.getDomNode && e.getDomNode().contains(el)) || eds[0];
if (ed) {{ ed.setValue(v); return 'monaco'; }}
const models = monaco.editor.getModels ? monaco.editor.getModels() : [];
if (models[0]) {{ models[0].setValue(v); return 'monaco'; }}
}} catch (e) {{}}
}}
if (tag === 'SELECT') {{ el.value = v; fire('input'); fire('change'); return 'select'; }}
if (el.isContentEditable) {{
el.textContent = v; fire('input', window.InputEvent || Event); fire('change');
try {{ el.blur(); }} catch (e) {{}} fire('focusout'); return true;
// ProseMirror / contenteditable: select-all then insertText fires
// beforeinput/input that PM and React composers listen for.
let ok = false;
try {{
const sel = window.getSelection();
const range = document.createRange();
range.selectNodeContents(el);
sel.removeAllRanges();
sel.addRange(range);
ok = document.execCommand('insertText', false, v);
}} catch (e) {{}}
if (!ok) {{ el.textContent = v; fire('input', window.InputEvent || Event); }}
fire('change');
try {{ el.blur(); }} catch (e) {{}}
fire('focusout');
return ok ? 'contenteditable' : 'contenteditable-fallback';
}}
const proto = tag === 'TEXTAREA' ? window.HTMLTextAreaElement.prototype
: window.HTMLInputElement.prototype;
const desc = Object.getOwnPropertyDescriptor(proto, 'value');
@@ -623,13 +654,13 @@ pub async fn fill(
fire('change');
try {{ el.blur(); }} catch (e) {{}}
fire('focusout'); // blur-triggered lookups/validation
return true;
return 'input';
}}"#,
val = serde_json::to_string(value).unwrap_or_default()
);
client
.send_command_typed::<_, Value>(
let result: EvaluateResult = client
.send_command_typed(
"Runtime.callFunctionOn",
&CallFunctionOnParams {
function_declaration: fill_js,
@@ -642,7 +673,11 @@ pub async fn fill(
)
.await?;
Ok(())
Ok(result
.result
.value
.and_then(|v| v.as_str().map(String::from))
.unwrap_or_else(|| "input".to_string()))
}
#[allow(clippy::too_many_arguments)]
+41 -5
View File
@@ -489,6 +489,17 @@ pub fn print_response_with_opts(resp: &Response, action: Option<&str>, opts: &Ou
println!("y: {}", y);
println!("width: {}", w);
println!("height: {}", h);
if let (Some(cx), Some(cy)) = (
obj.get("centerX").and_then(|v| v.as_i64()),
obj.get("centerY").and_then(|v| v.as_i64()),
) {
// Echoed in click-ready CSS px so the agent can paste straight
// into `click <centerX> <centerY>` (issue #43).
println!("center: {} {}", cx, cy);
}
if let Some(iv) = obj.get("inViewport").and_then(|v| v.as_bool()) {
println!("inViewport: {}", iv);
}
}
return;
}
@@ -1490,9 +1501,20 @@ Examples:
chrome-use fill - Clear and fill an input field
Usage: chrome-use fill <selector> <text>
chrome-use fill <selector> --file <path>
chrome-use fill <selector> --stdin
Clears the input field and fills it with the specified text.
This replaces any existing content in the field.
Clears the field and fills it with the text, replacing existing content.
Works on rich editors too (issue #41): CodeMirror 5, Monaco, ProseMirror and
plain contenteditable are detected and set via their own API / input events,
not a raw `.value` write and the response echoes which `engine` was used.
For framework inputs (React/Vue/Angular) the value goes through the native
setter so the form registers it (no more "pristine" Save no-ops).
Options:
--file <path> Read the value from a UTF-8 file (large/multiline text,
backticks/quotes/newlines/non-ASCII no shell escaping)
--stdin Read the value from stdin
Global Options:
--json Output as JSON
@@ -1501,7 +1523,8 @@ Global Options:
Examples:
chrome-use fill "#email" "user@example.com"
chrome-use fill @e3 "Hello World"
chrome-use fill "input[name='search']" "query"
chrome-use fill ".CodeMirror" --file ./article.md # set a CodeMirror editor
cat post.md | chrome-use fill @e7 --stdin
"##
}
"type" => {
@@ -1903,6 +1926,13 @@ Options:
--full, -f Capture full page (not just viewport)
[selector] Capture just an element (CSS or @ref), e.g. `screenshot ".header" h.png`
--clip <x,y,w,h> Capture a pixel region, e.g. `screenshot --clip 0,0,200,40 corner.png`
--max-width <px> Downscale so the image's width px (preserves aspect)
--max-height <px> Downscale so the image's height px
--scale <0..1> Downscale by a factor, e.g. 0.5 (DPR-1, so screenshot px
line up 1:1 with `click x y`)
Default: capped at 2000px longest edge unless overridden
(AGENT_BROWSER_SCREENSHOT_MAX_EDGE; 0 disables). Annotated
shots are never downscaled, so ref overlays stay aligned.
--annotate Overlay numbered labels on interactive elements.
Each label [N] corresponds to ref @eN from snapshot.
Prints a legend mapping labels to element roles/names.
@@ -1925,6 +1955,8 @@ Examples:
chrome-use screenshot --full ./full-page.png
chrome-use screenshot ".header .indicator" corner.png # just one element
chrome-use screenshot --clip 1600,0,200,40 corner.png # a pixel region
chrome-use screenshot --scale 0.5 ./half.png # DPR-1: screenshot px == click px
chrome-use screenshot --max-width 1400 ./shot.png # cap width for image readers
chrome-use screenshot --annotate # Labeled screenshot + legend
chrome-use screenshot --annotate ./page.png # Save annotated screenshot
chrome-use screenshot --annotate --json # JSON output with annotations
@@ -3263,7 +3295,8 @@ Core Commands:
click <sel|x y> Click element/@ref, or a viewport coordinate
dblclick <sel> Double-click element
type <sel> <text> Type into element
fill <sel> <text> Clear and fill
fill <sel> <text> Clear and fill (handles CodeMirror/Monaco/ProseMirror/
contenteditable; `--file <path>`/`--stdin` for large text)
press <key> [--hold <ms>] Press key (Enter, Tab, Control+a). --hold keeps it
down <ms> then releases precise (in-daemon), for
games/charge: `press d --hold 800`
@@ -3283,7 +3316,8 @@ Core Commands:
scroll <dir> [px] Scroll (up/down/left/right)
scrollintoview <sel> Scroll element into view
wait <sel|ms> Wait for element or time
screenshot [path] Take screenshot
screenshot [path] Take screenshot (auto-downscaled to 2000px long edge;
--max-width/--max-height/--scale to override)
pdf <path> Save as PDF
snapshot Accessibility tree with refs (for AI)
eval <js> Run JavaScript
@@ -3297,6 +3331,8 @@ Navigation:
Get Info: chrome-use get <what> [selector]
text, html, value, attr <name>, title, url, count, box, styles, cdp-url
box <sel> x,y,width,height,centerX,centerY,inViewport in CSS px (feed
centerX/centerY into `click x y`); value reads CodeMirror/Monaco too
text (no selector = whole page, all frames), text --main, frames (list)
Check State: chrome-use is <what> <selector>
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "chrome-use",
"version": "1.5.20",
"version": "1.5.21",
"description": "chrome-use — drive your real, logged-in Chrome from any AI agent, stealth by default",
"type": "module",
"packageManager": "pnpm@11.1.3",
+18
View File
@@ -59,6 +59,18 @@ next ref interaction.
> anyway.) Driving off pixels on the relay also risks a coordinate event drifting
> onto the user's foreground tab — refs never do. See issue #37.
> **Two different intents — only one is discouraged.** The rule above is about
> *screenshot-to-locate* (using a picture to find/hit an element) — that's the bug.
> *screenshot-to-capture* — saving a region or element to a file as a **reusable
> image asset** (maps, charts, og-images, visual-diff baselines, report figures) —
> is fully supported and encouraged: `screenshot [selector] [--clip x,y,w,h] <file>`.
> Capturing a rendered map region to a PNG for a blog post is the right tool, not a
> smell. Screenshots are auto-downscaled to ≤2000px (longest edge) so they fit an
> image reader and their pixels line up with `click x y`; override with
> `--max-width`/`--max-height`/`--scale`. To click something you couldn't hit by
> ref, `box @ref` gives the element's CSS-px box + `centerX/centerY` to feed
> straight into `click <centerX> <centerY>` — no screenshot needed.
## Before you automate: pick the cheapest tool
Driving a browser is the heavy option. chrome-use earns its keep when you
@@ -373,6 +385,12 @@ foreground, so prefer refs. For below-the-fold content in such a frame, scroll i
with `scroll down N --at x,y` (a pixel over the frame) or `--frame n`. For a
postal/autocomplete box inside the frame, `type @e "…" --key-events`.
> **Caveat: `find text "…"` can't reach into a cross-origin iframe** — it errors
> "Element not found" even though `snapshot -i` lists those nodes and
> `get text` reads them. Inside cross-origin iframes, target elements by their
> **snapshot `@ref`**, not by `find`. (`box @ref` also works on iframe refs when
> you need a coordinate fallback.)
### When refs don't work or you don't want to snapshot
Use semantic locators: