feat(site): auto-sync + auto-suggest adapters (auto-trigger)
Release binaries / Build macOS ARM64 (push) Has been cancelled
Release binaries / Build macOS x64 (push) Has been cancelled
Release binaries / Build Linux ARM64 (push) Has been cancelled
Release binaries / Build Linux musl ARM64 (push) Has been cancelled
Release binaries / Build Linux musl x64 (push) Has been cancelled
Release binaries / Build Linux x64 (push) Has been cancelled
Release binaries / Build Windows x64 (push) Has been cancelled
Release binaries / Attach binaries to GitHub Release (push) Has been cancelled

Make `site` trigger itself so an agent doesn't have to know adapters exist.

Auto-sync: the pack refreshes on first use and on a TTL (default 7d), both in
the `site` command path (blocking, fast) and as a non-blocking background task
on daemon startup — so ~/.chrome-use/sites/.index.json is always populated with
zero added latency. Tune via AGENT_BROWSER_SITES_TTL_DAYS; disable with
AGENT_BROWSER_SITES_NO_AUTO_UPDATE=1. `update` now writes .last_update + a
domain→adapters .index.json (read-only adapters ordered first).

Auto-suggest: `open`/`navigate`/`snapshot` onto a domain with adapters attaches
`siteAdapters: {domain, commands}` to the response; the CLI prints a
`💡 site adapters for <domain>` hint (stderr) and the field rides along in --json.
SKILL.md tells the agent to prefer the listed `site <name>/<cmd>` over scraping.
This keeps the 'never auto-disrupt user tabs' guarantee — it suggests, the agent
decides; nothing auto-runs on navigation.

site.rs: needs_refresh/adapters_for_domain/write_domain_index + timestamp/index
in update(). daemon.rs: background bootstrap. actions.rs: with_site_hint on
navigate + snapshot. output.rs: hint render. Verified live: open github.com →
hint leads with read-only github/issues; --json carries siteAdapters.
This commit is contained in:
leeguooooo
2026-06-17 17:15:34 +09:00
parent d81bc01645
commit 50b27ac0e0
11 changed files with 249 additions and 6 deletions
+40 -3
View File
@@ -2478,7 +2478,10 @@ async fn handle_navigate(cmd: &Value, state: &mut DaemonState) -> Result<Value,
wb.navigate(url).await?;
let new_url = wb.get_url().await.unwrap_or_else(|_| url.to_string());
let title = wb.get_title().await.unwrap_or_default();
return Ok(json!({ "url": new_url, "title": title }));
return Ok(with_site_hint(
json!({ "url": new_url, "title": title }),
url,
));
}
}
@@ -2545,7 +2548,7 @@ async fn handle_navigate(cmd: &Value, state: &mut DaemonState) -> Result<Value,
.unwrap_or(false)
{
if let Ok(Some(switched)) = mgr.reuse_tab_for_url(url).await {
return Ok(switched);
return Ok(with_site_hint(switched, url));
}
}
@@ -2553,7 +2556,34 @@ async fn handle_navigate(cmd: &Value, state: &mut DaemonState) -> Result<Value,
// Adaptive humanize: sample the freshly loaded page for known behavioural
// anti-bot vendors and escalate this session to Human if any are present.
detect_and_set_humanize(mgr).await;
Ok(result)
Ok(with_site_hint(result, url))
}
/// Annotate a navigation/snapshot result with the `site` adapters available for
/// the page's domain (auto-trigger): when you land on e.g. github.com, the
/// response carries `siteAdapters: { domain, commands: ["github/issues", …] }` so
/// the agent reaches for a structured-data adapter instead of scraping. No-op
/// when nothing matches or the pack isn't synced yet.
fn with_site_hint(mut result: Value, fallback_url: &str) -> Value {
let url = result
.get("url")
.and_then(|v| v.as_str())
.unwrap_or(fallback_url);
let host = url::Url::parse(url)
.ok()
.and_then(|u| u.host_str().map(String::from));
if let Some(host) = host {
let adapters = crate::site::adapters_for_domain(&host);
if !adapters.is_empty() {
if let Some(obj) = result.as_object_mut() {
obj.insert(
"siteAdapters".to_string(),
json!({ "domain": host, "commands": adapters }),
);
}
}
}
result
}
/// After navigation, probe the page for known anti-bot vendor fingerprints
@@ -2936,6 +2966,13 @@ async fn handle_snapshot(cmd: &Value, state: &mut DaemonState) -> Result<Value,
let ref_count = refs.len();
let mut out = json!({ "snapshot": tree, "origin": url, "refs": refs });
// Auto-trigger: if this domain has site adapters, surface them so the agent
// pulls structured data instead of walking the tree. `with_site_hint` reads
// the `url` field, so pass it under that key.
if let Some(hint) = with_site_hint(json!({ "url": url }), &url).get("siteAdapters") {
out["siteAdapters"] = hint.clone();
}
// Canvas/WebGL apps (games, map/3D viewers, drawing tools) paint to a
// <canvas> and expose almost no accessibility tree, so `snapshot` comes back
// near-empty and agents get stuck looking for refs that will never exist
+10
View File
@@ -21,6 +21,16 @@ pub async fn run_daemon(session: &str) {
// (via the ab-connect extension) land in a per-session Chrome tab group.
let _ = super::browser::DAEMON_SESSION.set(session.to_string());
// Bootstrap / refresh the site-adapter pack in the background (first-run +
// periodic TTL). This populates ~/.chrome-use/sites/.index.json so navigation
// can auto-suggest `site` commands for the page you land on, with zero added
// latency to any command. Best-effort; offline is a no-op.
if crate::site::needs_refresh() {
tokio::spawn(async {
let _ = crate::site::update().await;
});
}
let socket_dir = get_daemon_socket_dir();
if !socket_dir.exists() {
let _ = fs::create_dir_all(&socket_dir);