{"id":"a49246cf-d69a-47eb-9644-373da48c6c0d","entityType":"agent","slug":"clawhub-yicko-playwright-browser-use","name":"playwright-browser-use","canonicalUrl":"https://www.xpersona.co/agent/clawhub-yicko-playwright-browser-use","canonicalPath":"/agent/clawhub-yicko-playwright-browser-use","generatedAt":"2026-10-10T14:43:24.597Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T12:29:31.512Z","emptyReason":null},"description":"浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会将其与代码执行一并禁用）；(2) `eval` 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) `run Skill: playwright-browser-use Owner: yicko Summary: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— cookies / storage 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 PW_BROWSER_SAFE_MODE=1 会将其与代码执行一并禁用）；(2) eval 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) run Tags: latest:1.0.16 Version history: v1.0.16 | 2026-08-01T16:07:30.572Z | us","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s171gt9hm8qq7y2qtg3p5f639n84y0tw:playwright-browser-use","sourceUrl":"https://clawhub.ai/yicko/playwright-browser-use","homepage":"https://clawhub.ai/yicko/skills/playwright-browser-use","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/yicko/playwright-browser-use","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/yicko/skills/playwright-browser-use","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":63,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:29:31.512Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:29:31.512Z","emptyReason":null},"stars":null,"forks":null,"downloads":1432,"packageName":null,"latestVersion":"1.0.16","tractionLabel":"1.4K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:29:31.444Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T12:29:31.512Z","lastCrawledAt":"2026-10-10T12:29:31.444Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T12:29:31.444Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.16","createdAt":"2026-08-01T16:07:30.572Z","changelog":"## playwright-browser-use 1.0.16 Changelog - Removed redundant documentation file: `skill-card.md` - No functional changes; all core features and behaviors remain unchanged.","fileCount":19,"zipByteSize":77782},{"version":"1.0.15","createdAt":"2026-07-30T14:14:57.405Z","changelog":"- skill-card.md 文件已移除，不再随版本分发。 - 无其他变更。","fileCount":19,"zipByteSize":83945},{"version":"1.0.14","createdAt":"2026-07-28T01:32:26.248Z","changelog":"- 删除 skill-card.md 文件。 - 无功能变化，本更新仅为移除多余文档文件，未影响核心功能或稳定性。","fileCount":19,"zipByteSize":65819},{"version":"1.0.13","createdAt":"2026-07-27T03:44:55.921Z","changelog":"- Removed the sample file skill-card.md. - No functional or behavioral changes to the skill itself. - Documentation and core usage remain unchanged.","fileCount":19,"zipByteSize":62256},{"version":"1.0.12","createdAt":"2026-07-26T20:42:39.000Z","changelog":"- Switched dependency from Playwright full to playwright-core; no browser download required, now uses system Chrome/Edge only. - Added English documentation files: README.en.md and QUICKSTART.en.md for non-Chinese users. - Added CHANGELOG.md and MANIFEST.txt for versioning and packaging clarity. - Removed obsolete skill-card.md, consolidating documentation. - Updated SKILL.md to clarify dependency, browser detection, and availability of English docs.","fileCount":19,"zipByteSize":59837},{"version":"1.0.11","createdAt":"2026-07-25T16:37:00.884Z","changelog":"## Playwright-browser-use 1.0.11 Changelog - Added `package-lock.json` to the repository. - Removed `skill-card.md` file. - SKILL.md: Documented new `capabilities` section outlining supported features. - No core logic or runtime changes; updates are for documentation and repository metadata only.","fileCount":14,"zipByteSize":31583},{"version":"1.0.10","createdAt":"2026-07-25T05:38:34.896Z","changelog":"- 添加全面的安全声明：详细说明 `eval`（页面 JS 执行）与 `run-code`（Node/Playwright 沙箱代码执行）能力、攻击面、token 认证机制，以及通过 `PW_BROWSER_SAFE_MODE=1` 可强制禁用代码执行 - 更新 skill 描述，突出代码执行特性、守护进程架构、token 认证保护及安全模式 - 明确命令认证措施和适用边界，提示所有命令均需随机 token 验证，默认仅本地可信进程可用 - 新增「文档语言与本地化说明」小节，提示命令用法与可见文本的多语言适配问题 - 移除无关 skill-card.md，保持文档精简","fileCount":13,"zipByteSize":30395},{"version":"1.0.9","createdAt":"2026-07-25T05:00:24.371Z","changelog":"- 移除了 skill-card.md 文件，无功能性代码变更 - SKILL.md 描述大幅升级：强调 `eval`、`run-code` 两项代码执行能力、token 认证、`PW_BROWSER_SAFE_MODE` 安全选项和页面/节点隔离边界 - 明确说明 eval/run-code 的攻击面、认证机制和环境变量隔离，提升安全警示与用法风险提示 - 文档新增“文档语言与本地化说明”与多语言适配建议 - 其余交互规范、命令速查和恢复指南未更动，仅补充和强化原安全声明及适用场景","fileCount":13,"zipByteSize":30387}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171gt9hm8qq7y2qtg3p5f639n84y0tw:playwright-browser-use","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T14:43:24.591Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-yicko-playwright-browser-use/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T12:29:31.512Z","emptyReason":null},"readme":"Skill: playwright-browser-use\n\nOwner: yicko\n\nSummary: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会将其与代码执行一并禁用）；(2) `eval` 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) `run\n\nTags: latest:1.0.16\n\nVersion history:\n\nv1.0.16 | 2026-08-01T16:07:30.572Z | user\n\n## playwright-browser-use 1.0.16 Changelog\n\n- Removed redundant documentation file: `skill-card.md`\n- No functional changes; all core features and behaviors remain unchanged.\n\nv1.0.15 | 2026-07-30T14:14:57.405Z | user\n\n- skill-card.md 文件已移除，不再随版本分发。\n- 无其他变更。\n\nv1.0.14 | 2026-07-28T01:32:26.248Z | user\n\n- 删除 skill-card.md 文件。\n- 无功能变化，本更新仅为移除多余文档文件，未影响核心功能或稳定性。\n\nv1.0.13 | 2026-07-27T03:44:55.921Z | user\n\n- Removed the sample file skill-card.md.\n- No functional or behavioral changes to the skill itself.\n- Documentation and core usage remain unchanged.\n\nv1.0.12 | 2026-07-26T20:42:39.000Z | user\n\n- Switched dependency from Playwright full to playwright-core; no browser download required, now uses system Chrome/Edge only.\n- Added English documentation files: README.en.md and QUICKSTART.en.md for non-Chinese users.\n- Added CHANGELOG.md and MANIFEST.txt for versioning and packaging clarity.\n- Removed obsolete skill-card.md, consolidating documentation.\n- Updated SKILL.md to clarify dependency, browser detection, and availability of English docs.\n\nv1.0.11 | 2026-07-25T16:37:00.884Z | user\n\n## Playwright-browser-use 1.0.11 Changelog\n\n- Added `package-lock.json` to the repository.\n- Removed `skill-card.md` file.\n- SKILL.md: Documented new `capabilities` section outlining supported features.\n- No core logic or runtime changes; updates are for documentation and repository metadata only.\n\nv1.0.10 | 2026-07-25T05:38:34.896Z | user\n\n- 添加全面的安全声明：详细说明 `eval`（页面 JS 执行）与 `run-code`（Node/Playwright 沙箱代码执行）能力、攻击面、token 认证机制，以及通过 `PW_BROWSER_SAFE_MODE=1` 可强制禁用代码执行\n- 更新 skill 描述，突出代码执行特性、守护进程架构、token 认证保护及安全模式\n- 明确命令认证措施和适用边界，提示所有命令均需随机 token 验证，默认仅本地可信进程可用\n- 新增「文档语言与本地化说明」小节，提示命令用法与可见文本的多语言适配问题\n- 移除无关 skill-card.md，保持文档精简\n\nv1.0.9 | 2026-07-25T05:00:24.371Z | user\n\n- 移除了 skill-card.md 文件，无功能性代码变更\n- SKILL.md 描述大幅升级：强调 `eval`、`run-code` 两项代码执行能力、token 认证、`PW_BROWSER_SAFE_MODE` 安全选项和页面/节点隔离边界\n- 明确说明 eval/run-code 的攻击面、认证机制和环境变量隔离，提升安全警示与用法风险提示\n- 文档新增“文档语言与本地化说明”与多语言适配建议\n- 其余交互规范、命令速查和恢复指南未更动，仅补充和强化原安全声明及适用场景\n\nv1.0.8 | 2026-07-24T07:01:36.359Z | user\n\n- Removed the sample skill-card.md file.\n- No changes to functionality or user experience.\n\nv1.0.7 | 2026-07-23T19:35:13.520Z | user\n\nVersion 1.0.6\n\n- Added LICENSE file for clear licensing information.\n- Added README.md to provide usage instructions and documentation.\n\nv1.0.6 | 2026-07-23T19:19:38.684Z | user\n\n- Removed sample file `skill-card.md` from the project.\n- No changes to code or runtime behavior.\n- Documentation and usage remain the same.\n\nv1.0.5 | 2026-07-23T18:53:44.102Z | user\n\n- Removed the sample file skill-card.md.\n- No functional or interface changes to the skill itself.\n- Documentation (SKILL.md) was unchanged, aside from unrelated descriptive or formatting edits.\n\nv1.0.4 | 2026-07-23T18:05:01.282Z | user\n\nNo user-facing changes detected in this version.\n\n- No file changes were detected.\n- Behavior and documentation remain unchanged from the previous version.\n\nv1.0.3 | 2026-07-23T17:53:08.375Z | user\n\n## playwright-browser-use 1.0.3 — Changelog\n\n- **Major documentation update:** SKILL.md has been completely rewritten and expanded.\n- Skill documentation now includes explicit safety boundaries and security scenarios section, clarifying what environments are considered safe for use.\n- Login and human-in-the-loop verification workflows are fully documented, with step-by-step instructions for both users and agent.\n- All usage, workflow, and command examples updated to reflect absolute path conventions and local security measures.\n- Internal metadata file (skill-card.md) has been removed.\n\nv1.0.2 | 2026-07-23T06:06:40.326Z | user\n\n- Removed the file: skill-card.md\n- No changes to functionality, only documentation file cleanup.\n\nv1.0.1 | 2026-07-23T02:00:44.640Z | user\n\nInitial release: Playwright-based browser automation skill, fully independent of DuMate.\n\n- Provides a CLI tool (`pw-browser`) for browser automation using Playwright with Node.js, supporting all key actions like navigation, snap, click, fill, scroll, and tab management.\n- No DuMate agent or extension required—runs standalone, using system-installed Chrome or Edge (via Playwright channel).\n- Includes robust session/daemon management for persistent browser/browser state across commands.\n- Detailed semantic rules and workflow best practices are documented, ensuring reliable usage patterns.\n- Command overview and advanced strategies (wait, snap, run-code, pagination, SPA handling) included.\n\nv1.0.0 | 2026-07-23T01:46:46.164Z | auto\n\nInitial release — brings full Playwright-based browser automation with a standalone Node.js CLI.\n\n- Provides browser automation (open web pages, screenshots, clicks, form filling, navigation, and more) using Playwright, no DuMate environment required.\n- Uses a persistent local daemon for browser state, communicating over HTTP; supports Chrome/Edge via Playwright’s channel mechanism.\n- Exposes a CLI tool (`pw-browser`) with a wide range of commands for page interaction, advanced scripting, tab management, and automation of common patterns like paging and login handling.\n- Offers detailed workflow guidelines, robust error/daemon recovery procedures, and extensive documentation for CLI commands and operational rules.\n\nArchive index:\n\nArchive v1.0.16: 19 files, 77782 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), CHANGELOG.md (2514b), LICENSE (1062b), MANIFEST.txt (458b), package-lock.json (749b), package.json (1123b), pw-browser (266b), pw-browser.js (77857b), QUICKSTART.en.md (9139b), QUICKSTART.md (8662b), README.en.md (16514b), README.md (15399b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5570b), skill-card.md (2905b), SKILL.md (42369b), _meta.json (142b)\n\nFile v1.0.16:SKILL.md\n\n---\nname: playwright-browser-use\ndescription: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会将其与代码执行一并禁用）；(2) `eval` 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) `run-code` 在守护进程上下文执行 Playwright/Node 代码（vm 沙箱隔离）。全部经持久化本地守护进程（127.0.0.1:19223，浏览器状态跨命令保持）控制，受随机 token 认证保护；`PW_BROWSER_SAFE_MODE=1` 可彻底禁用代码执行与 cookies/storage 凭证读写（v1.3.2+）。仅在可信、用户可见的本地环境中授权使用；会话凭证落盘须遵循后文安全警告。\nallowed-tools: Bash(node:*), Bash(pw-browser:*), Bash(curl:*)\ncapabilities:\n  - \"code-execution: page-context (eval — arbitrary JS in current page)\"\n  - \"code-execution: daemon-vm (run-code — null-prototype VM sandbox, no host fs/process)\"\n  - \"network: arbitrary via browser/page context (credentialed requests possible)\"\n  - \"browser-state: persistent credentialed session across commands\"\n  - \"credential-access: direct read/export/import/clear/set of cookies & localStorage (session tokens) — no code execution needed; gated by daemon token, and fully DISABLED by PW_BROWSER_SAFE_MODE=1 (v1.3.2+)\"\n  - \"file: write to local disk (run-code can trigger downloads; cookies/storage export writes credential files)\"\n# --- Formal permission model (machine-readable) -------------------------\n# NOTE: This skill has NO built-in fine-grained permission system. The\n# fields below are declarative metadata for orchestrators/reviewers, NOT\n# an enforcement boundary. Real constraints come only from:\n#   - daemon token auth (gates NETWORK reachability, not per-capability)\n#   - PW_BROWSER_SAFE_MODE=1 (disables code-exec + credential primitives)\n#   - export/import path confinement to ~/.pw-browser (v1.3.1+); since v1.3.8\n#     the caller-supplied --unsafe flag alone can NOT lift it — the operator\n#     must also start the daemon with PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1\n#   - PW_BROWSER_CRED_PERSIST=off (v1.3.9) removes credential persistence only\n#     (export/import), keeping in-memory cookie/storage automation usable\n# \"capabilities\" lists the MINIMUM attack surface, not a maximum; the skill\n# can additionally drive any site the browser can reach, including ones where\n# the user is already authenticated. Least privilege is achieved at the\n# orchestration layer via deploy mode (safe mode + sandbox) + scoped ops.\npermissions:\n  model: \"none-formal\"          # no built-in RBAC / capability-dropping\n  enforcement:\n    - \"daemon-token\"            # gates network reachability to 127.0.0.1:19223\n    - \"safe-mode-env\"           # PW_BROWSER_SAFE_MODE=1 disables code-exec + creds\n    - \"path-confinement\"        # cookies/storage IO limited to ~/.pw-browser\n    - \"operator-gated-override\" # lifting confinement needs daemon env, not a caller flag\n    - \"cred-persist-killswitch\" # PW_BROWSER_CRED_PERSIST=off blocks export/import only\n    - \"secret-file-permissions\" # daemon token + credential dumps written 0600, dir 0700\n    - \"download-name-sanitised\" # page-supplied download filename forced to a basename inside --path\n  least-privilege:\n    default-mode: \"full\"        # full power; requires trusted, user-visible, local\n    scoped-mode: \"cred-persist-off\"      # PW_BROWSER_CRED_PERSIST=off — keeps automation,\n                                         # removes the privilege-persistence primitive\n    reduced-mode: \"safe-mode\"   # PW_BROWSER_SAFE_MODE=1 — for not-fully-trusted agents\n    untrusted-mode: \"safe-mode+sandbox\"  # add network isolation for untrusted input\n  metadata-is: \"attack-surface-description\"  # NOT a grant, NOT a sandbox boundary\ndisable: false\n---\n\n# 浏览器自动化 — Playwright 版（`pw-browser`）\n\n> `pw-browser` 是基于 playwright-core（Playwright 核心库）的浏览器自动化 CLI，仅依赖 Node.js 和 playwright-core，无需下载任何浏览器。\n>\n> ⚠️ **完整能力（含代码执行）**：本工具**不只是\"点开网页\"**——它包含 `eval`（页面上下文**任意 JavaScript** 执行）与 `run-code`（守护进程上下文执行 Playwright/Node 代码）两项代码执行能力，并通过**持久化本地守护进程**控制浏览器状态（状态跨命令保持）。所有代码执行均受 daemon token 认证保护，可用 `PW_BROWSER_SAFE_MODE=1` 彻底禁用。请先阅读下方「⚠️ 安全边界与使用场景」了解完整攻击面与适用边界，再决定是否授权。\n\n## ⚠️ 安全边界与使用场景\n\n本工具为 **本地 AI 助手对话环境**设计，运行在用户完全可视、可中断的场景中。\n\n| 场景 | 风险 | 说明 |\n|------|------|------|\n| AI 助手对话交互（推荐） | 🟡 低 | 用户全程可见浏览器操作，可随时中断 |\n| 本地开发/测试 | 🟡 低 | 在受控环境中操作测试页面 |\n| 手动触发的数据采集 | 🟡 低 | 用户明确指定的页面和操作 |\n\n以下场景 **不推荐**直接使用，需要额外安全措施：\n\n| 场景 | 风险 | 需要的额外措施 |\n|------|------|-------------|\n| 被不可信 agent 调用 | 🔴 高 | 已内置 daemon token 认证 + 可选 `PW_BROWSER_SAFE_MODE` 禁用代码执行与 cookies/storage 凭证原语（v1.3.2+）；若来源仍不可信，应进一步沙箱隔离 |\n| 作为公开 API 服务 | 🔴 极高 | 必须加认证 + 操作白名单 |\n| CI/CD 自动化流水线 | 🟡 中 | 需限定操作范围，禁止生产环境 |\n\n**核心能力声明：**\n- `eval`：在**浏览器上下文**（`page.evaluate`）执行**任意 JavaScript**——即完整的页面级代码执行能力。它无法访问 Node.js API（`require`/`fs`/`process`），作用域仅限于当前页面；但正因如此，它能读取 `document.cookie`/`localStorage`/`sessionStorage`、发起**带页面凭证的 `fetch` 请求**、操控 DOM 并触发页面内动作（点击、提交等）。**这是与 `run-code` 同级的\"代码执行\"能力，仅场景不同（页面 vs Node）**：同样受 daemon token 认证保护，同样在 `PW_BROWSER_SAFE_MODE=1` 下被禁用；只在用户明确指定、且页面可信时使用，不对高权限/来源不明页面执行\n- `run-code`：在 daemon 进程的**受限沙箱**（`vm` 模块）中执行 Playwright 代码，仅暴露 `page` 和安全 JS 全局；**无法直接调用** `fs`/`child_process`/`process`/`require`。但浏览器上下文可经 `download.saveAs`/`setInputFiles` 在本地磁盘读写文件、能发起任意网络请求（沙箱不阻止）。它仍拥有完整浏览器控制权（导航、读写存储、下载、提交表单、改页面内容），**仅限本地信任环境使用**\n- `cookies` / `storage`：**独立的会话凭证读写原语**，无需任何代码执行即可对 cookie 与 `localStorage` 做 `list` / `export` / `import` / `clear` / `set`。它可直接**提取**当前登录态（含 `HttpOnly` cookie、会话/Bearer 令牌、CSRF token），也可**注入**任意攻击者控制的状态，是凭据盗窃与账户接管的**独立高危面**——**不依赖** `eval` / `run-code`。普通模式下仅由 daemon token 认证保护；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会**整体禁用**全部 `cookies` / `storage` 子命令（与 eval/run-code 同级拦截）。落盘会话文件的处置见下方「会话持久化风险（Rogue Agent）」块：务必用完即删、不跨环境/账号复用。`export`/`import` 默认**限制在 `~/.pw-browser/` 目录内**以防凭证散落或加载外部攻击者构造文件；自 v1.3.8 起该限制**不可由调用方自行解除**——越界需**同时**满足「操作者以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon」+「调用方显式传 `--unsafe`」，否则返回 `UnsafeOverrideNotPermitted`\n- daemon：监听 `127.0.0.1:19223`，仅本机可访问，不暴露到公网；**所有命令（除 `/health` 存活探针）均要求随机 `token` 认证**，token 在 daemon 启动时生成并写入 `~/.pw-browser/daemon.json`（默认仅当前用户可读），CLI 自动携带，外部进程无法在未读取该文件的情况下调用\n- 安全模式：设置环境变量 `PW_BROWSER_SAFE_MODE=1` 启动 daemon 可**彻底禁用** `run-code` / `eval` **以及全部 `cookies` / `storage` 凭证原语**（v1.3.2+），仅保留 snap/click/fill 等白名单命令，适合不需要自定义代码、也不该触碰会话凭证的场景（如接入来源不完全可信的 agent）\n- 非 headless 模式：浏览器窗口始终可见，用户可直接监控所有操作\n- 文件下载：通过 `run-code` 触发页面下载（`download.saveAs`）会写入本地磁盘，注意目标路径\n- 下载文件名消毒（v1.3.10+）：`download` 命令的建议文件名由**被自动化的网页**决定而非操作者。当 `--path` 指向**目录**时，该名字会被强制取 `basename` 并校验解析后仍位于该目录内，越界返回 `PathTraversal`；恶意站点无法再用 `../../.bashrc` 之类的名字把文件写到目录之外。若 `--path` 显式指向一个**文件路径**，则按操作者意图原样使用\n\n## 🔐 权限模型与最低特权（形式化说明）\n\n> **关键澄清：`capabilities` 是攻击面描述，不是权限授予，也不是沙箱边界。**\n> frontmatter 里的 `capabilities` / `allowed-tools` 是给编排者与审核者的**自由文本元数据**：\n> - `capabilities` 列出的是本技能**能够做什么**（即真实攻击面的最小集合），**不代表**它被限制在这些能力内、也**不代表**它已获授权——拥有某项能力意味着技能*可以*行使它，而非*只能*行使它。\n> - `allowed-tools` 仅声明技能可调用哪些宿主工具（如 `Bash(node:*)`），**不构成**对技能行为的强制限制；真正的行为约束来自下方\"实际执行边界\"。\n> 审核时请以\"攻击面下限\"而非\"能力上限\"来读 `capabilities`：技能还可驱动浏览器能抵达的**任意站点**，包括用户**已登录**的站点，从而读取/操作该站点的认证状态。\n\n### 分级部署矩阵（最低特权选择）\n\n| 部署模式 | 代码执行 (`eval`/`run-code`) | 凭证原语 (`cookies`/`storage`) | 网络 | 适用场景 | 残余风险 |\n|----------|------------------------------|-------------------------------|------|----------|----------|\n| **全能力（默认）** | ✅ 开启 | ✅ 开启 | 经浏览器（含带凭证请求） | 可信本地、用户全程可见、临时/演示会话 | 高：可导出/注入会话、可发带凭证请求 |\n| **禁凭证落盘** `PW_BROWSER_CRED_PERSIST=off` | ✅ 开启 | ⚠️ 仅内存态（`export`/`import` 返回 `CredentialPersistenceDisabled`） | 经浏览器 | 需要完整自动化能力、但不允许会话状态跨运行留存的长驻 agent | 中高：仍可读取会话内容，但无法把它变成可复用的凭证文件 |\n| **安全模式** `PW_BROWSER_SAFE_MODE=1` | ⛔ 禁用 (返回 `Disabled`) | ⛔ 禁用 (返回 `Disabled`) | 仅 `snap`/`click`/`fill` 等白名单命令，不主动发请求 | 接入来源不完全可信的 agent、CI 只读巡检 | 中：仍可访问已打开页面的 DOM/可见内容 |\n| **沙箱 + 安全模式** | ⛔ 禁用 | ⛔ 禁用 | 额外做**网络隔离**（禁止出网 / 仅内网白名单） | 完全不可信或第三方输入驱动的自动化 | 低：沙箱逃逸前无法外联或触碰凭证 |\n\n**最低特权原则**：默认按\"全能力\"部署仅当满足——① 用户全程可见可中断；② 操作对象为用户明确指定的页面；③ 不在已登录高权限账户（银行/邮箱/工单系统）上执行未确认动作。否则应**优先启用安全模式**，并对不可信来源**进一步沙箱隔离**。\n\n### 实际执行边界（技能到底受什么约束）\n\n1. **daemon token 认证**：仅限制*谁能通过网络抵达 daemon*（须持有启动时生成的随机 token）。一旦本地进程持有 token，**所有能力全部可用**——token 是\"门禁\"而非\"按能力细分的权限\"。\n2. **安全模式（v1.3.2+）**：本技能**唯一内建的能力开关**，整体禁用代码执行与 `cookies`/`storage` 凭证原语；它是降低攻击面的主开关，但不是沙箱。\n3. **路径限制 + 操作者门禁（v1.3.1 / v1.3.8）**：`cookies`/`storage` 的 `export`/`import` 默认被限制在 `~/.pw-browser/` 内（符号链接经 realpath 解析，无法用软链逃逸）。**解除权归属操作者而非调用方**：越界须同时满足「daemon 以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动」+「调用方显式 `--unsafe`」；仅传 `--unsafe` 会被拒（`UnsafeOverrideNotPermitted`），因为调用方不能修改 daemon 进程的环境变量。这只约束落盘位置，不约束浏览器内的读/写行为。每次凭证路径访问都会写入 daemon stderr 审计行。\n4. **凭证持久化开关（v1.3.9）**：`PW_BROWSER_CRED_PERSIST=off` 启动 daemon 可**单独禁用** `cookies`/`storage` 的 `export`/`import`（返回 `CredentialPersistenceDisabled`），而 `list`/`get`/`set`/`clear` 等内存态操作照常可用。它比安全模式**粒度更细**：只切断\"会话状态落盘\"这一条特权持久化路径，保留正常自动化能力。同样是进程环境变量，调用方无法解除。\n5. **落盘文件权限（v1.3.9）**：`~/.pw-browser/` 以 `0700` 创建，daemon 认证 token（`daemon.json`）与导出的凭证文件均以 `0600` 写入（已存在的旧文件会被 chmod 收紧）。这防止**同机其他本地用户**读走 token 接管浏览器、或直接读取导出的会话。注意：POSIX 权限位在 Windows 上不由操作系统强制执行，Windows 下依赖用户目录的 ACL 继承。\n6. **宿主工具策略**：`allowed-tools` 由宿主平台在*调用层*决定是否放行技能发起的工具调用；它依赖平台实现，**不保证**能限制技能在已获准工具内的具体行为（例如在 `Bash(node:*)` 内仍可执行任意 Node 代码）。\n\n> 结论：本技能**没有**独立于上述六点的\"形式化权限系统\"。任何\"最小权限\"诉求都必须通过**部署模式选择（安全模式/沙箱）+ 操作范围约定**在编排层落实，而非依赖技能元数据自证安全。\n\n## 📝 文档语言与本地化说明\n\n- **文档语言**：本技能文档为**简体中文**。若你或下游 agent 的默认语言非中文，请以代码块中的命令、URL、CSS 选择器与 `snap` 返回的 `ref` 为准——这些是**与语言无关**的自动化锚点。\n- **界面文本匹配是启发式的**：识别分页 / 按钮类型时，文档列出的中文、英文关键词（如\"下一页\"/\"Next\"、\"更新\"/\"保存\"/\"发布\"）只是**识别信号示例，并非穷举**；非中文页面的实际文案会不同。\n- **优先用 DOM 锚点，而非可见文字**：跨语言页面请尽量用 `snap` 得到的 `ref` 或 CSS 选择器（`page.locator('.xxx')`）定位元素，避免依赖本地化后的可见文本，以防因文案不同导致误点 / 误操作。\n- **适用区域**：技能本身不限定网站区域；文档示例与中文 UI 关键词面向中文环境，各\"识别信号\"表已并列给出英文界面关键词。\n- **英文文档**：面向非中文 agent/用户，提供 [`README.en.md`](./README.en.md)（英文 README）与 [`QUICKSTART.en.md`](./QUICKSTART.en.md)（英文端到端示例）。`SKILL.md` 本身保持中文，但其内的命令、URL、CSS 选择器、`snap` 返回的 `ref` 均为语言无关锚点，非中文 agent 可直接据此执行。\n\n## 前置条件（首次使用）\n\n本 Skill 所在目录需已执行 `npm install`（将安装 `playwright-core`）。**无需单独下载浏览器** — daemon 启动时自动检测并使用系统的 Chrome 或 Edge。\n\n> 下文所有命令中的 `{SKILL_DIR}` 请替换为实际的 skill 安装目录路径。\n\n## 架构\n\n`pw-browser` 采用 daemon + client 架构：\n\n```\n┌──────────────┐     HTTP (localhost:19223)     ┌──────────────┐\n│  pw-browser  │ ──────────────────────────────→│   Daemon     │\n│  (CLI 客户端) │                                │  (浏览器进程)  │\n└──────────────┘                                └──────┬───────┘\n                                                       │\n                                                       ├─ Playwright\n                                                       ├─ Chromium 浏览器\n                                                       └─ 页面状态持久化\n```\n\n**daemon 启动后持续运行**，浏览器和页面状态跨命令保持。CLI 每次通过 HTTP 调用 daemon。\n\n## 启动 Daemon\n\n**每次会话开始前**，在后台启动 daemon。以下命令使用 Skill 所在目录的绝对路径和 shim 脚本：\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" daemon &\nsleep 4\n```\n\n> daemon 会在 `127.0.0.1:19223` 监听，首次启动会用 Playwright 的 `channel: 'chrome'` 自动连接系统 Chrome 浏览器（如已安装了 Edge 也会尝试）。无需下载额外的 Chromium。\n\n验证 daemon 可用：\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" init\n```\n\n**关闭 daemon：**\n\n```bash\npw-browser close --all\n```\n\n## 核心工作流\n\n> **注意**：下面所有 `pw-browser` 命令都需要设置 `NODE_PATH`。Agent 执行时应使用完整形式：\n> ```bash\n> SKILL_DIR=\"{SKILL_DIR}\"\n> NODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" <cmd> [args] [--json]\n> ```\n> 为简洁起见，下文示例省略前缀，用 `pw-browser` 表示。\n\n```bash\n# 1. 启动 daemon（会话开始一次）\npw-browser daemon &\n\n# 2. 打开页面\npw-browser open https://www.baidu.com\n\n# 3. 获取页面快照（必须！每次交互前都要 snap）\npw-browser snap\n\n# 4. 交互 — 基于快照中的 e0, e1, e2... ref 引用\npw-browser click e8          # 点击 ref=e8 的元素\npw-browser fill e5 \"hello\"   # 在 ref=e5 的输入框填入文本\npw-browser press Enter       # 键盘按键\n\n# 5. 等待\npw-browser wait-for \"text=加载完成\" --timeout 8000\npw-browser wait-for \"url:https://example.com/*\"\npw-browser wait-for \"state:networkidle\"\n\n# 6. Tab 管理\npw-browser tab list\npw-browser tab select 1\npw-browser tab close 0\n\n# 7. 关闭\npw-browser close             # 关闭当前页面\npw-browser close --all       # 关闭浏览器 + daemon\n```\n\n## 语义规则（必须遵守）\n\n### 规则 1：先观察再操作\n\nCLI **不会**在 open/click 后自动获取快照。**以下情况必须执行 `pw-browser snap`**：\n\n- **每次 `open` / `goto` / `go-back` / `go-forward` / `reload` 之后**（页面已变更）\n- **每次点击可能触发导航的元素之后**（链接、提交按钮等）\n- **首次进入陌生页面时**\n\n**同一页面内连续操作**（click → fill → press → click）可复用已获取的 ref，无需每步 snap。\n\n```\n正确: pw-browser open URL → pw-browser snap → pw-browser click e5 → pw-browser fill e8 \"text\"\n错误: pw-browser open URL → pw-browser click e5（open 后缺 snap）\n正确: pw-browser snap → pw-browser click e5 → pw-browser click e7（同一页面连续操作，ref 稳定）\n错误: pw-browser snap → pw-browser click e5 → pw-browser snap → pw-browser click e5（同一页面内多余 snap）\n```\n\n> ⚠️ **何时必须重新 snap：** 页面发生导航（URL 变化、新页面打开）、或 `act` 命令返回了 `interrupted`（DOM 已被修改）。此时旧 ref 可能指向错误元素，必须 re-snap 获取最新快照。\n\n### 规则 2：点击链接后处理导航\n\n点击可能触发导航的链接（`<a>` 标签、按钮等）后：\n1. `pw-browser snap` — 检查页面是否已变化\n2. 如有新 tab → `pw-browser tab list` → `pw-browser tab select <idx>`\n3. `pw-browser snap` — 获取新页面内容\n\n### 规则 3：页面内容不全\n\n如果快照中元素不全（列表不完整等）：\n- `pw-browser mousewheel 0 500` 滚动\n- 或点击\"加载更多\"/\"下一页\"\n- 重新 `pw-browser snap`\n\n### 规则 4：登录与验证码（人机协作）\n\npw-browser 使用非 headless 模式打开实体 Chrome 窗口，用户可直接看到并操作浏览器。遇到需要人工介入的认证场景时，**不要用 `fill`/`click` 盲目尝试**，应按以下流程交接：\n\n#### 触发条件\n\n从 `snap` 中发现以下任一信号时，启动人工协作流程：\n\n- 页面 title 为「登录」/「Login」/「Sign In」\n- 页面 `url` 包含 `/login`、`/auth`、`/signin`\n- 快照中出现「登录」按钮 + 用户名/密码输入框\n- 快照中出现「验证码」「短信验证」「扫码登录」「滑块验证」等关键字\n- `open` 后自动跳转到登录页（URL 变化）\n\n#### 协作流程\n\n```\n第1步：通报用户\n  告知当前页面需要登录/验证，简明描述页面内容（输入框、验证码类型等）\n\n第2步：询问凭据（可选）\n  如果用户无凭据 → 跳过，直接等用户操作\n  如果用户提供凭据 → 用 fill/click 填入账号密码，点击登录按钮\n\n第3步：等待用户完成验证\n  明确告诉用户\"请在浏览器中完成验证码/二次验证\"\n  用户说\"好了\"\"完成了\"\"继续\"之后才继续\n\n第4步：验证登录状态\n  执行 pw-browser snap\n  检查是否进入目标页面 → 如果还是登录页，询问用户是否还需要操作\n  如果已进入 → 继续自动化流程\n```\n\n#### 示例对话\n\n```\nAgent: 页面跳转到了登录页 (https://xxx.com/login)，页面上有：\n       用户名输入框、密码输入框、登录按钮、滑块验证码。\n       需要我帮你填入账号密码吗？还是你在浏览器里自己操作？\n\nUser:  我来操作\n\nAgent: 好的，Chrome 窗口已打开 — 请完成登录后告诉我。\n\nUser:  好了\n\nAgent: [执行 snap]\n       登录成功！当前是「工作台」页面，左侧菜单有...\n```\n\n#### 重要约束\n\n- **不猜测凭据**：永远不要尝试默认密码或遍历登录\n- **不绕过验证码**：遇到验证码/滑块/短信验证时，立即交给用户\n- **不过度等待**：用户说继续后立即 snap，不额外 sleep\n- **登录失败回环**：snap 后发现仍在登录页 → 告知用户\"看起来还没登录成功，密码错误或验证未通过，请再试试\"\n\n### 规则 5：不要主动新建 tab\n\n点击导致新 tab 时用 `tab list/select/close` 处理。没有 `tab new` 命令。\n\n### 规则 6：翻页前读取策略\n\n涉及翻页、统计、收集、遍历时，参考下面的\"分页策略\"章节。\n\n### 规则 7：不手动读取快照文件\n\n快照通过 `pw-browser snap` 命令获取，不要直接读 `~/.pw-browser/snap.yml`。\n\n### 规则 8：SPA / 富文本编辑器\n\n遇到知识库、文档系统、CMS 等 SPA 页面，参考下面的\"SPA 与富文本编辑器\"章节。\n\n### 规则 9：daemon 故障恢复\n\n如果 CLI 返回连接错误：\n\n```bash\n# 删除旧的 daemon 状态文件\nrm -rf ~/.pw-browser/daemon.json\n```\n\n**杀掉占用端口 19223 的旧进程：**\n\n```bash\n# Windows (PowerShell)\npowershell -Command \"Get-NetTCPConnection -LocalPort 19223 -ErrorAction SilentlyContinue | ForEach-Object { Stop-Process -Id \\$_.OwningProcess -Force }\"\n\n# macOS / Linux\nlsof -ti:19223 | xargs kill -9 2>/dev/null\n# 或: fuser -k 19223/tcp 2>/dev/null\n```\n\n**重新启动：**\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" daemon &\nsleep 4\n```\n\n---\n\n## 命令速查\n\n### 生命周期\n| 命令 | 说明 |\n|------|------|\n| `pw-browser init` | 连接 daemon，确认浏览器可用 |\n| `pw-browser open <url>` | 导航到 URL |\n| `pw-browser close` | 关闭当前页面 |\n| `pw-browser close --all` | 关闭浏览器 + daemon |\n| `pw-browser recover` | 重启浏览器连接 |\n\n### 状态感知\n| 命令 | 说明 |\n|------|------|\n| `pw-browser snap` | 获取页面快照（含 ref 引用）。**ref 在同一页面会话内保持稳定**（同一逻辑元素同 ref）；**每次导航后（新 URL / 新页面）必须重新 snap**，不可复用旧 ref |\n| `pw-browser wait-for <target> [--timeout ms]` | 等待条件满足 |\n\n> **Shadow DOM / iframe 支持**：快照会递归进入 **open shadow root** 与 **同源 iframe**，这些元素同样出现在 ref 表中并可直接 `click`/`fill`/`upload`/`drag`。`snap` 输出的 ref 信息带 `inShadow: true`（shadow 内）或 `frameChain`（iframe 链）标记；定位由 `css >>>` 穿透 + `frameLocator` 自动完成，外部 agent 无需关心。`--annotate` 截图仅标注主文档元素（shadow/iframe 元素无法用 xpath 定位标注，但文本快照中仍可点）。跨域 iframe 不可访问，自动跳过。\n\n> **超大页面快照保护**：对元素极多的页面（数万节点），`snap` 默认最多收集 **3000** 个可交互元素，超出即停止收集并在输出标记 `⚠ snapshot truncated`，JSON 返回 `truncated: true`。这是为防止超大型 DOM 拖慢/撑爆快照的兜底；可用环境变量 `PW_BROWSER_SNAP_LIMIT=<N>` 调大上限（设为 `0` 关闭上限，但仍会遍历整棵树），或先交互缩小页面范围再 `snap`。\n\n### 交互\n| 命令 | 说明 |\n|------|------|\n| `pw-browser click <ref>` | 点击元素 |\n| `pw-browser fill <ref> \"text\"` | 填入文本 |\n| `pw-browser type \"text\"` | 键盘输入 |\n| `pw-browser press <key>` | 按下按键（Enter, Escape, Tab 等） |\n| `pw-browser hover <ref>` | 悬停 |\n| `pw-browser select <ref> <option>` | 选择下拉选项 |\n| `pw-browser check <ref>` | 勾选复选框 |\n| `pw-browser uncheck <ref>` | 取消勾选 |\n| `pw-browser upload <ref> <file1> [file2 ...]` | 文件上传（`<input type=\"file\">`，支持多文件，逗号或空格分隔） |\n| `pw-browser drag <ref源> <ref目标>` | 拖拽（把源元素拖到目标元素，基于 Playwright `dragTo`） |\n| `pw-browser download <ref> [--path dir] [--timeout ms]` | 文件下载（对称于 `upload`）：可选点击 `<ref>` 触发下载，保存到 `--path`（默认当前目录）；也可作为 `act` 动作 `{\"action\":\"download\",\"ref\":\"eN\",\"path\":\"/tmp/x.csv\"}` |\n\n### 页面导航\n| 命令 | 说明 |\n|------|------|\n| `pw-browser goto <url>` | 同 open |\n| `pw-browser go-back` | 后退 |\n| `pw-browser go-forward` | 前进 |\n| `pw-browser reload` | 刷新 |\n\n### 高级\n| 命令 | 说明 |\n|------|------|\n| `pw-browser screenshot [ref] [--path file] [--annotate]` | 截图；`--annotate` 在可交互元素上叠加与 snap ref 对应的编号框，供多模态 agent 直接读编号定位 |\n| `pw-browser mousewheel <dx> <dy>` | 滚动 |\n| `pw-browser eval \"<expr>\" [ref]` | ⚠️ 执行**任意 JavaScript**（页面上下文，完整页面级代码执行：可读 cookie/存储、发起带凭证请求；受 token 认证保护，safe-mode 下禁用）<br>**作用域语义**：不带 `ref` → 表达式在页面全局求值（返回 `scope: \"page\"`）；带 `ref` → 表达式在该元素上求值，标识符 **`el`** 绑定到对应 DOM 节点（返回 `scope: \"element\"`），例：`eval \"el.textContent\" e3`。<br>返回值必须可 JSON 序列化，**不能返回 DOM 节点或循环结构**。 |\n| `pw-browser run-code \"<code>\"` | ⚠️ 执行 Playwright 代码（受限沙箱：无直接 Node fs/process 权限，但可经浏览器下载/上传读写本地文件） |\n| `pw-browser dialog-accept [text]` | 确认对话框 |\n| `pw-browser dialog-dismiss` | 取消对话框 |\n\n### Tab\n| 命令 | 说明 |\n|------|------|\n| `pw-browser tab list` | 列出所有 tab |\n| `pw-browser tab select <idx>` | 切换到指定 tab (0-based) |\n| `pw-browser tab close <idx>` | 关闭指定 tab |\n\n### 延时\n| 命令 | 说明 |\n|------|------|\n| `pw-browser sleep <seconds>` | 等待 N 秒 |\n\n### 批量动作与历史（借鉴 browser-use 的 multi-act / 自纠错）\n| 命令 | 说明 |\n|------|------|\n| `pw-browser act '<json>'` | 批量执行动作序列（JSON 数组），如 `[{\"action\":\"fill\",\"ref\":\"e3\",\"text\":\"hello\"},{\"action\":\"click\",\"ref\":\"e5\"}]`。每步后自动检测 DOM 变化，若页面出现新元素则**中断序列并自动 re-snap** 返回最新快照；失败步附带诊断（元素是否仍存在/相似 ref 建议） |\n| `pw-browser history [--limit N] [--clear]` | 查询 daemon 记录的操作历史（每条命令、参数、耗时、结果），`--clear` 清空 |\n\n`act` 支持的动作：`click` / `fill` / `type` / `press` / `hover` / `select` / `check` / `uncheck` / `upload`（对象带 `files: [\"/path\"]`）/ `drag`（对象带 `target: \"eN\"`）/ `goto` / `screenshot`，动作对象形如 `{\"action\":\"...\",\"ref\":\"eN\",\"text\":\"...\",\"key\":\"...\",\"option\":\"...\",\"url\":\"...\",\"files\":[\"...\"],\"target\":\"eN\"}`。\n\n---\n\n## 省 Token 用法（默认即高效，别退回 browser-use 的反模式）\n\n本 skill 是**确定性执行器 + 持久 daemon**，大模型（外部 AI）只负责规划、不内嵌在 skill 里。因此**没有「每步都调 LLM」的 token 黑洞**——token 只在外部 AI 主动调用时产生，且完全可控。请保持以下用法以持续省 token：\n\n- **用文本 `snap` 规划，而不是每步截图喂视觉模型**。`snap` 返回的是紧凑的 ref 文本表（`e1 button 提交`），几十 token；`screenshot` 一张图是数百 KB 的 base64，贵 1~2 个数量级。\n- **同一页面内复用稳定 ref，不必每步重新 `snap`**（无导航变化时）。同一逻辑元素的 ref 跨多次 snap 保持不变（见上「状态感知」表），外部 AI 可直接拿上一步的 ref 点 `click e1` / `fill e2`。**若页面发生导航（URL 变化、新页面打开），必须重新 `snap`**，不可复用旧 ref。\n- **把动作攒成 `act` 一次性发**。登录等一连串操作写成 `[{...},{...}]` 一次调用，daemon 内部自纠错，中间不回模型。理想形态：`1 次 snap → 1 次 act → 完事`。\n- **`screenshot --annotate` 是 opt-in**，仅在真有视觉歧义、需要多模态定位时才用；不要默认每步截图。\n\n> ⚠️ 若外部 AI 被 prompt 成「每步都 `screenshot --annotate` 丢给视觉模型」，就会复刻 browser-use 的烧钱循环。token 成本的责任在编排层，不在 skill。\n\n## Cookie 与本地存储（一等命令）\n\n不再需要靠 `eval` 曲线救国，直接用以下命令读写 cookie 与 `localStorage`：\n\n### `cookies`\n| 子命令 | 说明 |\n|--------|------|\n| `pw-browser cookies list` | 列出当前上下文全部 cookie |\n| `pw-browser cookies export [--path file]` | 导出 cookie 到 JSON 文件（默认 `~/.pw-browser/cookies.json`） |\n| `pw-browser cookies import <file>` | 从 JSON 文件导入 cookie |\n| `pw-browser cookies clear` | 清空全部 cookie |\n| `pw-browser cookies set <name> <value> [--domain d] [--path p]` | 设置一个 cookie；`--domain` 省略时取当前页面域名 |\n\n### `storage`（localStorage）\n| 子命令 | 说明 |\n|--------|------|\n| `pw-browser storage get [key]` | 读取某个 key（省略 key 则返回全部，以对象形式返回） |\n| `pw-browser storage set <key> <value>` | 写入 key/value |\n| `pw-browser storage clear` | 清空 localStorage |\n| `pw-browser storage export [--path file]` | 导出 localStorage 到 JSON 文件 |\n| `pw-browser storage import <file>` | 从 JSON 文件导入（逐 key 写入） |\n\n> ⚠️ `cookies` / `storage` 依赖真实页面源（http/https）。`file://` 与 `data:` 页面不支持 cookie，`localStorage` 行为也不可靠——请先 `open` 一个真实 URL 再操作。\n\n> 🔒 **会话持久化风险（Rogue Agent / 中危）**：`cookies` / `storage` 的 `export` / `import` 让登录态可落盘备份、跨运行恢复——对正常用户是免登录便利，但**在失控或恶意 Agent 场景下，这正是\"特权访问持久化\"的典型手段**：导出的会话文件等同一份可复用的身份凭证，可被用来跳过认证、长期驻留。缓解：① 导出的会话文件**等同密钥**，用完即 `rm`；暂存请留在 `~/.pw-browser/` 内——自 v1.3.9 起该目录以 `0700` 创建、凭证文件与 daemon token 均以 `0600` 写入（**无需再手工 `chmod`**，旧文件也会被自动收紧；Windows 上 POSIX 权限位不由系统强制，依赖用户目录 ACL）；② **不要**在自动化流程里默认把凭证持久化到磁盘——自 v1.3.9 起这不再只是建议：操作者可用 `PW_BROWSER_CRED_PERSIST=off` 启动 daemon，**强制禁用** `export`/`import`（返回 `CredentialPersistenceDisabled`）而保留 `list`/`get`/`set`/`clear` 内存态操作，在不牺牲自动化能力的前提下切断特权持久化路径；③ 不需要时尽快 `cookies clear` / `storage clear` 并 `pw-browser shutdown` 关闭 daemon，利用空闲自动退出（`PW_BROWSER_IDLE_MS`，默认 15min）缩短凭证在内存中的驻留窗口；④ 对来源不可信的调用方，用 `PW_BROWSER_SAFE_MODE=1` 启动 daemon——自 v1.3.2 起它会**整体禁用**全部 `cookies` / `storage` 子命令（凭证原语与代码执行同级拦截），必要时再配合沙箱隔离或限定操作范围；⑤ **代码层兜底（v1.3.1 / v1.3.8）**：`export`/`import` 默认被限制在 `~/.pw-browser/` 目录内（realpath 解析，软链无法逃逸），从路径层面降低凭证被散落或加载外部攻击者构造文件的可能；且**该限制不可由被约束方自行解除**——越界须由操作者以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon 后、调用方再显式 `--unsafe`，缺一不可，全部访问均写 stderr 审计。详见 QUICKSTART 示例 4 安全提醒。\n\n## Daemon 生命周期（为什么任务结束后浏览器还在）\n\ndaemon 是**故意持久化**的：它跨命令持有同一个浏览器实例，避免每次交互都重开 Chrome。因此：\n\n- 你看到「任务结束 Chrome 还在」是正常的——daemon 进程还活着、抱着浏览器。\n- **显式停止**：`pw-browser shutdown` 会关掉浏览器并退出 daemon（已加固：即使 `browser.close()` 卡住也会超时兜底退出，不会再退不出/卡客户端）。\n- **空闲自动退出**：daemon 默认 **15 分钟无命令** 就自动关浏览器并退出（环境变量 `PW_BROWSER_IDLE_MS` 可改，设为 `0` 关闭该特性）。所以走开后不用手动 `shutdown`，它自己会清理，Chrome 不会一直挂着。\n- **监听端口可配 + 冲突避让**：默认 `127.0.0.1:19223`，可用环境变量 `PW_BROWSER_PORT` 覆盖。启动时若目标端口已有 daemon 存活（`/health` 返回 200），当前进程会直接退出（避免双开）；若被其它进程占用（`EADDRINUSE`），则自动递增端口直到可用，并把实际端口写回 `~/.pw-browser/daemon.json`。客户端读取该文件里的端口，无需手动指定。\n- **别直接杀进程**：用任务管理器 / `Stop-Process` 强杀 daemon 可能导致浏览器子进程残留（Windows 上 Playwright 的 job object 通常会回收，但不保证）。优先用 `shutdown` 或等空闲自动退出。\n\n---\n\n## 等待策略\n\n`wait-for` 支持多种目标格式：\n\n```bash\n# 等待 URL 匹配\npw-browser wait-for \"url:**/dashboard\"\n\n# 等待文本出现\npw-browser wait-for \"text=加载完成\"\n\n# 等待页面加载状态（load / domcontentloaded / networkidle）\npw-browser wait-for \"state:networkidle\"\n\n# 等待 CSS 选择器\npw-browser wait-for \".result-list\" --timeout 15000\n```\n\n---\n\n## 运行自定义代码（`run-code`）⚠️ 高级功能\n\n> ⚠️ **安全警告：** `run-code` 在 daemon 进程的 **受限沙箱（`vm` 模块）** 中执行 Playwright 代码。它**无法直接调用** Node.js 系统 API（`fs` / `child_process` / `process` / `require`）；但浏览器上下文本身可经下载（`download.saveAs`）或文件上传（`setInputFiles`）在本地磁盘读写文件、并能发起任意网络请求（沙箱不阻止），因此**仍能持久化数据到本地磁盘**。仅在用户明确指定的任务中使用，不要执行来源不明的代码片段。\n\n当内置命令不够用时，用 `run-code` 执行自定义 Playwright 代码：\n\n```bash\n# 获取页面标题\npw-browser run-code \"return await page.title();\"\n\n# 获取页面 HTML\npw-browser run-code \"return await page.content();\"\n\n# 在页面中执行 JS\npw-browser run-code \"return await page.evaluate(() => document.title);\"\n\n# 等待网络空闲\npw-browser run-code \"await page.waitForLoadState('networkidle');\"\n\n# 复杂场景：提取列表数据\npw-browser run-code \"\n  const items = await page.locator('.product-item').all();\n  const results = [];\n  for (const item of items) {\n    results.push({\n      title: await item.locator('.title').textContent(),\n      price: await item.locator('.price').textContent()\n    });\n  }\n  return JSON.stringify(results);\n\"\n```\n\n**注意：**\n- `run-code` 中直接使用 Playwright Page API\n- 代码在 async 函数中执行，`page` 对象已注入\n- 返回值自动序列化为字符串\n- ⚠️ 此命令可以触发实际的业务操作（提交订单、发送消息、删除数据等），执行前确认用户意图\n\n---\n\n## 分页策略\n\n### 步骤 1：识别分页类型\n\n> ⚠️ 下表关键词为**识别启发式**：中文 / 英文示例（\"下一页\"/\"Next\"、\"加载更多\"/\"Load more\"）并不穷举，非中文页面的文案会不同。实际定位请优先用 `snap` 的 `ref` 或 CSS 选择器，勿仅依赖可见文本。\n\n从 snap 判断：\n\n| 类型 | 识别信号 | 翻页方式 |\n|------|---------|---------|\n| **页码分页** | 底部有 1/2/3...页码、\"下一页\"/\"Next\"/\">\" | 点击页码或\"下一页\" |\n| **无限滚动** | 底部无分页控件，内容随滚动增加 | `mousewheel` 滚动 |\n| **加载更多** | 底部有\"加载更多\"/\"Load more\"/\"查看更多\" | 点击该按钮 |\n\n### 步骤 2：执行翻页\n\n**页码分页：**\n```bash\npw-browser snap                          # 找到\"下一页\"按钮的 ref\npw-browser click e42                     # 点击\npw-browser sleep 2 && pw-browser snap    # 验证\n```\n\n**无限滚动：**\n```bash\npw-browser mousewheel 0 800\npw-browser sleep 2 && pw-browser snap\n```\n\n**加载更多按钮：**\n```bash\npw-browser click <ref>\npw-browser sleep 2 && pw-browser snap\n```\n\n### 步骤 3：判断翻页成功\n\n| 方式 | 成功信号 | 失败/结束信号 |\n|------|---------|-------------|\n| 页码 | 内容更新，URL 变化 | \"下一页\"按钮 disabled 或消失 |\n| 滚动 | 内容增加，新元素出现 | 内容不变，\"没有更多了\" |\n| 按钮 | 新内容加载，按钮仍可点击 | \"已加载全部\"，按钮消失 |\n\n---\n\n## SPA 与富文本编辑器 ⚠️ 破坏性操作\n\n> ⚠️ **警告：** 以下操作会**真实修改**网页内容（知识库文档、CMS 页面等）。执行前确认当前处于编辑/草稿状态、修改内容已经用户确认。保存/发布操作不可逆。\n\n处理知识库、文档系统、CMS 等 SPA 页面的编辑操作：\n\n### 识别信号\n\n- 点击\"编辑\"后 URL 不变但按钮变化\n- snap 中出现 `contenteditable`、编辑器 toolbar\n- 不是普通 `input/textarea`，而是复杂编辑器\n\n### 编辑流程\n\n1. **进入编辑态：** `pw-browser click <编辑按钮的ref>`\n2. **验证进入：** `pw-browser snap` — 检查是否出现\"更新\"/\"保存\"按钮\n3. **写入内容（RTE）：**\n```bash\npw-browser run-code \"\n  const editor = page.locator('[contenteditable=\\\"true\\\"]').first();\n  await editor.click();\n  await page.keyboard.press('Control+A');\n  await page.keyboard.type('要写入的内容');\n  await page.waitForTimeout(500);\n\"\n```\n4. **保存：**\n```bash\npw-browser run-code \"\n  await page.evaluate(() => {\n    const btn = Array.from(document.querySelectorAll('button'))\n      .find(b => ['更新','保存','发布'].includes(b.textContent.trim()));\n    btn?.click();\n  });\n  await page.waitForTimeout(3000);\n\"\n```\n5. **验证：** `pw-browser snap` — 确认保存成功、内容正确\n\n> **不要**直接用 `innerText`/`textContent` 修改 RTE 内容。Playwright 的 `keyboard.type` 和 `fill` 是正确方式。\n\n---\n\n## 结构化输出\n\n所有命令在 daemon 端返回 JSON：\n\n```json\n{\"ok\": true, \"data\": {...}, \"elapsedMs\": 123}\n{\"ok\": false, \"error\": {\"kind\": \"ElementNotFound\", \"message\": \"...\"}, \"elapsedMs\": 50}\n```\n\nCLI 客户端默认以人类可读格式输出；加 `--json` 标志输出原始 JSON。\n\n## 错误处理\n\n| 错误类型 | 原因 | 处理 |\n|---------|------|------|\n| `ElementNotFound` | snap 后 ref 已失效 | 重新 snap 获取新 ref |\n| `NavigationTimeout` | 页面加载超时 | 先 snap 检查实际状态 |\n| 连接拒绝 | daemon 未运行 | 重新启动 daemon |\n| 空快照 | 页面未加载完成 | wait-for state:load 后重新 snap |\n\n---\n\n## 完整示例：百度搜索\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nPW=\"NODE_PATH=${SKILL_DIR}/node_modules node ${SKILL_DIR}/pw-browser.js\"\n\n# 启动 daemon（首次）\n$PW daemon &\nsleep 4\n\n# 打开百度\n$PW open https://www.baidu.com\n\n# 快照 → 找到搜索框和按钮的 ref\n$PW snap\n# 例如：e12 = textarea（搜索框），e13 = button（百度一下）\n\n# 填搜索关键词\n$PW fill e12 \"天气预报\"\n\n# 点击搜索\n$PW click e13\nsleep 2\n\n# 检查搜索结果\n$PW snap | head -30\n\n# 清理\n$PW close --all\n```\n\n---\n\n## 深入参考\n\n| 场景 | 文件 |\n|------|------|\n| 端到端示例（表单/上传下载/shadow-iframe/cookie/act） | `QUICKSTART.md`（中文）/ `QUICKSTART.en.md`（英文） |\n| 翻页策略详解 | `references/pagination.md` |\n| 富文本编辑器策略 | `references/rich-text-editor.md` |\n| 运行自定义代码 | `references/running-code.md` |\n\nFile v1.0.16:README.md\n\n# pw-browser\n\n> 基于 Playwright 的浏览器自动化 CLI — daemon + client 架构，直接使用系统 Chrome/Edge，无需下载额外浏览器。\n\n> 📘 简体中文文档（本文件）。English documentation: [`README.en.md`](./README.en.md)。端到端示例见 [`QUICKSTART.md`](./QUICKSTART.md)（含 [`QUICKSTART.en.md`](./QUICKSTART.en.md)）。\n\n> ⚠️ **完整能力（含代码执行与凭证操控）**：本工具**不只是\"点开网页\"**——它内置 `eval`（页面上下文**任意 JavaScript** 执行）与 `run-code`（守护进程上下文执行 Playwright/Node 代码）两项**代码执行**能力，以及 `cookies` / `storage` **凭证读写原语**（可**无需代码执行**即提取/注入登录态与会话令牌）。全部经持久化本地守护进程（127.0.0.1:19223）控制，受 daemon token 认证保护；`PW_BROWSER_SAFE_MODE=1` 可**彻底禁用**代码执行与凭证读写（v1.3.2+）。请先阅读下方「🔐 权限模型与最低特权」与「⚠️ 安全边界与使用场景」了解完整攻击面与适用边界，再决定是否授权。\n\n## 特性\n\n- **系统浏览器复用**：通过 Playwright `channel: 'chrome'` 连接系统已安装的 Chrome 或 Edge，不下载 Chromium\n- **跨平台**：支持 Windows、macOS、Linux\n- **持久化会话**：daemon 保持浏览器状态跨命令存活，无需每次重启\n- **可访问性快照**：`snap` 命令生成页面元素树，带 ref 引用，无需写 CSS 选择器\n- **非 headless**：浏览器窗口始终可见，用户可实时监控所有操作\n- **人机协作**：遇到登录/验证码时自动交接给用户操作\n\n## 安装\n\n```bash\ngit clone <repo-url> pw-browser\ncd pw-browser\nnpm install\n```\n\n> 前提：系统已安装 [Node.js](https://nodejs.org/) 18+ 和 Chrome 或 Edge 浏览器。\n\n## 快速开始\n\n```bash\n# 1. 启动 daemon（后台运行）\nnode pw-browser.js daemon &\n\n# 2. 等待 daemon 就绪\nsleep 4\n\n# 3. 打开页面\nnode pw-browser.js open https://www.baidu.com\n\n# 4. 获取页面快照（必须！每次交互前都要 snap）\nnode pw-browser.js snap\n\n# 5. 交互 — 基于快照中的 e0, e1, e2... ref\nnode pw-browser.js fill e12 \"天气预报\"\nnode pw-browser.js click e13\n\n# 6. 查看结果\nsleep 2\nnode pw-browser.js snap\n\n# 7. 关闭\nnode pw-browser.js close --all\n```\n\n## 命令一览\n\n| 类别 | 命令 | 说明 |\n|------|------|------|\n| 生命周期 | `init` / `open <url>` / `close` / `close --all` / `recover` | 启动、导航、关闭、恢复 |\n| 状态感知 | `snap` / `wait-for <target>` | 快照（**ref 跨多次 snap 保持稳定**，同元素同 ref） |\n| 交互 | `click <ref>` / `fill <ref> \"text\"` / `type \"text\"` / `press <key>` / `hover <ref>` / `select <ref> <opt>` / `check <ref>` / `uncheck <ref>` / `upload <ref> <file...>` / `drag <ref源> <ref目标>` / `download <ref> [--path dir]` | 点击、填写、按键、上传、拖拽、下载等 |\n| Cookie/存储 | `cookies list` / `export [--path f]` / `import <file>` / `clear` / `set <name> <value> [--domain d]` · `storage get [key]` / `set <k> <v>` / `clear` / `export [--path f]` / `import <file>` | 读写 cookie 与 localStorage（需先 `open` 真实 http/https 页面） |\n| 导航 | `goto <url>` / `go-back` / `go-forward` / `reload` | 页面导航控制 |\n| Tab | `tab list` / `tab select <idx>` / `tab close <idx>` | 多标签页管理 |\n| 批量 | `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` | 一次性执行动作序列，中途检测 DOM 变化自动中断重规划 |\n| 历史 | `history [--limit N] [--clear]` | 查看/清空操作历史 |\n| 高级 | `screenshot [--annotate]` / `mousewheel <dx> <dy>` / `eval \"<expr>\" [ref]` ⚠️ / `run-code \"<code>\"` ⚠️ | 截图（`--annotate` 叠加与 ref 对应编号框）、滚动、代码执行。`eval` 不带 `ref` 时在页面全局求值（`scope: \"page\"`）；带 `ref` 时标识符 **`el`** 绑定到该 DOM 节点（`scope: \"element\"`），如 `eval \"el.textContent\" e3`。返回值须可 JSON 序列化 |\n| 延时 | `sleep <seconds>` | 等待 |\n| 对话框 | `dialog-accept [text]` / `dialog-dismiss` | 处理原生 alert/confirm/prompt 对话框 |\n| 守护进程 | `shutdown` | 关闭持久化 daemon，释放浏览器进程 |\n\n> ⚠️ `eval` 在**浏览器上下文**执行 JS（无 Node 权限）；`run-code` 在 daemon 的**受限沙箱（`vm`）**中执行 Playwright 代码（无直接 Node `fs`/`process`/`child_process` 权限，但可经浏览器下载/上传读写本地文件、发起任意网络请求）。二者均拥有完整浏览器控制权。仅在用户明确指定的任务中使用。\n\n## 架构\n\n```\n┌──────────────┐     HTTP (localhost:19223)     ┌──────────────┐\n│  pw-browser  │ ──────────────────────────────→│   Daemon     │\n│  (CLI 客户端) │                                │  (浏览器进程)  │\n└──────────────┘                                └──────┬───────┘\n                                                       │\n                                                       ├─ Playwright\n                                                       ├─ Chrome / Edge\n                                                       └─ 页面状态持久化\n```\n\ndaemon 启动后持续运行，浏览器和页面状态跨命令保持。CLI 每次通过 HTTP 调用 daemon。\n\n## 工作流核心规则\n\n1. **先观察再操作**：每次交互前必须 `snap`，基于快照 ref 操作\n2. **登录交接**：遇到登录/验证码页面时，告知用户手动操作，用户确认后继续\n3. **翻页前先识别类型**：页码分页 / 无限滚动 / 加载更多，对应不同翻页方式\n4. **daemon 故障恢复**：连接错误时杀端口 19223 进程 → 删 daemon.json → 重启\n\n详细规则和示例见 `SKILL.md`。\n\n## 增强能力（借鉴 browser-use）\n\n本工具的设计定位是「**安全可控的浏览器 CLI 基座 + 持久 daemon，规划交给外部 AI**」，而非 browser-use 那种「大脑+手一体」的 Agent 框架。我们从 browser-use 借鉴了四项对 CLI 基座同样有价值的能力：\n\n### A. 稳定的元素引用（stable ref）\n- `snap` 不再每次重排 ref，而是按元素的**语义身份**（有文本/placeholder/aria-label 用 `tag|text|placeholder|aria`；否则回退到 DOM 分支路径哈希）建立 `stableKey → ref` 持久映射。\n- **同一逻辑元素跨多次 snap 保持同一 ref**，外部 AI 不必每步重新 `snap`，可直接复用之前的 ref 继续交互（典型场景：填完表单再点提交，e13 还是那个提交按钮）。\n- `findElement` 新增 **Strategy 0：xpath 精确定位优先**——用快照里记录的精确 xpath 钉住元素，彻底解决纯语义查找「同名元素误点」的歧义。\n\n### B. 视觉辅助截图（`screenshot --annotate`）\n- `screenshot --annotate` 会在页面上注入覆盖层，为**每个与 snap ref 对应的元素**叠加红色边框 + 编号 label（用 `document.evaluate(xpath)` 精确定位）。\n- 截图后自动移除覆盖层。多模态 AI 可「看图读编号」直接定位 `e13` 在页面哪个位置，弥补纯文本快照缺少空间信息的短板。\n\n### C. 动作序列 + 自纠错（`act`）\n- `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` 一次性下发动作序列，复用 `executeSingle` 逐个执行。\n- 每个非首动作执行前重新 `snap` 并比对 `branchPathHash` 集合：若检测到页面出现**新元素**（DOM 变化），立即 `interrupted=true` 并返回最新快照，交由外部 AI 重新规划——这对应 browser-use 的 `multi_act` 中途中断重规划。\n- 失败的动作会附带 `diagnosis`（ref 是否仍在、相似 ref 建议），便于自愈。\n\n### D. 结构化输入 + 操作历史（`history`）\n- `act` 接受结构化 JSON 动作数组（而非零散子命令），降低外部 AI 拼 CLI 参数的出错率。\n- `history` 记录每次操作的 `{ ts, cmd, params（已过滤 token）, ok, elapsedMs }`，上限 500 条，可 `--clear`。外部 AI 可回溯「刚才点了什么、哪步失败」，实现上下文压缩与复盘——对应 browser-use 用 Mem0 压缩历史上下文的思路（此处用轻量本地历史替代向量库）。\n\n### E. Shadow DOM / iframe 穿透\n- 快照递归进入 **open shadow root** 与 **同源 iframe**，这些元素同样出现在 ref 表中并可直接 `click`/`fill`/`upload`/`drag`。\n- ref 信息带 `inShadow: true`（shadow 内）或 `frameChain`（iframe 链）标记；定位由 `css >>>` 穿透 + `frameLocator` 自动完成，外部 AI 无感知。\n- 跨域 iframe 不可访问，自动跳过；`--annotate` 截图仅标注主文档元素（shadow/iframe 元素无法用 xpath 定位标注，但文本快照中仍可点）。\n\n### F. 文件上传与拖拽\n- `upload <ref> <file1> [file2 ...]`：对 `<input type=\"file\">` 设置文件（支持多文件）。\n- `drag <ref源> <ref目标>`：基于 Playwright `dragTo` 实现拖拽。\n- `download <ref> [--path dir]`：对称于 `upload`——可选点击 `<ref>` 触发下载，保存到 `--path`（默认当前目录）；亦可作为 `act` 动作。\n  > 🔒 v1.3.10+：网页给出的建议文件名会被强制取 `basename` 并限制在 `--path` 目录内（越界返回 `PathTraversal`），恶意页面无法借 `../../` 之类的文件名写到目录之外。`--path` 若指向具体文件则按原样使用。\n\n### G. Cookie 与本地存储（一等命令）\n- `cookies list|export|import|clear|set`：读写当前上下文 cookie，不再需要靠 `eval` 曲线救国。\n- `storage get|set|clear|export|import`：基于 `localStorage` 的读写（依赖真实 http/https 页面源）。\n- 典型用法：登录后 `cookies export` 备份会话，下次 `cookies import` 直接恢复，免去重复登录。\n- 📁 **路径限制（操作者门禁）**：`export`/`import` 默认限制在 `~/.pw-browser/` 目录内（防凭证散落 `/tmp` 或加载外部攻击者构造文件；路径经 realpath 解析，软链无法逃逸）。自 v1.3.8 起**该限制不可由调用方自行解除**——越界需**同时**满足：① 操作者以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon；② 调用方显式传 `--unsafe`。仅传 `--unsafe` 会返回 `UnsafeOverrideNotPermitted`。命令返回带 `confined` 与 `warning` 字段，daemon stderr 记录每次凭证路径访问。\n- 🔒 **会话持久化风险（Rogue Agent / 中危）**：`export`/`import` 让登录态可落盘、跨运行恢复——便利，但失控/恶意 Agent 可用它做\"特权访问持久化\"。导出的会话文件等同密钥：用完即删；暂存留在 `~/.pw-browser/` 内即可——自 v1.3.9 起该目录 `0700`、凭证文件与 daemon token 均 `0600` 写入，无需手工 `chmod`（Windows 上依赖用户目录 ACL）。不想让凭证落盘的场景，用 `PW_BROWSER_CRED_PERSIST=off` 启动 daemon：**强制禁用** `export`/`import`（`CredentialPersistenceDisabled`），保留 `list`/`get`/`set`/`clear` 内存态操作——比安全模式粒度更细，适合需要完整自动化但禁止会话跨运行留存的长驻 agent。不用时 `cookies clear`/`storage clear` 并 `shutdown` 关 daemon（空闲默认 15min 自动退出）。对不完全可信的调用方，用 `PW_BROWSER_SAFE_MODE=1` 启动 daemon——自 v1.3.2 起它会**整体禁用**全部 cookies/storage 子命令（凭证原语与代码执行同级拦截），必要时再配合沙箱隔离。\n\n## 省 Token 与 Daemon 生命周期\n\n**没有「每步调 LLM」的 token 黑洞**：本 skill 只是确定性执行器 + 持久 daemon，大模型只在外部编排层、且仅在主动调用时产生 token。保持以下用法即持续省 token：\n\n- 用文本 `snap`（ref 表，几十 token）规划，而非每步 `screenshot`（数百 KB base64）喂视觉模型；\n- 复用跨 snap 稳定的 ref，不必每步重新 `snap`；\n- 把一连串操作攒成一次 `act '<json>'`，daemon 内部自纠错，中间不回模型。\n\n**daemon 是故意持久化的**（跨命令复用同一浏览器，避免每次重开 Chrome）。因此：\n\n- 任务结束后浏览器窗口还在是正常的——daemon 进程仍持有它；\n- 显式停止：`pw-browser shutdown`（已加固，`browser.close()` 卡住也会超时兜底退出，不会退不出/卡客户端）；\n- **空闲自动退出**：默认 **15 分钟无命令** 即自动关浏览器并退出，环境变量 `PW_BROWSER_IDLE_MS` 可调整（设为 `0` 关闭）。走开后无需手动清理；\n- **监听端口可配 + 冲突避让**：默认 `127.0.0.1:19223`，可用 `PW_BROWSER_PORT` 覆盖。启动时若目标端口已有 daemon 存活则直接退出（避免双开）；若被其它进程占用则自动递增端口，并把实际端口写回 `~/.pw-browser/daemon.json`，客户端自动读取。\n- 优先用 `shutdown` 或等空闲退出，避免直接强杀进程导致浏览器子进程残留。\n\n## 安全说明\n\n- daemon 监听 `127.0.0.1:19223`，仅本机可访问\n- **命令认证**：除 `/health` 存活探针外，所有 daemon 命令都要求随机 `token`。token 在 daemon 启动时生成，写入 `~/.pw-browser/daemon.json`（默认仅当前用户可读），CLI 自动携带。未读取该文件的本地进程无法调用——这关闭了\"无认证 HTTP 端点 = RCE 界面\"的缺口\n- **安全模式**：`PW_BROWSER_SAFE_MODE=1 node pw-browser.js daemon` 可彻底禁用 `run-code` / `eval` **以及全部 `cookies` / `storage` 凭证原语**（v1.3.2+），仅保留 snap/click/fill 等白名单命令——适合接入来源不完全可信的 agent\n- `eval` 运行在浏览器上下文；`run-code` 运行在 `vm` 受限沙箱（无直接 Node 系统 API，但可经浏览器下载/上传读写本地文件），仅限本地信任环境使用\n- 浏览器以非 headless 模式运行，用户可实时监控\n- 不推荐作为公开 API 服务暴露，如需请加认证和操作白名单\n- **权限模型说明**：frontmatter 的 `capabilities` / `allowed-tools` 是**自由文本元数据**，描述本技能的攻击面，**不是权限授予或沙箱边界**——拥有某项能力意味着技能*可以*做它，而非*只能*做它。本技能**没有**独立于 daemon token / 安全模式 / 路径限制之外的形式化权限系统；任何\"最小权限\"诉求须通过**部署模式选择（安全模式 + 沙箱）+ 操作范围约定**在编排层落实。详见 SKILL.md「🔐 权限模型与最低特权」。\n\n### 依赖安全说明（CVE-2025-59288）\n\n本项目直接依赖 `playwright-core@1.61.1`（Playwright 核心库，无浏览器下载逻辑，≥ 1.55.1）。CVE-2025-59288 影响的是完整 `playwright` 包**下载并安装浏览器**的安装脚本（`curl -k` 未校验证书）；本工具仅依赖 `playwright-core`，**根本不包含浏览器下载代码**，且运行时通过 `channel: 'chrome'` 复用系统已安装的 Chrome/Edge，因此该漏洞在本工具的使用路径上完全不可触发。\n\n## 作为 AI Skill 使用\n\n本工具可作为 AI 助手的 skill 使用。将整个目录放入 skill 安装路径，AI 助手通过 `SKILL.md` 中的规则指导自动化操作。\n\n## License\n\n[MIT](LICENSE)\n\nFile v1.0.16:_meta.json\n\n{\n  \"ownerId\": \"kn74y8h6hjpvfa40rgznyt8e3184yaqe\",\n  \"slug\": \"playwright-browser-use\",\n  \"version\": \"1.0.16\",\n  \"publishedAt\": 1785600450572\n}\n\nFile v1.0.16:references/pagination.md\n\n# 分页策略\n\n翻页前必须先**识别页面分页类型**，选对翻页方式。\n\n> 📝 **文档语言与本地化**：本文件为**简体中文**。识别信号表中的中文 / 英文关键词（\"下一页\"/\"Next\"、\"加载更多\"/\"Load more\"）仅为**启发式示例，并非穷举**——非中文页面的实际文案会不同。跨语言页面请优先用 `snap` 返回的 `ref` 或 CSS 选择器（`page.locator('.xxx')`）定位，避免依赖可见文本。完整语言说明见 `SKILL.md`「📝 文档语言与本地化说明」。\n\n## 步骤 1：识别分页类型\n\n从 `pw-browser snap` 的输出判断：\n\n| 类型 | 识别信号 | 翻页方式 |\n|------|---------|---------|\n| **页码分页** | 底部有页码（1,2,3...）、\"下一页\"/\"Next\"/\">\" | `click` 页码或\"下一页\" |\n| **无限滚动** | 底部无分页控件，内容随滚动增加 | `mousewheel` |\n| **加载更多** | 底部有\"加载更多\"/\"Load more\"/\"查看更多\" | `click` 该按钮 |\n\n### 常见网站参考\n\n| 网站 | 分页类型 | 翻页方式 |\n|------|---------|---------|\n| 百度搜索 | 页码分页 | 点击页码 |\n| 淘宝搜索 | 页码分页 | 点击页码 |\n| 京东搜索 | 页码分页 | 点击页码 |\n| 知乎 | 页码分页 | 点击页码 |\n| 小红书 | 无限滚动 | 滚动加载 |\n| 抖音 | 无限滚动 | 滚动加载 |\n\n## 步骤 2：执行翻页\n\n### A. 页码分页\n\n```bash\n# 从 snap 找到\"下一页\"按钮 ref\npw-browser snap\npw-browser click e42          # 点击\"下一页\"\npw-browser sleep 2\npw-browser snap               # 验证\n```\n\n备选方式 — 直接点页码：\n\n```bash\npw-browser snap\n# 找到页码数字（如 \"2\"）对应的 ref\npw-browser click e50\npw-browser sleep 2 && pw-browser snap\n```\n\n### B. 无限滚动\n\n```bash\npw-browser mousewheel 0 800\npw-browser sleep 2\npw-browser snap\n```\n\n连续多次滚动直到内容不再增加。\n\n### C. 加载更多按钮\n\n```bash\npw-browser snap\n# 找到按钮 ref\npw-browser click e30\npw-browser sleep 2\npw-browser snap\n```\n\n## 步骤 3：判断翻页成功\n\n| 分页方式 | 成功信号 | 结束信号 |\n|---------|---------|---------|\n| 页码 | snap 内容变化，URL 可能变化 | \"下一页\"按钮消失或 disabled |\n| 滚动 | snap 出现新元素 | 内容不变，出现\"没有更多了\" |\n| 按钮 | 新内容加载 | 出现\"已加载全部\"，按钮消失 |\n\n## 批量翻页提取\n\n```bash\n# 使用 run-code 批量翻页\npw-browser run-code \"\n  const allResults = [];\n  let hasNext = true;\n  while (hasNext) {\n    const items = await page.locator('.item').all();\n    for (const item of items) {\n      allResults.push(await item.textContent());\n    }\n    const nextBtn = page.locator('text=下一页');\n    if (await nextBtn.count() === 0 || await nextBtn.isDisabled()) {\n      hasNext = false;\n    } else {\n      await nextBtn.click();\n      await page.waitForTimeout(2000);\n    }\n  }\n  return JSON.stringify(allResults);\n\"\n```\n\n## 常见问题\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| 滚动后不加载新内容 | 实际上是页码分页 | 检查 snap 底部是否有页码，改为 click |\n| 点击页码没反应 | 按钮 disabled 或需要等待 | `sleep 1` 后再点击 |\n| 翻页后内容相同 | AJAX 加载，需要等待 | 延长 sleep 时间或用 `wait-for` |\n| 页码按钮被遮挡 | 需先滚动到底部 | `mousewheel 0 1000` 再 snap |\n\nFile v1.0.16:references/rich-text-editor.md\n\n# SPA 与富文本编辑器 ⚠️\n\n> ⚠️ **破坏性操作警告：** 以下操作会真实修改网页内容（知识库文档、CMS 页面等）。执行前确认：① 处于编辑/草稿状态而非已发布内容；② 修改内容已经用户确认；③ 保存/发布操作不可逆。\n\n处理知识库、文档系统、CMS（如 Notion/语雀/飞书类页面）中的 SPA 编辑态和富文本编辑器写入。\n\n## 何时使用\n\n满足**任一条件**时，遵循本指南：\n\n- URL/页面属于知识库、文档、笔记、CMS 类站点\n- 任务要求创建/编辑/保存文档正文\n- 点击\"编辑\"按钮后 URL 不变但页面状态变化\n- snap 中出现 `contenteditable`、编辑器 toolbar、\"插入\"/\"正文\"等\n- 表单不是普通 input/textarea，而是复杂编辑器\n\n## 核心原则\n\n- **不要**直接用 `innerText`/`textContent` 写 RTE——不会被编辑器状态机接受\n- 先确认进入编辑态，再用 Playwright 键盘输入\n- 保存后验证内容而不是只看按钮状态\n\n## 流程\n\n### 1. 进入编辑态\n\n```bash\n# 先 snap 找到编辑按钮\npw-browser snap\n\n# 点击编辑按钮\npw-browser click <编辑ref>\n\n# 验证进入编辑态\npw-browser run-code \"\n  return await page.evaluate(() => ({\n    hasUpdate: Array.from(document.querySelectorAll('button'))\n      .some(b => b.textContent.trim() === '更新'),\n    editableCount: document.querySelectorAll('[contenteditable]').length\n  }));\n\"\n```\n\n如果 `hasUpdate=true` 或 `editableCount > 0`，继续；否则尝试重试点击。\n\n### 2. 写入内容（RTE）\n\n```bash\npw-browser run-code \"\n  const editor = page.locator('[contenteditable=\\\"true\\\"], [contenteditable=\\\"plaintext-only\\\"]').first();\n  await editor.click();\n  await page.keyboard.press('Control+A');\n  await page.keyboard.type('要写入的文本内容');\n  await page.waitForTimeout(1000);\n  const text = await editor.textContent();\n  return text;\n\"\n```\n\n### 3. 保存 ⚠️\n\n> ⚠️ 保存/发布操作不可逆，确认内容无误后再执行。\n\n```bash\npw-browser run-code \"\n  await page.evaluate(() => {\n    const btn = Array.from(document.querySelectorAll('button'))\n      .find(b => ['更新','保存','发布','完成'].includes(b.textContent.trim()));\n    btn?.click();\n  });\n  await page.waitForTimeout(3000);\n\"\n```\n\n### 4. 验证\n\n```bash\npw-browser run-code \"\n  const title = document.querySelector('h1, [class*=title]')?.textContent || document.title;\n  const main = document.querySelector('main') || document.body;\n  const blocks = Array.from(main.querySelectorAll('p, h1, h2, h3, li'))\n    .map(b => b.textContent.trim().slice(0, 80))\n    .filter(Boolean);\n  return JSON.stringify({ title, sampleBlocks: blocks.slice(0, 10) });\n\"\n```\n\n成功标准：\n- `hasUpdate=false`（编辑态已退出）\n- 正文包含目标文本\n- 标题未被误改或清空\n- 没有重复写入的文本\n\n## 弹窗处理\n\n- DOM 浮层（弹窗、抽屉、popover）：通过 snap 识别并 click 关闭按钮\n- 原生 JS dialog（alert/confirm/prompt）：用 `pw-browser dialog-accept` / `dialog-dismiss`\n\n## 失败处理\n\n- 点击\"编辑\"超时后，先检查是否已进入编辑态，不要重复点击\n- 后续命令超时，执行 `pw-browser recover` 恢复 daemon\n- 恢复后如果已在编辑态，继续输入和保存\n\nFile v1.0.16:references/running-code.md\n\n# 运行自定义代码（`run-code`）⚠️\n\n> ⚠️ **安全警告：** `run-code` 在 daemon 进程的 **受限沙箱（`vm` 模块）** 中执行代码，**无法直接调用** Node.js 系统 API（`fs` / `child_process` / `process` / `require`）。但**浏览器上下文本身可经下载（`download.saveAs`）或文件上传（`setInputFiles`）在本地磁盘读写文件，并能发起任意网络请求**——这些由注入的 `page` 句柄提供，沙箱不阻止。因此它**仍能把数据持久化到本地磁盘**（见下方\"文件下载\"小节）。仅在用户明确指定任务中使用，不要执行来源不明的代码片段。\n\n当内置命令不够用时，用 `pw-browser run-code` 执行 Playwright 代码。\n\n## `eval`：页面上下文任意 JavaScript 执行 ⚠️\n\n> ⚠️ **安全警告：** `eval` 在**当前页面的 JavaScript 上下文**中执行你提供的任意代码（等价于在浏览器开发者工具控制台里直接输入并执行）。它能读取 `document.cookie`、`localStorage`/`sessionStorage`、发起**携带当前页面凭证**的 `fetch`/`XMLHttpRequest`，并直接操控 DOM、触发点击与表单提交。**它不等于\"执行一个无害的 JS 表达式\"——它是完整的页面级代码执行（page-context RCE）。**\n>\n> `eval` 与 `run-code` 是同一类\"代码执行\"能力，只是作用域不同：\n> - **`eval`** → 页面上下文，能触及页面里的所有数据与会话凭证\n> - **`run-code`** → Node 沙箱上下文，能驱动浏览器但拿不到宿主机 `fs`/`process`\n>\n> 两者都受 daemon **token 认证**保护（未持 token 的外部进程无法调用），且都在 `PW_BROWSER_SAFE_MODE=1` 启动时**被禁用**。仅在用户明确指定的任务、且目标页面可信时使用；不要对来源不明或高权限页面执行。\n\n```bash\n# 读取当前页面所有 cookie（含会话令牌）\npw-browser eval \"document.cookie\"\n\n# 读取 localStorage\npw-browser eval \"JSON.stringify(localStorage)\"\n\n# 在指定元素上下文执行（ref 来自 snap）\npw-browser eval \"el.innerText\" e5\n\n# 发起带页面凭证的请求（可被滥用于 CSRF / 数据外泄，慎用）\npw-browser eval \"await (await fetch('/api/me')).text()\"\n```\n\n## 语法\n\n```bash\npw-browser run-code \"<code>\"\n```\n\n代码在 daemon 的 `vm` 沙箱中执行，`page` 对象已注入（标准的 Playwright Page）。沙箱仅暴露 `page` 和安全 JS 全局，不提供 Node 系统模块。\n\n## 等待策略\n\n```bash\n# 等待网络空闲\npw-browser run-code \"await page.waitForLoadState('networkidle');\"\n\n# 等待元素出现\npw-browser run-code \"await page.locator('.loading').waitFor({ state: 'hidden' });\"\n\n# 等待自定义条件\npw-browser run-code \"await page.waitForFunction(() => window.appReady === true);\"\n\n# 带超时的等待\npw-browser run-code \"await page.locator('.result').waitFor({ timeout: 10000 });\"\n```\n\n## 页面信息\n\n```bash\n# 获取标题\npw-browser run-code \"return await page.title();\"\n\n# 获取 URL\npw-browser run-code \"return page.url();\"\n\n# 获取整个 HTML\npw-browser run-code \"return await page.content();\"\n\n# 视口大小\npw-browser run-code \"return JSON.stringify(page.viewportSize());\"\n```\n\n## 在页面中执行 JS（evaluate）\n\n```bash\n# 获取 userAgent\npw-browser run-code \"return await page.evaluate(() => navigator.userAgent);\"\n\n# 获取所有链接\npw-browser run-code \"\n  return await page.evaluate(() =>\n    [...document.querySelectorAll('a')].map(a => ({ text: a.textContent.trim(), href: a.href }))\n  );\n\"\n\n# 获取 localStorage\npw-browser run-code \"return await page.evaluate(() => JSON.stringify(localStorage));\"\n```\n\n## Iframe 操作\n\n```bash\npw-browser run-code \"\n  const frame = page.frameLocator('iframe#my-iframe');\n  await frame.locator('button.submit').click();\n\"\n```\n\n## 文件下载 ⚠️\n\n> ⚠️ 文件将写入本地磁盘，注意目标路径，避免覆盖已有文件。\n\n```bash\npw-browser run-code \"\n  const [download] = await Promise.all([\n    page.waitForEvent('download'),\n    page.locator('text=下载').click()\n  ]);\n  await download.saveAs('./downloaded-file.pdf');\n  return download.suggestedFilename();\n\"\n```\n\n## 错误处理\n\n```bash\npw-browser run-code \"\n  try {\n    await page.locator('button.submit').click({ timeout: 3000 });\n    return 'clicked';\n  } catch (e) {\n    return 'element not found: ' + e.message;\n  }\n\"\n```\n\n## 复杂场景：多页数据采集\n\n```bash\npw-browser run-code \"\n  const results = [];\n  for (let i = 1; i <= 5; i++) {\n    await page.goto('https://example.com/page/' + i);\n    const items = await page.locator('.item').allTextContents();\n    results.push(...items);\n  }\n  return JSON.stringify(results);\n\"\n```\n\n## 复杂场景：表单填写 ⚠️\n\n> ⚠️ 此操作会真实提交表单，可能触发实际的业务操作（注册账号、下单、发送消息等）。执行前确认目标页面和表单内容已经用户确认。\n\n```bash\npw-browser run-code \"\n  await page.fill('#name', '张三');\n  await page.fill('#email', 'test@example.com');\n  await page.selectOption('#city', '北京');\n  await page.check('#agree');\n  await page.locator('button[type=submit]').click();\n  await page.waitForURL('**/success');\n  return 'form submitted';\n\"\n```\n\n## 注意事项\n\n- `run-code` 中直接使用 Playwright API，无需额外的 `page.evaluate` 包装\n- 客户端请求超时为 120 秒，长耗时操作请合理拆分\n- 返回值自动序列化为字符串，复杂对象请用 `JSON.stringify()`\n- 如果代码中有引号冲突，优先用单引号包裹 JS 字符串\n\nFile v1.0.16:CHANGELOG.md\n\n# Changelog\n\nAll notable changes to `pw-browser` are documented here. This project follows semver-ish versioning (`MAJOR.MINOR.PATCH`).\n\n## [1.3.11] — 2026-07-31\n\n### Bug fixes\n\n- **`act` DOM 变化检测遗漏\"元素消失\"**：此前只检测新元素出现（`appeared`），页面已有元素被移除或文本变更后 `act` 序列仍会用旧 ref 继续执行，导致 `ElementNotFound` 或误操作。现同时检测 `disappeared`（旧 ref 在新快照中消失），任一条件满足即中断序列并返回最新快照。\n- **`run-code` 的 `vm.Script` 编译失败未返回结构化错误**：用户传入非法 JS 时，`new vm.Script()` 抛出的 SyntaxError 会被外层 `catch` 吞掉，daemon 返回 500 而非带 `kind: 'SyntaxError'` 的 JSON 错误。现在将编译步骤提前到独立 try/catch，失败时立即通过 `json()` 返回结构化错误并记录 `elapsedMs`。\n- **客户端请求超时后 `res.on('end')` 可能双重回调**：`req.destroy()` 触发 `timeout` 事件后，`res` 端仍可能发来数据并触发 `end` 事件，导致 `resolve` 和 `reject` 都被调用（unhandled rejection 或错误结果）。引入 `settled` 守卫，确保 `resolve`/`reject` 仅被调用一次。\n\n### Performance\n\n- **`findElement` Strategy 1+2 合并**：`getByRole` 有 name 和无 name 原本是两段独立 if 块，会构造两次 Playwright locator。合并为一次查询（有 name 且 < 100 字符时带 name，否则不带），减少不必要的 locator 构造。\n- **`ensurePage()` 过滤已关闭页面**：`context.pages()[0]` 可能拿到刚被关闭但尚未从 context 中移除的页面，后续操作会报错。改为 `filter(p => !p.isClosed())` 确保拿到的是有效页面。\n\n### 文档\n\n- **规则 1 澄清**：明确哪些情况必须 snap（每次导航/页面变更后），哪些情况可复用 ref（同一页面内连续操作），消除「每步 snap」与「不必每步 snap」之间的歧义\n- **状态感知表 snap 描述修正**：ref 跨 snap 稳定的前提是「同一页面会话」，每次导航后必须重新 snap，旧 ref 不可复用\n- **省 Token 建议补充导航条件**：「复用稳定 ref，不必每步重新 snap」改为「同一页面内复用稳定 ref，若页面发生导航（URL 变化、新页面打开），必须重新 snap」\n- **storage 风险文档增强**：补充 SAFE_MODE=1 作为完全禁用 cookies/storage 的兜底措施说明，与代码实现在 v1.3.2+ 的行为对齐\n\n## [1.3.10] — 2026-07-30\n\nFile v1.0.16:QUICKSTART.en.md\n\n# pw-browser Quickstart Cookbook\n\n> End-to-end examples for AI agents / scripts. Every example can be copied and run as-is.\n> Convention: `pw-browser` stands for the full command\n> `NODE_PATH=\"<SKILL_DIR>/node_modules\" node \"<SKILL_DIR>/pw-browser.js\"`.\n> All examples assume a daemon is already running in the background (`pw-browser daemon &`).\n\n---\n\n## Example 1: Fill a form and submit (most common)\n\nGoal: open a page, locate the input and button, type text and submit.\n\n```bash\n# Open the page\npw-browser open https://example.com/login\n\n# You MUST snap first to get the ref table\npw-browser snap\n# → e1 input placeholder=\"Username\"\n# → e2 input placeholder=\"Password\"\n# → e5 button \"Sign in\"\n\n# Fill and click (using the refs from snap)\npw-browser fill e1 \"alice\"\npw-browser fill e2 \"s3cret\"\npw-browser click e5\n\n# Verify\npw-browser wait-for \"text=Welcome\" --timeout 8000\npw-browser snap\n```\n\nKey point: `snap` must come before `click`/`fill`, and refs stay stable across snaps (e1 this time is still e1 next time).\n\n---\n\n## Example 2: File upload + download round-trip\n\nGoal: upload a local file, then trigger a download and confirm it landed on disk.\n\n```bash\npw-browser open https://example.com/upload\n\npw-browser snap\n# → e3 input type=file \"Choose file\"\n# → e4 button \"Upload\"\n\n# Upload (multiple files supported, space- or comma-separated)\npw-browser upload e3 /tmp/report.pdf /tmp/appendix.xlsx\npw-browser click e4\n\n# Wait for upload to finish\npw-browser wait-for \"text=Upload complete\" --timeout 10000\n\n# Trigger download: click a link/button that downloads, save to a dir\npw-browser snap\n# → e9 a \"Export CSV\"\npw-browser download e9 --path /tmp/downloads --timeout 30000\n# → { \"ok\": true, \"savedPath\": \"/tmp/downloads/export.csv\", \"suggestedFilename\": \"export.csv\" }\n```\n\nNote: `download` is symmetric to `upload`; without `--path` it saves to the current working directory. It can also be an `act` action:\n`{\"action\":\"download\",\"ref\":\"e9\",\"path\":\"/tmp/downloads\"}`.\n\n> 🔒 v1.3.10+: `suggestedFilename` is chosen by the web page and is untrusted input. When `--path` is a directory, the name is reduced to a `basename` and confined inside it; anything escaping returns `PathTraversal`.\n\n---\n\n## Example 3: Shadow DOM / iframe piercing\n\nGoal: some elements live inside a shadow root or a same-origin iframe where normal selectors can't reach — the tool pierces them automatically.\n\n```bash\npw-browser open https://example.com/widget\n\npw-browser snap\n# → e2 button \"Inner button\"  inShadow:true\n# → e7 button \"iframe submit\"  frameChain:[{sel:\"iframe#frame1\"}]\n\n# Just click! Location uses css >>> piercing + frameLocator automatically, invisible to the caller\npw-browser click e2\npw-browser click e7\n```\n\nKey points:\n- `inShadow: true` means the element is inside a shadow root; a non-empty `frameChain` means it's inside an iframe.\n- `--annotate` screenshots only mark main-document elements (shadow/iframe elements can't be located by xpath for annotation, but remain clickable in the text snapshot).\n- Cross-origin iframes are inaccessible and are skipped automatically.\n\n---\n\n## Example 4: Cookie / localStorage extraction and session restore\n\nGoal: back up the session after login, then restore it next time to skip re-login.\n\n```bash\npw-browser open https://example.com/dashboard\n# (complete login manually or via Example 1 first)\n\n# Back up cookies (defaults to ~/.pw-browser/cookies.json, kept inside the confined dir)\npw-browser cookies export\n# → { \"ok\": true, \"exported\": \"~/.pw-browser/cookies.json\", \"count\": 12,\n#     \"warning\": \"SECURITY: this file holds live session credentials ...\" }\n\n# Back up localStorage (defaults to ~/.pw-browser/localStorage.json)\npw-browser storage export\n\n# —— next session ——\npw-browser open https://example.com/dashboard\npw-browser cookies import          # reads the default ~/.pw-browser/cookies.json\npw-browser storage import          # reads the default ~/.pw-browser/localStorage.json\n\n# Already authenticated after reload\npw-browser reload\npw-browser snap\n```\n\n> ⚠️ `cookies` / `storage` depend on a real http/https page origin; `file://` and `data:` pages don't support cookies, and `localStorage` behaviour there is unreliable.\n\n> 📁 **Path confinement:** `cookies` / `storage` `export`/`import` are **confined to `~/.pw-browser/` by default** — a deliberate safety measure so credentials can't be silently scattered into `/tmp` or loaded from attacker-controlled paths elsewhere. A custom path must also live under `~/.pw-browser/` (paths are resolved through `realpath`, so symlinks can't escape either).\n>\n> 🛡️ **Waiving the confinement is the operator's call, not the caller's (v1.3.8):** earlier versions let any caller lift the guard just by adding `--unsafe` — handing the key to the very party the guard constrains. Escaping now requires **both**: ① the operator (a human) starts the daemon with `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` (a process env var, which a party sending HTTP commands cannot set), and ② the caller passes `--unsafe` (intent). `--unsafe` alone is rejected with `UnsafeOverrideNotPermitted`. Responses carry `confined` and `warning` fields, and the daemon logs every credential-path access to stderr for auditing.\n\n> 🔒 **Security warning (read this):** the exported `cookies.json` / `localStorage.json` contain your **full authenticated session** — potentially `HttpOnly` cookies, bearer/session tokens, CSRF tokens, and other sensitive state. Anyone who obtains the file can **impersonate you** on the site. This tool's daemon **persists browser state across commands**, and the tool also exposes `eval` / `run-code` (which can read cookies/storage from the page context), so the saved material carries a higher risk of being misread, misused, or exfiltrated from your local environment.\n> - **Do not** commit these files to git, upload them, or share them.\n> - Delete them when done (`rm`); if you must keep them briefly, just leave them in the default `~/.pw-browser/` — since v1.3.9 that directory is created `0700` and credential files (plus the daemon auth token) are written `0600`, so **no manual `chmod` is needed**. On Windows POSIX mode bits aren't enforced by the OS; user-profile ACLs apply instead.\n> - If credentials must never touch the disk at all, the operator can start the daemon with `PW_BROWSER_CRED_PERSIST=off`: `export`/`import` are then **refused** (`CredentialPersistenceDisabled`) while in-memory `cookies list` / `storage get|set|clear` keep working — finer-grained than `PW_BROWSER_SAFE_MODE=1`, which disables the credential primitives and code execution wholesale.\n> - Restore a session file **only on your own machine and for the same site** — never reuse it across environments or accounts.\n> - **Importing (restore) is just as risky:** restoring a session grants logged-in / privileged access to that site. Never auto-import from untrusted paths. In agent / automated workflows, **do not** persist credential material across runs by default — only restore deliberately and in a controlled way, to avoid unintended privilege persistence or session theft.\n> - **Serving not-fully-trusted agents:** start the daemon with `PW_BROWSER_SAFE_MODE=1` — since v1.3.2 safe mode **disables all `cookies` / `storage` subcommands entirely** (blocked at the same level as `eval`/`run-code`), closing the session-credential surface at the root.\n\n---\n\n## Example 5: `act` multi-step + self-correction (recommended for agents)\n\nGoal: send a whole sequence of actions to the daemon at once; if the page DOM changes mid-sequence (e.g. clicking a button pops a new menu), the daemon interrupts automatically and returns a fresh snapshot for you to re-plan.\n\n```bash\npw-browser open https://example.com/form\n\npw-browser snap\n# → e1 input \"Title\"\n# → e2 button \"Next\"   (clicking dynamically reveals new fields e3/e4)\n\n# Send the action sequence in one shot\npw-browser act '[\n  {\"action\":\"fill\",\"ref\":\"e1\",\"text\":\"Monthly report\"},\n  {\"action\":\"click\",\"ref\":\"e2\"}\n]'\n# If clicking e2 reveals new elements → returns { interrupted: true, snap: <fresh snapshot> }\n# Continue with the new refs from the returned snapshot:\n#   pw-browser act '[{\"action\":\"fill\",\"ref\":\"e3\",\"text\":\"...\"},{\"action\":\"click\",\"ref\":\"e4\"}]'\n\n# A failed action comes with a diagnosis (whether the ref still exists, similar ref hints) for self-healing\n```\n\nToken-saving tip: `1 snap → 1 act → done`, no round-trips back to the model in between.\n\n---\n\n## Troubleshooting quick reference\n\n| Symptom | Fix |\n|---------|-----|\n| Connection refused | daemon not running → `pw-browser daemon &`; or stale process → clear `~/.pw-browser/daemon.json` + kill port 19223, then restart |\n| `ElementNotFound` | ref expired → re-`snap` for a fresh ref |\n| `NavigationTimeout` | page loads slowly → `snap` to see actual state, or raise `--timeout` |\n| snapshot truncated (`⚠ snapshot truncated`) | page has too many elements, hit the cap → raise `PW_BROWSER_SNAP_LIMIT` or interact first to narrow the page |\n\nFull rules and the complete command table are in `SKILL.md` and `README.md`.\n\nFile v1.0.16:QUICKSTART.md\n\n# pw-browser 快速上手 Cookbook\n\n> 面向 AI Agent / 脚本的端到端示例集。每个示例都可直接照抄执行。\n> 约定：下文用 `pw-browser` 表示完整命令\n> `NODE_PATH=\"<SKILL_DIR>/node_modules\" node \"<SKILL_DIR>/pw-browser.js\"`。\n> 所有示例默认 daemon 已在后台启动（`pw-browser daemon &`）。\n\n---\n\n## 示例 1：填写表单并提交（最常用）\n\n目标：打开一个页面，定位输入框与按钮，填入文本并提交。\n\n```bash\n# 打开页面\npw-browser open https://example.com/login\n\n# 必须先 snap，拿到 ref 表\npw-browser snap\n# → e1 input placeholder=\"用户名\"\n# → e2 input placeholder=\"密码\"\n# → e5 button \"登录\"\n\n# 填写并点击（基于 snap 给的 ref）\npw-browser fill e1 \"alice\"\npw-browser fill e2 \"s3cret\"\npw-browser click e5\n\n# 验证结果\npw-browser wait-for \"text=欢迎\" --timeout 8000\npw-browser snap\n```\n\n要点：`snap` 之后才能 `click`/`fill`，且 ref 跨多次 snap 保持稳定（本次的 e1 下次仍是 e1）。\n\n---\n\n## 示例 2：文件上传 + 下载闭环\n\n目标：上传一个本地文件，再触发一次下载，确认文件落盘。\n\n```bash\npw-browser open https://example.com/upload\n\npw-browser snap\n# → e3 input type=file \"选择文件\"\n# → e4 button \"开始上传\"\n\n# 上传（支持多文件，空格或逗号分隔）\npw-browser upload e3 /tmp/report.pdf /tmp/appendix.xlsx\npw-browser click e4\n\n# 等待上传完成\npw-browser wait-for \"text=上传成功\" --timeout 10000\n\n# 触发下载：点击某个会下载的链接/按钮，保存到指定目录\npw-browser snap\n# → e9 a \"导出 CSV\"\npw-browser download e9 --path /tmp/downloads --timeout 30000\n# → { \"ok\": true, \"savedPath\": \"/tmp/downloads/export.csv\", \"suggestedFilename\": \"export.csv\" }\n```\n\n说明：`download` 对称于 `upload`；不传 `--path` 时存到当前工作目录。也可作为 `act` 动作：\n`{\"action\":\"download\",\"ref\":\"e9\",\"path\":\"/tmp/downloads\"}`。\n\n> 🔒 v1.3.10+：`suggestedFilename` 由网页决定，属不可信输入。当 `--path` 是目录时，该名字会被取 `basename` 并限制在目录内，越界返回 `PathTraversal`。\n\n---\n\n## 示例 3：Shadow DOM / iframe 穿透\n\n目标：页面里有些元素藏在 shadow root 或同源 iframe 内，普通选择器够不着——本工具会自动穿透。\n\n```bash\npw-browser open https://example.com/widget\n\npw-browser snap\n# → e2 button \"内部按钮\"  inShadow:true\n# → e7 button \"iframe 里的提交\"  frameChain:[{sel:\"iframe#frame1\"}]\n\n# 直接点！定位由 css >>> 穿透 + frameLocator 自动完成，外部无感知\npw-browser click e2\npw-browser click e7\n```\n\n要点：\n- `inShadow: true` 表示元素在 shadow root 内；`frameChain` 非空表示在 iframe 内。\n- `--annotate` 截图**只**标注主文档元素（shadow/iframe 元素无法用 xpath 定位标注，但文本快照里照常可点）。\n- 跨域 iframe 不可访问，会自动跳过。\n\n---\n\n## 示例 4：Cookie / localStorage 提取与会话恢复\n\n目标：登录后把会话备份下来，下次免登录直接恢复。\n\n```bash\npw-browser open https://example.com/dashboard\n# （先手动或通过示例 1 完成登录）\n\n# 备份 cookie（默认写入 ~/.pw-browser/cookies.json，受目录限制保护）\npw-browser cookies export\n# → { \"ok\": true, \"exported\": \"~/.pw-browser/cookies.json\", \"count\": 12,\n#     \"warning\": \"SECURITY: this file holds live session credentials ...\" }\n\n# 备份 localStorage（默认写入 ~/.pw-browser/localStorage.json）\npw-browser storage export\n\n# —— 下次新会话 ——\npw-browser open https://example.com/dashboard\npw-browser cookies import          # 读取默认路径 ~/.pw-browser/cookies.json\npw-browser storage import          # 读取默认路径 ~/.pw-browser/localStorage.json\n\n# 刷新后已是登录态\npw-browser reload\npw-browser snap\n```\n\n> ⚠️ `cookies` / `storage` 依赖真实 http/https 页面源；`file://` 与 `data:` 页面不支持 cookie，`localStorage` 行为也不可靠。\n\n> 📁 **路径限制**：`cookies` / `storage` 的 `export`/`import` 默认被限制在 `~/.pw-browser/` 目录内——这是刻意的安全设计，防止凭证被静默散落到 `/tmp` 或加载来自任意路径的攻击者构造文件。自定义路径也必须位于 `~/.pw-browser/` 之下（路径经 realpath 解析，用符号链接也逃不出去）。\n>\n> 🛡️ **越界解除权归操作者，不归调用方（v1.3.8）**：早期版本只要调用方自己加 `--unsafe` 就能解除限制——等于把护栏的钥匙交给被约束的一方。现在必须**同时**满足两个条件才放行：① 操作者（人）以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon（进程环境变量，发 HTTP 命令的一方改不了）；② 调用方显式传 `--unsafe`（表达意图）。只传 `--unsafe` 会被拒绝并返回 `UnsafeOverrideNotPermitted`。命令返回带 `confined`（是否仍在受限目录内）与 `warning` 字段，daemon stderr 会记录每一次凭证路径访问，便于事后审计。\n\n> 🔒 **安全提醒（务必读完）**：导出的 `cookies.json` / `localStorage.json` 含有**完整登录会话凭证**——可能包括 `HttpOnly` cookie、Bearer/会话令牌、CSRF token 等敏感状态。文件一旦泄露，任何人拿到即可**冒用你的身份**登录对应站点。本工具的 daemon 会在多次命令间**持久化浏览器状态**，且工具另提供 `eval` / `run-code` 代码执行能力（可在页面上下文读取 cookie/storage），因此落盘的会话文件被误读、被错用或在本地环境外泄的风险更高。\n> - **不要**把导出的文件提交进 git、上传到网盘或分享给他人。\n> - 用完即删（`rm`）；若需暂存，放在默认的 `~/.pw-browser/` 内即可——自 v1.3.9 起该目录以 `0700` 创建、凭证文件以 `0600` 写入（连 daemon 认证 token 也是），**不必再手工 `chmod`**。Windows 上 POSIX 权限位不由系统强制，依赖用户目录 ACL。\n> - 若整条流程都不该让凭证落盘：操作者用 `PW_BROWSER_CRED_PERSIST=off` 启动 daemon，`export`/`import` 会被**强制拒绝**（`CredentialPersistenceDisabled`），而 `cookies list` / `storage get|set|clear` 等内存态操作不受影响。这比 `PW_BROWSER_SAFE_MODE=1`（整体禁用凭证原语与代码执行）粒度更细。\n> - 同一份会话文件**只应在你自己的本机、同一站点**恢复，切勿跨环境/跨账号复用。\n> - **导入（恢复）同样高风险**：恢复会话即赋予对应站点的登录态/特权访问。切勿从不可信路径自动导入；在 Agent / 自动化流程里，**不要**默认把凭证材料在多次运行之间持久化，以免意外固化特权访问或导致会话被盗用——只在明确需要、且受控的前提下才恢复。\n> - **面向不完全可信的 agent**：用 `PW_BROWSER_SAFE_MODE=1` 启动 daemon——自 v1.3.2 起安全模式会**整体禁用**全部 `cookies` / `storage` 子命令（与 `eval`/`run-code` 同级拦截），从根上关闭会话凭证读写面。\n\n---\n\n## 示例 5：`act` 多步 + 自纠错（推荐用于 Agent）\n\n目标：把一连串操作一次性发给 daemon，过程中若页面 DOM 变化（例如点了按钮弹出新菜单），daemon 自动中断并返回最新快照，供你重新规划。\n\n```bash\npw-browser open https://example.com/form\n\npw-browser snap\n# → e1 input \"标题\"\n# → e2 button \"下一步\"   (点了会动态出现 e3/e4 新字段)\n\n# 一次性下发动作序列\npw-browser act '[\n  {\"action\":\"fill\",\"ref\":\"e1\",\"text\":\"月度报告\"},\n  {\"action\":\"click\",\"ref\":\"e2\"}\n]'\n# 若点击 e2 后页面出现新元素 → 返回 { interrupted: true, snap: <最新快照> }\n# 此时用返回的 snap 里的新 ref 继续：\n#   pw-browser act '[{\"action\":\"fill\",\"ref\":\"e3\",\"text\":\"...\"},{\"action\":\"click\",\"ref\":\"e4\"}]'\n\n# 失败的动作会附带 diagnosis（ref 是否还在、相似 ref 建议），便于自愈\n```\n\n省 token 提示：`1 次 snap → 1 次 act → 完事`，中间不回模型。\n\n---\n\n## 故障排查速查\n\n| 现象 | 处理 |\n|------|------|\n| 连接拒绝 | daemon 没起 → `pw-browser daemon &`；或旧进程残留 → 清 `~/.pw-browser/daemon.json` + 杀 19223 端口后重启 |\n| `ElementNotFound` | ref 失效 → 重新 `snap` 拿新 ref |\n| `NavigationTimeout` | 页面加载慢 → 先 `snap` 看实际状态，或调大 `--timeout` |\n| 快照被截断（`⚠ snapshot truncated`） | 页面元素过多触发上限 → 调大 `PW_BROWSER_SNAP_LIMIT` 或先交互缩小页面范围 |\n\n详细规则与完整命令表见 `SKILL.md` 与 `README.md`。\n\nFile v1.0.16:README.en.md\n\n# pw-browser\n\n> A Playwright-based browser automation CLI — daemon + client architecture that drives your system's Chrome/Edge directly, with no extra browser download required.\n\n> 📘 This is the English documentation. 中文文档见 [`README.md`](./README.md). The agent-facing instructions live in [`SKILL.md`](./SKILL.md), which is currently Chinese-only; a Chinese-reading agent (or the user) can follow it. For end-to-end examples see [`QUICKSTART.en.md`](./QUICKSTART.en.md).\n\n> ⚠️ **Full capabilities (includes code execution & credential manipulation)**: this tool is **NOT just \"open web pages\"** — it ships two **code-execution** abilities, `eval` (arbitrary in-page **JavaScript**) and `run-code` (Playwright/Node code in the daemon context), plus `cookies` / `storage` **credential read/write primitives** that can extract/inject login state and session tokens **without any code execution**. All are controlled by a persistent local daemon (127.0.0.1:19223), gated by a daemon token; `PW_BROWSER_SAFE_MODE=1` **fully disables** code-exec and credential access (v1.3.2+). Read \"🔐 Permission Model & Least Privilege\" and \"⚠️ Security Boundaries\" below to understand the full attack surface and applicable boundaries before authorizing.\n\n## Features\n\n- **System browser reuse**: connects to your installed Chrome or Edge via Playwright's `channel: 'chrome'`, no Chromium download.\n- **Cross-platform**: Windows, macOS, Linux.\n- **Persistent session**: the daemon keeps browser state alive across commands — no restart per call.\n- **Accessibility snapshot**: `snap` produces a page element tree with stable `ref` references, no CSS selectors to write.\n- **Non-headless**: the browser window is always visible so the user can monitor every action in real time.\n- **Human-in-the-loop**: automatically hands off to the user for login / CAPTCHA.\n\n## Installation\n\n```bash\ngit clone <repo-url> pw-browser\ncd pw-browser\nnpm install\n```\n\n> Requires: [Node.js](https://nodejs.org/) 18+ and Chrome or Edge installed on the system.\n\n## Quick start\n\n```bash\n# 1. Start the daemon (background)\nnode pw-browser.js daemon &\n\n# 2. Wait for the daemon to be ready\nsleep 4\n\n# 3. Open a page\nnode pw-browser.js open https://www.baidu.com\n\n# 4. Take a snapshot (required! snap before every interaction)\nnode pw-browser.js snap\n\n# 5. Interact — using the e0, e1, e2... refs from the snapshot\nnode pw-browser.js fill e12 \"weather forecast\"\nnode pw-browser.js click e13\n\n# 6. Check the result\nsleep 2\nnode pw-browser.js snap\n\n# 7. Close\nnode pw-browser.js close --all\n```\n\n## Command reference\n\n| Category | Command | Description |\n|----------|---------|-------------|\n| Lifecycle | `init` / `open <url>` / `close` / `close --all` / `recover` | start, navigate, close, recover |\n| State | `snap` / `wait-for <target>` | snapshot (**refs stay stable across snaps**, same element = same ref) |\n| Interaction | `click <ref>` / `fill <ref> \"text\"` / `type \"text\"` / `press <key>` / `hover <ref>` / `select <ref> <opt>` / `check <ref>` / `uncheck <ref>` / `upload <ref> <file...>` / `drag <ref-from> <ref-to>` / `download <ref> [--path dir]` | click, fill, press, upload, drag, download, etc. |\n| Cookie/Storage | `cookies list` / `export [--path f]` / `import <file>` / `clear` / `set <name> <value> [--domain d]` · `storage get [key]` / `set <k> <v>` / `clear` / `export [--path f]` / `import <file>` | read/write cookies and localStorage (open a real http/https page first) |\n| Navigation | `goto <url>` / `go-back` / `go-forward` / `reload` | page navigation |\n| Tabs | `tab list` / `tab select <idx>` / `tab close <idx>` | multi-tab management |\n| Batch | `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` | run an action sequence at once; auto-interrupts and re-snaps on DOM change |\n| History | `history [--limit N] [--clear]` | view/clear operation history |\n| Advanced | `screenshot [--annotate]` / `mousewheel <dx> <dy>` / `eval \"<expr>\" [ref]` ⚠️ / `run-code \"<code>\"` ⚠️ | screenshot (`--annotate` overlays numbered boxes matched to refs), scroll, code execution. Without `ref`, `eval` runs against the page global scope (`scope: \"page\"`); with `ref`, the identifier **`el`** is bound to that DOM node (`scope: \"element\"`), e.g. `eval \"el.textContent\" e3`. Return values must be JSON-serializable |\n| Delay | `sleep <seconds>` | wait |\n| Dialog | `dialog-accept [text]` / `dialog-dismiss` | handle native alert/confirm/prompt dialogs |\n| Daemon | `shutdown` | stop the persistent daemon and free the browser process |\n\n> ⚠️ `eval` runs JS in the **browser context** (no Node access); `run-code` runs Playwright code in a **restricted `vm` sandbox** in the daemon (no direct Node `fs`/`process`/`child_process`, but it can read/write local files via browser download/upload and issue arbitrary network requests). Both have full browser control. Use only for explicitly authorized tasks.\n\n## Architecture\n\n```\n┌──────────────┐     HTTP (localhost:19223)     ┌──────────────┐\n│  pw-browser  │ ──────────────────────────────→│   Daemon     │\n│  (CLI client)│                                │  (browser)    │\n└──────────────┘                                └──────┬───────┘\n                                                       │\n                                                       ├─ Playwright\n                                                       ├─ Chrome / Edge\n                                                       └─ persistent page state\n```\n\nThe daemon runs continuously after start; browser and page state persist across commands. The CLI calls the daemon over HTTP each time.\n\n## Core workflow rules\n\n1. **Observe before you act**: `snap` before every interaction, act on the snapshot refs.\n2. **Login handoff**: on a login/CAPTCHA page, tell the user to act manually; continue after they confirm.\n3. **Identify pagination type before paging**: numbered pages / infinite scroll / \"load more\" each need a different approach.\n4. **Daemon recovery**: on connection error, kill the process on port 19223 → delete `daemon.json` → restart.\n\nSee `SKILL.md` for detailed rules and examples.\n\n## Enhanced capabilities (inspired by browser-use)\n\nThis tool is positioned as a \"**secure, controllable browser CLI base + persistent daemon; planning is left to the external AI**\", rather than an \"all-in-one brain+hand\" agent framework like browser-use. We borrowed four capabilities from browser-use that are also valuable for a CLI base:\n\n### A. Stable element references (stable ref)\n- `snap` no longer re-numbers refs each time; it builds a persistent `stableKey → ref` map by the element's **semantic identity** (text/placeholder/aria-label when present, else a DOM branch-path hash).\n- **The same logical element keeps the same ref across snaps**, so the external AI can reuse a previous ref without re-snapping every step (e.g. after filling a form, e13 is still that submit button).\n- `findElement` adds **Strategy 0: precise xpath first** — pins the element by the exact xpath recorded in the snapshot, eliminating the \"same-name element mis-click\" ambiguity of pure semantic lookup.\n\n### B. Visual-aided screenshot (`screenshot --annotate`)\n- Injects an overlay drawing a red border + numbered label for **each element matched to a snap ref** (`document.evaluate(xpath)` for precise positioning).\n- The overlay is removed after the screenshot. A multimodal AI can \"read the numbers off the picture\" to locate `e13` spatially,弥补ing the lack of spatial info in a plain-text snapshot.\n\n### C. Action sequence + self-correction (`act`)\n- `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` sends a sequence at once, reusing `executeSingle` step by step.\n- Before each non-first action it re-snaps and compares the `branchPathHash` set: if the page shows **new elements** (DOM change), it sets `interrupted=true` and returns a fresh snapshot for the external AI to re-plan — mirroring browser-use's `multi_act` mid-sequence interrupt.\n- A failed action carries a `diagnosis` (whether the ref still exists, similar-ref hints) for self-healing.\n\n### D. Structured input + operation history (`history`)\n- `act` takes a structured JSON action array (instead of scattered subcommands), lowering the external AI's CLI-arg error rate.\n- `history` records `{ ts, cmd, params (token filtered), ok, elapsedMs }` per operation, capped at 500, with `--clear`. The external AI can review \"what I clicked, which step failed\" — compressing context and enabling review, echoing browser-use's Mem0-style history compression (here with a lightweight local history instead of a vector store).\n\n### E. Shadow DOM / iframe piercing\n- The snapshot recurses into **open shadow roots** and **same-origin iframes**; those elements appear in the ref table and can be directly `click`/`fill`/`upload`/`drag`.\n- A ref carries `inShadow: true` (inside shadow) or `frameChain` (iframe chain); location is done automatically via `css >>>` piercing + `frameLocator`, invisible to the external AI.\n- Cross-origin iframes are inaccessible and skipped; `--annotate` screenshots only mark main-document elements (shadow/iframe elements can't be located by xpath for annotation, but remain clickable in the text snapshot).\n\n### F. File upload and drag\n- `upload <ref> <file1> [file2 ...]`: sets files on an `<input type=\"file\">` (multiple supported).\n- `drag <ref-from> <ref-to>`: drag via Playwright `dragTo`.\n- `download <ref> [--path dir]`: symmetric to `upload` — optionally click `<ref>` to trigger a download, saved to `--path` (default cwd); also usable as an `act` action.\n  > 🔒 v1.3.10+: the page-supplied suggested filename is reduced to a bare `basename` and confined inside `--path` (out-of-bounds returns `PathTraversal`), so a hostile page can no longer use a name like `../../.bashrc` to write outside the directory. If `--path` names an explicit file, it is used verbatim.\n\n### G. Cookies and local storage (first-class commands)\n- `cookies list|export|import|clear|set`: read/write the current context's cookies, no more `eval` workarounds.\n- `storage get|set|clear|export|import`: read/write based on `localStorage`.\n- Typical use: `cookies export` after login to back up the session, `cookies import` next time to skip re-login.\n- 📁 **Path confinement (operator-gated):** `export`/`import` are confined to `~/.pw-browser/` by default (so credentials can't be scattered into `/tmp` or loaded from attacker-controlled paths elsewhere; paths are resolved through `realpath`, so symlinks can't escape). Since v1.3.8 the confinement **cannot be waived by the caller**: escaping it requires **both** ① the operator starting the daemon with `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1`, and ② the caller passing `--unsafe`. `--unsafe` alone returns `UnsafeOverrideNotPermitted`. Responses carry `confined` and `warning` fields, and every credential-path access is logged to the daemon's stderr.\n- 🔒 **Session-persistence risk (Rogue Agent / Medium):** `export`/`import` let the login state be written to disk and restored across runs — convenient, but a rogue/misbehaving agent can use it for **privilege persistence**. Treat exported session files as secrets: delete when done; if kept, just leave them under `~/.pw-browser/` — since v1.3.9 that directory is created `0700` and both credential dumps and the daemon token file are written `0600`, so no manual `chmod` is needed (on Windows POSIX mode bits aren't enforced; user-profile ACLs apply). To forbid credentials on disk entirely, start the daemon with `PW_BROWSER_CRED_PERSIST=off`: `export`/`import` are **blocked** (`CredentialPersistenceDisabled`) while `list`/`get`/`set`/`clear` keep working — finer-grained than safe mode, for long-running agents that need full automation but must not leave restorable sessions behind. `cookies clear` / `storage clear` and `shutdown` the daemon when not needed (idle auto-exit defaults to 15 min). For not-fully-trusted callers, start the daemon with `PW_BROWSER_SAFE_MODE=1` — since v1.3.2 it **disables all cookies/storage subcommands entirely** (credential primitives blocked at the same level as code execution); pair with sandbox isolation if needed.\n\n## Token saving & daemon lifecycle\n\n**No \"LLM call per step\" token black hole**: this skill is just a deterministic executor + persistent daemon; the model only lives in the external orchestration layer and only spends tokens when it actively calls. Keep these habits to stay token-efficient:\n\n- Plan with a text `snap` (ref table, tens of tokens) instead of a `screenshot` (hundreds of KB base64) fed to a vision model every step.\n- Reuse stable refs across snaps instead of re-snapping every step.\n- Batch a string of actions into one `act '<json>'`; the daemon self-corrects internally without returning to the model.\n\n**The daemon is intentionally persistent** (reuses one browser across commands, avoiding reopening Chrome each time). Therefore:\n\n- It's normal to still see the browser window after a task ends — the daemon still holds it.\n- Stop explicitly: `pw-browser shutdown` (hardened — even if `browser.close()` hangs it times out and exits, never stuck/hanging the client).\n- **Idle auto-exit**: by default the daemon **auto-closes the browser and exits after 15 minutes idle**; `PW_BROWSER_IDLE_MS` adjusts this (`0` disables). No manual cleanup needed after you leave.\n- **Configurable port + conflict avoidance**: default `127.0.0.1:19223`, override with `PW_BROWSER_PORT`. On startup, if a daemon is already alive on the target port it exits immediately (no double-start); if the port is taken by something else it auto-increments and writes the actual port back to `~/.pw-browser/daemon.json`, which the client reads automatically.\n- Prefer `shutdown` or idle exit over force-killing the process to avoid orphaned browser children.\n\n## Security\n\n- The daemon listens on `127.0.0.1:19223`, localhost-only.\n- **Command auth**: every daemon command except the `/health` probe requires a random `token`, generated at daemon start and written to `~/.pw-browser/daemon.json` (default readable only by the current user). The CLI carries it automatically; a local process that hasn't read that file cannot call in — closing the \"unauthenticated HTTP endpoint = RCE surface\" gap.\n- **Safe mode**: `PW_BROWSER_SAFE_MODE=1 node pw-browser.js daemon` fully disables `run-code` / `eval` **and all `cookies` / `storage` credential primitives** (v1.3.2+), keeping only the whitelisted snap/click/fill commands — suitable when serving not-fully-trusted agents.\n- `eval` runs in the browser context; `run-code` runs in a `vm` restricted sandbox (no direct Node system APIs, but it can read/write local files via browser download/upload). Local trusted environments only.\n- The browser runs non-headless so the user can monitor in real time.\n- Not recommended as a public API service; if you must, add auth and an operation whitelist.\n- **Permission model note**: the frontmatter `capabilities` / `allowed-tools` are **free-form metadata** describing this skill's attack surface — **not a permission grant or a sandbox boundary**. Having a capability means the skill *can* do it, not that it is *limited* to it. There is **no** formal permission system beyond the daemon token / safe mode / path confinement; any \"least privilege\" need must be met at the orchestration layer via **deployment mode (safe mode + sandbox) + scoped operation agreements**. See SKILL.md \"🔐 Permission Model & Least Privilege\".\n\n### Dependency security note (CVE-2025-59288)\n\nThis project depends directly on `playwright-core@1.61.1` (Playwright core, no browser-download logic, ≥ 1.55.1). CVE-2025-59288 affects the full `playwright` package's install script that downloads and installs a browser (`curl -k` without cert verification). This tool depends only on `playwright-core`, **contains no browser-download code at all**, and at runtime reuses the system's installed Chrome/Edge via `channel: 'chrome'`. The vulnerability is therefore completely untriggerable on this tool's usage path.\n\n## Using as an AI Skill\n\nThis tool can be used as a skill by an AI assistant. Drop the whole directory into the skill install path and the assistant follows the rules in `SKILL.md` to automate.\n\n## License\n\n[MIT](LICENSE)\n\nFile v1.0.16:skill-card.md\n\n## Description:\n\nPlaywright Browser Use provides a local Playwright-based browser automation CLI for opening pages, capturing snapshots, clicking, filling forms, managing tabs, handling cookies and storage, executing page JavaScript, and running Playwright code through a token-protected local daemon.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[yicko](https://clawhub.ai/user/yicko)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers and agent operators use this skill to automate visible local browser workflows such as navigation, form entry, page inspection, screenshots, downloads, tab management, and controlled Playwright scripting.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Default operation gives a local daemon token broad browser automation, page JavaScript execution, Playwright code execution, file upload or download, and persistent session authority.\n\nMitigation: Install only for trusted local use where the browser is visible and interruptible; prefer PW_BROWSER_SAFE_MODE=1 for agent-driven or untrusted-page workflows.\n\nRisk: Cookie and storage commands can expose, export, import, clear, or set session credentials and authentication state.\n\nMitigation: Use PW_BROWSER_CRED_PERSIST=off when session files should never be written, avoid sensitive logged-in accounts, clear cookies and storage when finished, and shut down the daemon.\n\nRisk: Untrusted automation targets can combine browser reachability with authenticated page state to submit actions or make credentialed requests.\n\nMitigation: Scope operations to user-approved pages, review proposed actions before execution, and add sandbox or network isolation when handling untrusted inputs.\n\n## Reference(s):\n\n- [English README](README.en.md)\n- [English Quickstart](QUICKSTART.en.md)\n- [Pagination Strategy](references/pagination.md)\n- [Rich Text Editor Strategy](references/rich-text-editor.md)\n- [Running Custom Code](references/running-code.md)\n- [Node.js](https://nodejs.org/)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance, files]\n\n**Output Format:** [Markdown guidance with inline shell commands, plus CLI text or JSON responses and generated browser artifacts such as screenshots or downloads.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs may reflect live browser state, authenticated sessions, local files selected for upload, and pages reachable from the user's machine.]\n\n## Skill Version(s):\n\n1.0.16 (source: server release metadata; artifact package.json and CHANGELOG report 1.3.11)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.16:package-lock.json\n\n{\n  \"name\": \"pw-browser\",\n  \"version\": \"1.3.11\",\n  \"lockfileVersion\": 3,\n  \"requires\": true,\n  \"packages\": {\n    \"\": {\n    \"name\": \"pw-browser\",\n    \"version\": \"1.3.11\",\n      \"license\": \"MIT\",\n      \"dependencies\": {\n        \"playwright-core\": \"1.61.1\"\n      },\n      \"engines\": {\n        \"node\": \">=18\"\n      }\n    },\n    \"node_modules/playwright-core\": {\n      \"version\": \"1.61.1\",\n      \"resolved\": \"https://registry.npmjs.org/playwright-core/-/playwright-core-1.61.1.tgz\",\n      \"integrity\": \"sha512-h7Qlt6m4REp25qvIdvbDtVmD4LqVXfpRxhORv9L0jzETM05p4fuPJ3dKyuSXQxDSbXnmS79HAgi9589lGSpLkg==\",\n      \"license\": \"Apache-2.0\",\n      \"bin\": {\n        \"playwright-core\": \"cli.js\"\n      },\n      \"engines\": {\n        \"node\": \">=18\"\n      }\n    }\n  }\n}\n\nArchive v1.0.15: 19 files, 83945 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), CHANGELOG.md (17311b), LICENSE (1062b), MANIFEST.txt (453b), package-lock.json (749b), package.json (1123b), pw-browser (266b), pw-browser.js (76871b), QUICKSTART.en.md (9139b), QUICKSTART.md (8662b), README.en.md (16514b), README.md (15399b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5570b), skill-card.md (2870b), SKILL.md (41372b), _meta.json (142b)\n\nFile v1.0.15:SKILL.md\n\n---\nname: playwright-browser-use\ndescription: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会将其与代码执行一并禁用）；(2) `eval` 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) `run-code` 在守护进程上下文执行 Playwright/Node 代码（vm 沙箱隔离）。全部经持久化本地守护进程（127.0.0.1:19223，浏览器状态跨命令保持）控制，受随机 token 认证保护；`PW_BROWSER_SAFE_MODE=1` 可彻底禁用代码执行与 cookies/storage 凭证读写（v1.3.2+）。仅在可信、用户可见的本地环境中授权使用；会话凭证落盘须遵循后文安全警告。\nallowed-tools: Bash(node:*), Bash(pw-browser:*), Bash(curl:*)\ncapabilities:\n  - \"code-execution: page-context (eval — arbitrary JS in current page)\"\n  - \"code-execution: daemon-vm (run-code — null-prototype VM sandbox, no host fs/process)\"\n  - \"network: arbitrary via browser/page context (credentialed requests possible)\"\n  - \"browser-state: persistent credentialed session across commands\"\n  - \"credential-access: direct read/export/import/clear/set of cookies & localStorage (session tokens) — no code execution needed; gated by daemon token, and fully DISABLED by PW_BROWSER_SAFE_MODE=1 (v1.3.2+)\"\n  - \"file: write to local disk (run-code can trigger downloads; cookies/storage export writes credential files)\"\n# --- Formal permission model (machine-readable) -------------------------\n# NOTE: This skill has NO built-in fine-grained permission system. The\n# fields below are declarative metadata for orchestrators/reviewers, NOT\n# an enforcement boundary. Real constraints come only from:\n#   - daemon token auth (gates NETWORK reachability, not per-capability)\n#   - PW_BROWSER_SAFE_MODE=1 (disables code-exec + credential primitives)\n#   - export/import path confinement to ~/.pw-browser (v1.3.1+); since v1.3.8\n#     the caller-supplied --unsafe flag alone can NOT lift it — the operator\n#     must also start the daemon with PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1\n#   - PW_BROWSER_CRED_PERSIST=off (v1.3.9) removes credential persistence only\n#     (export/import), keeping in-memory cookie/storage automation usable\n# \"capabilities\" lists the MINIMUM attack surface, not a maximum; the skill\n# can additionally drive any site the browser can reach, including ones where\n# the user is already authenticated. Least privilege is achieved at the\n# orchestration layer via deploy mode (safe mode + sandbox) + scoped ops.\npermissions:\n  model: \"none-formal\"          # no built-in RBAC / capability-dropping\n  enforcement:\n    - \"daemon-token\"            # gates network reachability to 127.0.0.1:19223\n    - \"safe-mode-env\"           # PW_BROWSER_SAFE_MODE=1 disables code-exec + creds\n    - \"path-confinement\"        # cookies/storage IO limited to ~/.pw-browser\n    - \"operator-gated-override\" # lifting confinement needs daemon env, not a caller flag\n    - \"cred-persist-killswitch\" # PW_BROWSER_CRED_PERSIST=off blocks export/import only\n    - \"secret-file-permissions\" # daemon token + credential dumps written 0600, dir 0700\n    - \"download-name-sanitised\" # page-supplied download filename forced to a basename inside --path\n  least-privilege:\n    default-mode: \"full\"        # full power; requires trusted, user-visible, local\n    scoped-mode: \"cred-persist-off\"      # PW_BROWSER_CRED_PERSIST=off — keeps automation,\n                                         # removes the privilege-persistence primitive\n    reduced-mode: \"safe-mode\"   # PW_BROWSER_SAFE_MODE=1 — for not-fully-trusted agents\n    untrusted-mode: \"safe-mode+sandbox\"  # add network isolation for untrusted input\n  metadata-is: \"attack-surface-description\"  # NOT a grant, NOT a sandbox boundary\ndisable: false\n---\n\n# 浏览器自动化 — Playwright 版（`pw-browser`）\n\n> `pw-browser` 是基于 playwright-core（Playwright 核心库）的浏览器自动化 CLI，仅依赖 Node.js 和 playwright-core，无需下载任何浏览器。\n>\n> ⚠️ **完整能力（含代码执行）**：本工具**不只是\"点开网页\"**——它包含 `eval`（页面上下文**任意 JavaScript** 执行）与 `run-code`（守护进程上下文执行 Playwright/Node 代码）两项代码执行能力，并通过**持久化本地守护进程**控制浏览器状态（状态跨命令保持）。所有代码执行均受 daemon token 认证保护，可用 `PW_BROWSER_SAFE_MODE=1` 彻底禁用。请先阅读下方「⚠️ 安全边界与使用场景」了解完整攻击面与适用边界，再决定是否授权。\n\n## ⚠️ 安全边界与使用场景\n\n本工具为 **本地 AI 助手对话环境**设计，运行在用户完全可视、可中断的场景中。\n\n| 场景 | 风险 | 说明 |\n|------|------|------|\n| AI 助手对话交互（推荐） | 🟡 低 | 用户全程可见浏览器操作，可随时中断 |\n| 本地开发/测试 | 🟡 低 | 在受控环境中操作测试页面 |\n| 手动触发的数据采集 | 🟡 低 | 用户明确指定的页面和操作 |\n\n以下场景 **不推荐**直接使用，需要额外安全措施：\n\n| 场景 | 风险 | 需要的额外措施 |\n|------|------|-------------|\n| 被不可信 agent 调用 | 🔴 高 | 已内置 daemon token 认证 + 可选 `PW_BROWSER_SAFE_MODE` 禁用代码执行与 cookies/storage 凭证原语（v1.3.2+）；若来源仍不可信，应进一步沙箱隔离 |\n| 作为公开 API 服务 | 🔴 极高 | 必须加认证 + 操作白名单 |\n| CI/CD 自动化流水线 | 🟡 中 | 需限定操作范围，禁止生产环境 |\n\n**核心能力声明：**\n- `eval`：在**浏览器上下文**（`page.evaluate`）执行**任意 JavaScript**——即完整的页面级代码执行能力。它无法访问 Node.js API（`require`/`fs`/`process`），作用域仅限于当前页面；但正因如此，它能读取 `document.cookie`/`localStorage`/`sessionStorage`、发起**带页面凭证的 `fetch` 请求**、操控 DOM 并触发页面内动作（点击、提交等）。**这是与 `run-code` 同级的\"代码执行\"能力，仅场景不同（页面 vs Node）**：同样受 daemon token 认证保护，同样在 `PW_BROWSER_SAFE_MODE=1` 下被禁用；只在用户明确指定、且页面可信时使用，不对高权限/来源不明页面执行\n- `run-code`：在 daemon 进程的**受限沙箱**（`vm` 模块）中执行 Playwright 代码，仅暴露 `page` 和安全 JS 全局；**无法直接调用** `fs`/`child_process`/`process`/`require`。但浏览器上下文可经 `download.saveAs`/`setInputFiles` 在本地磁盘读写文件、能发起任意网络请求（沙箱不阻止）。它仍拥有完整浏览器控制权（导航、读写存储、下载、提交表单、改页面内容），**仅限本地信任环境使用**\n- `cookies` / `storage`：**独立的会话凭证读写原语**，无需任何代码执行即可对 cookie 与 `localStorage` 做 `list` / `export` / `import` / `clear` / `set`。它可直接**提取**当前登录态（含 `HttpOnly` cookie、会话/Bearer 令牌、CSRF token），也可**注入**任意攻击者控制的状态，是凭据盗窃与账户接管的**独立高危面**——**不依赖** `eval` / `run-code`。普通模式下仅由 daemon token 认证保护；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会**整体禁用**全部 `cookies` / `storage` 子命令（与 eval/run-code 同级拦截）。落盘会话文件的处置见下方「会话持久化风险（Rogue Agent）」块：务必用完即删、不跨环境/账号复用。`export`/`import` 默认**限制在 `~/.pw-browser/` 目录内**以防凭证散落或加载外部攻击者构造文件；自 v1.3.8 起该限制**不可由调用方自行解除**——越界需**同时**满足「操作者以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon」+「调用方显式传 `--unsafe`」，否则返回 `UnsafeOverrideNotPermitted`\n- daemon：监听 `127.0.0.1:19223`，仅本机可访问，不暴露到公网；**所有命令（除 `/health` 存活探针）均要求随机 `token` 认证**，token 在 daemon 启动时生成并写入 `~/.pw-browser/daemon.json`（默认仅当前用户可读），CLI 自动携带，外部进程无法在未读取该文件的情况下调用\n- 安全模式：设置环境变量 `PW_BROWSER_SAFE_MODE=1` 启动 daemon 可**彻底禁用** `run-code` / `eval` **以及全部 `cookies` / `storage` 凭证原语**（v1.3.2+），仅保留 snap/click/fill 等白名单命令，适合不需要自定义代码、也不该触碰会话凭证的场景（如接入来源不完全可信的 agent）\n- 非 headless 模式：浏览器窗口始终可见，用户可直接监控所有操作\n- 文件下载：通过 `run-code` 触发页面下载（`download.saveAs`）会写入本地磁盘，注意目标路径\n- 下载文件名消毒（v1.3.10+）：`download` 命令的建议文件名由**被自动化的网页**决定而非操作者。当 `--path` 指向**目录**时，该名字会被强制取 `basename` 并校验解析后仍位于该目录内，越界返回 `PathTraversal`；恶意站点无法再用 `../../.bashrc` 之类的名字把文件写到目录之外。若 `--path` 显式指向一个**文件路径**，则按操作者意图原样使用\n\n## 🔐 权限模型与最低特权（形式化说明）\n\n> **关键澄清：`capabilities` 是攻击面描述，不是权限授予，也不是沙箱边界。**\n> frontmatter 里的 `capabilities` / `allowed-tools` 是给编排者与审核者的**自由文本元数据**：\n> - `capabilities` 列出的是本技能**能够做什么**（即真实攻击面的最小集合），**不代表**它被限制在这些能力内、也**不代表**它已获授权——拥有某项能力意味着技能*可以*行使它，而非*只能*行使它。\n> - `allowed-tools` 仅声明技能可调用哪些宿主工具（如 `Bash(node:*)`），**不构成**对技能行为的强制限制；真正的行为约束来自下方\"实际执行边界\"。\n> 审核时请以\"攻击面下限\"而非\"能力上限\"来读 `capabilities`：技能还可驱动浏览器能抵达的**任意站点**，包括用户**已登录**的站点，从而读取/操作该站点的认证状态。\n\n### 分级部署矩阵（最低特权选择）\n\n| 部署模式 | 代码执行 (`eval`/`run-code`) | 凭证原语 (`cookies`/`storage`) | 网络 | 适用场景 | 残余风险 |\n|----------|------------------------------|-------------------------------|------|----------|----------|\n| **全能力（默认）** | ✅ 开启 | ✅ 开启 | 经浏览器（含带凭证请求） | 可信本地、用户全程可见、临时/演示会话 | 高：可导出/注入会话、可发带凭证请求 |\n| **禁凭证落盘** `PW_BROWSER_CRED_PERSIST=off` | ✅ 开启 | ⚠️ 仅内存态（`export`/`import` 返回 `CredentialPersistenceDisabled`） | 经浏览器 | 需要完整自动化能力、但不允许会话状态跨运行留存的长驻 agent | 中高：仍可读取会话内容，但无法把它变成可复用的凭证文件 |\n| **安全模式** `PW_BROWSER_SAFE_MODE=1` | ⛔ 禁用 (返回 `Disabled`) | ⛔ 禁用 (返回 `Disabled`) | 仅 `snap`/`click`/`fill` 等白名单命令，不主动发请求 | 接入来源不完全可信的 agent、CI 只读巡检 | 中：仍可访问已打开页面的 DOM/可见内容 |\n| **沙箱 + 安全模式** | ⛔ 禁用 | ⛔ 禁用 | 额外做**网络隔离**（禁止出网 / 仅内网白名单） | 完全不可信或第三方输入驱动的自动化 | 低：沙箱逃逸前无法外联或触碰凭证 |\n\n**最低特权原则**：默认按\"全能力\"部署仅当满足——① 用户全程可见可中断；② 操作对象为用户明确指定的页面；③ 不在已登录高权限账户（银行/邮箱/工单系统）上执行未确认动作。否则应**优先启用安全模式**，并对不可信来源**进一步沙箱隔离**。\n\n### 实际执行边界（技能到底受什么约束）\n\n1. **daemon token 认证**：仅限制*谁能通过网络抵达 daemon*（须持有启动时生成的随机 token）。一旦本地进程持有 token，**所有能力全部可用**——token 是\"门禁\"而非\"按能力细分的权限\"。\n2. **安全模式（v1.3.2+）**：本技能**唯一内建的能力开关**，整体禁用代码执行与 `cookies`/`storage` 凭证原语；它是降低攻击面的主开关，但不是沙箱。\n3. **路径限制 + 操作者门禁（v1.3.1 / v1.3.8）**：`cookies`/`storage` 的 `export`/`import` 默认被限制在 `~/.pw-browser/` 内（符号链接经 realpath 解析，无法用软链逃逸）。**解除权归属操作者而非调用方**：越界须同时满足「daemon 以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动」+「调用方显式 `--unsafe`」；仅传 `--unsafe` 会被拒（`UnsafeOverrideNotPermitted`），因为调用方不能修改 daemon 进程的环境变量。这只约束落盘位置，不约束浏览器内的读/写行为。每次凭证路径访问都会写入 daemon stderr 审计行。\n4. **凭证持久化开关（v1.3.9）**：`PW_BROWSER_CRED_PERSIST=off` 启动 daemon 可**单独禁用** `cookies`/`storage` 的 `export`/`import`（返回 `CredentialPersistenceDisabled`），而 `list`/`get`/`set`/`clear` 等内存态操作照常可用。它比安全模式**粒度更细**：只切断\"会话状态落盘\"这一条特权持久化路径，保留正常自动化能力。同样是进程环境变量，调用方无法解除。\n5. **落盘文件权限（v1.3.9）**：`~/.pw-browser/` 以 `0700` 创建，daemon 认证 token（`daemon.json`）与导出的凭证文件均以 `0600` 写入（已存在的旧文件会被 chmod 收紧）。这防止**同机其他本地用户**读走 token 接管浏览器、或直接读取导出的会话。注意：POSIX 权限位在 Windows 上不由操作系统强制执行，Windows 下依赖用户目录的 ACL 继承。\n6. **宿主工具策略**：`allowed-tools` 由宿主平台在*调用层*决定是否放行技能发起的工具调用；它依赖平台实现，**不保证**能限制技能在已获准工具内的具体行为（例如在 `Bash(node:*)` 内仍可执行任意 Node 代码）。\n\n> 结论：本技能**没有**独立于上述六点的\"形式化权限系统\"。任何\"最小权限\"诉求都必须通过**部署模式选择（安全模式/沙箱）+ 操作范围约定**在编排层落实，而非依赖技能元数据自证安全。\n\n## 📝 文档语言与本地化说明\n\n- **文档语言**：本技能文档为**简体中文**。若你或下游 agent 的默认语言非中文，请以代码块中的命令、URL、CSS 选择器与 `snap` 返回的 `ref` 为准——这些是**与语言无关**的自动化锚点。\n- **界面文本匹配是启发式的**：识别分页 / 按钮类型时，文档列出的中文、英文关键词（如\"下一页\"/\"Next\"、\"更新\"/\"保存\"/\"发布\"）只是**识别信号示例，并非穷举**；非中文页面的实际文案会不同。\n- **优先用 DOM 锚点，而非可见文字**：跨语言页面请尽量用 `snap` 得到的 `ref` 或 CSS 选择器（`page.locator('.xxx')`）定位元素，避免依赖本地化后的可见文本，以防因文案不同导致误点 / 误操作。\n- **适用区域**：技能本身不限定网站区域；文档示例与中文 UI 关键词面向中文环境，各\"识别信号\"表已并列给出英文界面关键词。\n- **英文文档**：面向非中文 agent/用户，提供 [`README.en.md`](./README.en.md)（英文 README）与 [`QUICKSTART.en.md`](./QUICKSTART.en.md)（英文端到端示例）。`SKILL.md` 本身保持中文，但其内的命令、URL、CSS 选择器、`snap` 返回的 `ref` 均为语言无关锚点，非中文 agent 可直接据此执行。\n\n## 前置条件（首次使用）\n\n本 Skill 所在目录需已执行 `npm install`（将安装 `playwright-core`）。**无需单独下载浏览器** — daemon 启动时自动检测并使用系统的 Chrome 或 Edge。\n\n> 下文所有命令中的 `{SKILL_DIR}` 请替换为实际的 skill 安装目录路径。\n\n## 架构\n\n`pw-browser` 采用 daemon + client 架构：\n\n```\n┌──────────────┐     HTTP (localhost:19223)     ┌──────────────┐\n│  pw-browser  │ ──────────────────────────────→│   Daemon     │\n│  (CLI 客户端) │                                │  (浏览器进程)  │\n└──────────────┘                                └──────┬───────┘\n                                                       │\n                                                       ├─ Playwright\n                                                       ├─ Chromium 浏览器\n                                                       └─ 页面状态持久化\n```\n\n**daemon 启动后持续运行**，浏览器和页面状态跨命令保持。CLI 每次通过 HTTP 调用 daemon。\n\n## 启动 Daemon\n\n**每次会话开始前**，在后台启动 daemon。以下命令使用 Skill 所在目录的绝对路径和 shim 脚本：\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" daemon &\nsleep 4\n```\n\n> daemon 会在 `127.0.0.1:19223` 监听，首次启动会用 Playwright 的 `channel: 'chrome'` 自动连接系统 Chrome 浏览器（如已安装了 Edge 也会尝试）。无需下载额外的 Chromium。\n\n验证 daemon 可用：\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" init\n```\n\n**关闭 daemon：**\n\n```bash\npw-browser close --all\n```\n\n## 核心工作流\n\n> **注意**：下面所有 `pw-browser` 命令都需要设置 `NODE_PATH`。Agent 执行时应使用完整形式：\n> ```bash\n> SKILL_DIR=\"{SKILL_DIR}\"\n> NODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" <cmd> [args] [--json]\n> ```\n> 为简洁起见，下文示例省略前缀，用 `pw-browser` 表示。\n\n```bash\n# 1. 启动 daemon（会话开始一次）\npw-browser daemon &\n\n# 2. 打开页面\npw-browser open https://www.baidu.com\n\n# 3. 获取页面快照（必须！每次交互前都要 snap）\npw-browser snap\n\n# 4. 交互 — 基于快照中的 e0, e1, e2... ref 引用\npw-browser click e8          # 点击 ref=e8 的元素\npw-browser fill e5 \"hello\"   # 在 ref=e5 的输入框填入文本\npw-browser press Enter       # 键盘按键\n\n# 5. 等待\npw-browser wait-for \"text=加载完成\" --timeout 8000\npw-browser wait-for \"url:https://example.com/*\"\npw-browser wait-for \"state:networkidle\"\n\n# 6. Tab 管理\npw-browser tab list\npw-browser tab select 1\npw-browser tab close 0\n\n# 7. 关闭\npw-browser close             # 关闭当前页面\npw-browser close --all       # 关闭浏览器 + daemon\n```\n\n## 语义规则（必须遵守）\n\n### 规则 1：先观察再操作\n\nCLI **不会**在 open/click 后自动获取快照。**每次交互前，Agent 必须主动执行 `pw-browser snap`**，基于最新快照选择 ref。\n\n```\n正确: pw-browser open URL → pw-browser snap → pw-browser click e5\n错误: pw-browser open URL → pw-browser click e5（缺少 snap）\n```\n\n### 规则 2：点击链接后处理导航\n\n点击可能触发导航的链接（`<a>` 标签、按钮等）后：\n1. `pw-browser snap` — 检查页面是否已变化\n2. 如有新 tab → `pw-browser tab list` → `pw-browser tab select <idx>`\n3. `pw-browser snap` — 获取新页面内容\n\n### 规则 3：页面内容不全\n\n如果快照中元素不全（列表不完整等）：\n- `pw-browser mousewheel 0 500` 滚动\n- 或点击\"加载更多\"/\"下一页\"\n- 重新 `pw-browser snap`\n\n### 规则 4：登录与验证码（人机协作）\n\npw-browser 使用非 headless 模式打开实体 Chrome 窗口，用户可直接看到并操作浏览器。遇到需要人工介入的认证场景时，**不要用 `fill`/`click` 盲目尝试**，应按以下流程交接：\n\n#### 触发条件\n\n从 `snap` 中发现以下任一信号时，启动人工协作流程：\n\n- 页面 title 为「登录」/「Login」/「Sign In」\n- 页面 `url` 包含 `/login`、`/auth`、`/signin`\n- 快照中出现「登录」按钮 + 用户名/密码输入框\n- 快照中出现「验证码」「短信验证」「扫码登录」「滑块验证」等关键字\n- `open` 后自动跳转到登录页（URL 变化）\n\n#### 协作流程\n\n```\n第1步：通报用户\n  告知当前页面需要登录/验证，简明描述页面内容（输入框、验证码类型等）\n\n第2步：询问凭据（可选）\n  如果用户无凭据 → 跳过，直接等用户操作\n  如果用户提供凭据 → 用 fill/click 填入账号密码，点击登录按钮\n\n第3步：等待用户完成验证\n  明确告诉用户\"请在浏览器中完成验证码/二次验证\"\n  用户说\"好了\"\"完成了\"\"继续\"之后才继续\n\n第4步：验证登录状态\n  执行 pw-browser snap\n  检查是否进入目标页面 → 如果还是登录页，询问用户是否还需要操作\n  如果已进入 → 继续自动化流程\n```\n\n#### 示例对话\n\n```\nAgent: 页面跳转到了登录页 (https://xxx.com/login)，页面上有：\n       用户名输入框、密码输入框、登录按钮、滑块验证码。\n       需要我帮你填入账号密码吗？还是你在浏览器里自己操作？\n\nUser:  我来操作\n\nAgent: 好的，Chrome 窗口已打开 — 请完成登录后告诉我。\n\nUser:  好了\n\nAgent: [执行 snap]\n       登录成功！当前是「工作台」页面，左侧菜单有...\n```\n\n#### 重要约束\n\n- **不猜测凭据**：永远不要尝试默认密码或遍历登录\n- **不绕过验证码**：遇到验证码/滑块/短信验证时，立即交给用户\n- **不过度等待**：用户说继续后立即 snap，不额外 sleep\n- **登录失败回环**：snap 后发现仍在登录页 → 告知用户\"看起来还没登录成功，密码错误或验证未通过，请再试试\"\n\n### 规则 5：不要主动新建 tab\n\n点击导致新 tab 时用 `tab list/select/close` 处理。没有 `tab new` 命令。\n\n### 规则 6：翻页前读取策略\n\n涉及翻页、统计、收集、遍历时，参考下面的\"分页策略\"章节。\n\n### 规则 7：不手动读取快照文件\n\n快照通过 `pw-browser snap` 命令获取，不要直接读 `~/.pw-browser/snap.yml`。\n\n### 规则 8：SPA / 富文本编辑器\n\n遇到知识库、文档系统、CMS 等 SPA 页面，参考下面的\"SPA 与富文本编辑器\"章节。\n\n### 规则 9：daemon 故障恢复\n\n如果 CLI 返回连接错误：\n\n```bash\n# 删除旧的 daemon 状态文件\nrm -rf ~/.pw-browser/daemon.json\n```\n\n**杀掉占用端口 19223 的旧进程：**\n\n```bash\n# Windows (PowerShell)\npowershell -Command \"Get-NetTCPConnection -LocalPort 19223 -ErrorAction SilentlyContinue | ForEach-Object { Stop-Process -Id \\$_.OwningProcess -Force }\"\n\n# macOS / Linux\nlsof -ti:19223 | xargs kill -9 2>/dev/null\n# 或: fuser -k 19223/tcp 2>/dev/null\n```\n\n**重新启动：**\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" daemon &\nsleep 4\n```\n\n---\n\n## 命令速查\n\n### 生命周期\n| 命令 | 说明 |\n|------|------|\n| `pw-browser init` | 连接 daemon，确认浏览器可用 |\n| `pw-browser open <url>` | 导航到 URL |\n| `pw-browser close` | 关闭当前页面 |\n| `pw-browser close --all` | 关闭浏览器 + daemon |\n| `pw-browser recover` | 重启浏览器连接 |\n\n### 状态感知\n| 命令 | 说明 |\n|------|------|\n| `pw-browser snap` | 获取页面快照（含 ref 引用）。**ref 跨多次 snap 保持稳定**（同一逻辑元素同 ref），外部 agent 不必每步重新 snap |\n| `pw-browser wait-for <target> [--timeout ms]` | 等待条件满足 |\n\n> **Shadow DOM / iframe 支持**：快照会递归进入 **open shadow root** 与 **同源 iframe**，这些元素同样出现在 ref 表中并可直接 `click`/`fill`/`upload`/`drag`。`snap` 输出的 ref 信息带 `inShadow: true`（shadow 内）或 `frameChain`（iframe 链）标记；定位由 `css >>>` 穿透 + `frameLocator` 自动完成，外部 agent 无需关心。`--annotate` 截图仅标注主文档元素（shadow/iframe 元素无法用 xpath 定位标注，但文本快照中仍可点）。跨域 iframe 不可访问，自动跳过。\n\n> **超大页面快照保护**：对元素极多的页面（数万节点），`snap` 默认最多收集 **3000** 个可交互元素，超出即停止收集并在输出标记 `⚠ snapshot truncated`，JSON 返回 `truncated: true`。这是为防止超大型 DOM 拖慢/撑爆快照的兜底；可用环境变量 `PW_BROWSER_SNAP_LIMIT=<N>` 调大上限（设为 `0` 关闭上限，但仍会遍历整棵树），或先交互缩小页面范围再 `snap`。\n\n### 交互\n| 命令 | 说明 |\n|------|------|\n| `pw-browser click <ref>` | 点击元素 |\n| `pw-browser fill <ref> \"text\"` | 填入文本 |\n| `pw-browser type \"text\"` | 键盘输入 |\n| `pw-browser press <key>` | 按下按键（Enter, Escape, Tab 等） |\n| `pw-browser hover <ref>` | 悬停 |\n| `pw-browser select <ref> <option>` | 选择下拉选项 |\n| `pw-browser check <ref>` | 勾选复选框 |\n| `pw-browser uncheck <ref>` | 取消勾选 |\n| `pw-browser upload <ref> <file1> [file2 ...]` | 文件上传（`<input type=\"file\">`，支持多文件，逗号或空格分隔） |\n| `pw-browser drag <ref源> <ref目标>` | 拖拽（把源元素拖到目标元素，基于 Playwright `dragTo`） |\n| `pw-browser download <ref> [--path dir] [--timeout ms]` | 文件下载（对称于 `upload`）：可选点击 `<ref>` 触发下载，保存到 `--path`（默认当前目录）；也可作为 `act` 动作 `{\"action\":\"download\",\"ref\":\"eN\",\"path\":\"/tmp/x.csv\"}` |\n\n### 页面导航\n| 命令 | 说明 |\n|------|------|\n| `pw-browser goto <url>` | 同 open |\n| `pw-browser go-back` | 后退 |\n| `pw-browser go-forward` | 前进 |\n| `pw-browser reload` | 刷新 |\n\n### 高级\n| 命令 | 说明 |\n|------|------|\n| `pw-browser screenshot [ref] [--path file] [--annotate]` | 截图；`--annotate` 在可交互元素上叠加与 snap ref 对应的编号框，供多模态 agent 直接读编号定位 |\n| `pw-browser mousewheel <dx> <dy>` | 滚动 |\n| `pw-browser eval \"<expr>\" [ref]` | ⚠️ 执行**任意 JavaScript**（页面上下文，完整页面级代码执行：可读 cookie/存储、发起带凭证请求；受 token 认证保护，safe-mode 下禁用）<br>**作用域语义**：不带 `ref` → 表达式在页面全局求值（返回 `scope: \"page\"`）；带 `ref` → 表达式在该元素上求值，标识符 **`el`** 绑定到对应 DOM 节点（返回 `scope: \"element\"`），例：`eval \"el.textContent\" e3`。<br>返回值必须可 JSON 序列化，**不能返回 DOM 节点或循环结构**。 |\n| `pw-browser run-code \"<code>\"` | ⚠️ 执行 Playwright 代码（受限沙箱：无直接 Node fs/process 权限，但可经浏览器下载/上传读写本地文件） |\n| `pw-browser dialog-accept [text]` | 确认对话框 |\n| `pw-browser dialog-dismiss` | 取消对话框 |\n\n### Tab\n| 命令 | 说明 |\n|------|------|\n| `pw-browser tab list` | 列出所有 tab |\n| `pw-browser tab select <idx>` | 切换到指定 tab (0-based) |\n| `pw-browser tab close <idx>` | 关闭指定 tab |\n\n### 延时\n| 命令 | 说明 |\n|------|------|\n| `pw-browser sleep <seconds>` | 等待 N 秒 |\n\n### 批量动作与历史（借鉴 browser-use 的 multi-act / 自纠错）\n| 命令 | 说明 |\n|------|------|\n| `pw-browser act '<json>'` | 批量执行动作序列（JSON 数组），如 `[{\"action\":\"fill\",\"ref\":\"e3\",\"text\":\"hello\"},{\"action\":\"click\",\"ref\":\"e5\"}]`。每步后自动检测 DOM 变化，若页面出现新元素则**中断序列并自动 re-snap** 返回最新快照；失败步附带诊断（元素是否仍存在/相似 ref 建议） |\n| `pw-browser history [--limit N] [--clear]` | 查询 daemon 记录的操作历史（每条命令、参数、耗时、结果），`--clear` 清空 |\n\n`act` 支持的动作：`click` / `fill` / `type` / `press` / `hover` / `select` / `check` / `uncheck` / `upload`（对象带 `files: [\"/path\"]`）/ `drag`（对象带 `target: \"eN\"`）/ `goto` / `screenshot`，动作对象形如 `{\"action\":\"...\",\"ref\":\"eN\",\"text\":\"...\",\"key\":\"...\",\"option\":\"...\",\"url\":\"...\",\"files\":[\"...\"],\"target\":\"eN\"}`。\n\n---\n\n## 省 Token 用法（默认即高效，别退回 browser-use 的反模式）\n\n本 skill 是**确定性执行器 + 持久 daemon**，大模型（外部 AI）只负责规划、不内嵌在 skill 里。因此**没有「每步都调 LLM」的 token 黑洞**——token 只在外部 AI 主动调用时产生，且完全可控。请保持以下用法以持续省 token：\n\n- **用文本 `snap` 规划，而不是每步截图喂视觉模型**。`snap` 返回的是紧凑的 ref 文本表（`e1 button 提交`），几十 token；`screenshot` 一张图是数百 KB 的 base64，贵 1~2 个数量级。\n- **复用稳定 ref，不必每步重新 `snap`**。同一逻辑元素的 ref 跨多次 snap 保持不变（见上「状态感知」表），外部 AI 可直接拿上一步的 ref 点 `click e1` / `fill e2`。\n- **把动作攒成 `act` 一次性发**。登录等一连串操作写成 `[{...},{...}]` 一次调用，daemon 内部自纠错，中间不回模型。理想形态：`1 次 snap → 1 次 act → 完事`。\n- **`screenshot --annotate` 是 opt-in**，仅在真有视觉歧义、需要多模态定位时才用；不要默认每步截图。\n\n> ⚠️ 若外部 AI 被 prompt 成「每步都 `screenshot --annotate` 丢给视觉模型」，就会复刻 browser-use 的烧钱循环。token 成本的责任在编排层，不在 skill。\n\n## Cookie 与本地存储（一等命令）\n\n不再需要靠 `eval` 曲线救国，直接用以下命令读写 cookie 与 `localStorage`：\n\n### `cookies`\n| 子命令 | 说明 |\n|--------|------|\n| `pw-browser cookies list` | 列出当前上下文全部 cookie |\n| `pw-browser cookies export [--path file]` | 导出 cookie 到 JSON 文件（默认 `~/.pw-browser/cookies.json`） |\n| `pw-browser cookies import <file>` | 从 JSON 文件导入 cookie |\n| `pw-browser cookies clear` | 清空全部 cookie |\n| `pw-browser cookies set <name> <value> [--domain d] [--path p]` | 设置一个 cookie；`--domain` 省略时取当前页面域名 |\n\n### `storage`（localStorage）\n| 子命令 | 说明 |\n|--------|------|\n| `pw-browser storage get [key]` | 读取某个 key（省略 key 则返回全部，以对象形式返回） |\n| `pw-browser storage set <key> <value>` | 写入 key/value |\n| `pw-browser storage clear` | 清空 localStorage |\n| `pw-browser storage export [--path file]` | 导出 localStorage 到 JSON 文件 |\n| `pw-browser storage import <file>` | 从 JSON 文件导入（逐 key 写入） |\n\n> ⚠️ `cookies` / `storage` 依赖真实页面源（http/https）。`file://` 与 `data:` 页面不支持 cookie，`localStorage` 行为也不可靠——请先 `open` 一个真实 URL 再操作。\n\n> 🔒 **会话持久化风险（Rogue Agent / 中危）**：`cookies` / `storage` 的 `export` / `import` 让登录态可落盘备份、跨运行恢复——对正常用户是免登录便利，但**在失控或恶意 Agent 场景下，这正是\"特权访问持久化\"的典型手段**：导出的会话文件等同一份可复用的身份凭证，可被用来跳过认证、长期驻留。缓解：① 导出的会话文件**等同密钥**，用完即 `rm`；暂存请留在 `~/.pw-browser/` 内——自 v1.3.9 起该目录以 `0700` 创建、凭证文件与 daemon token 均以 `0600` 写入（**无需再手工 `chmod`**，旧文件也会被自动收紧；Windows 上 POSIX 权限位不由系统强制，依赖用户目录 ACL）；② **不要**在自动化流程里默认把凭证持久化到磁盘——自 v1.3.9 起这不再只是建议：操作者可用 `PW_BROWSER_CRED_PERSIST=off` 启动 daemon，**强制禁用** `export`/`import`（返回 `CredentialPersistenceDisabled`）而保留 `list`/`get`/`set`/`clear` 内存态操作，在不牺牲自动化能力的前提下切断特权持久化路径；③ 不需要时尽快 `cookies clear` / `storage clear` 并 `pw-browser shutdown` 关闭 daemon，利用空闲自动退出（`PW_BROWSER_IDLE_MS`，默认 15min）缩短凭证在内存中的驻留窗口；④ 对来源不可信的调用方，用 `PW_BROWSER_SAFE_MODE=1` 启动 daemon——自 v1.3.2 起它会**整体禁用**全部 `cookies` / `storage` 子命令（凭证原语与代码执行同级拦截），必要时再配合沙箱隔离或限定操作范围；⑤ **代码层兜底（v1.3.1 / v1.3.8）**：`export`/`import` 默认被限制在 `~/.pw-browser/` 目录内（realpath 解析，软链无法逃逸），从路径层面降低凭证被散落或加载外部攻击者构造文件的可能；且**该限制不可由被约束方自行解除**——越界须由操作者以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon 后、调用方再显式 `--unsafe`，缺一不可，全部访问均写 stderr 审计。详见 QUICKSTART 示例 4 安全提醒。\n\n## Daemon 生命周期（为什么任务结束后浏览器还在）\n\ndaemon 是**故意持久化**的：它跨命令持有同一个浏览器实例，避免每次交互都重开 Chrome。因此：\n\n- 你看到「任务结束 Chrome 还在」是正常的——daemon 进程还活着、抱着浏览器。\n- **显式停止**：`pw-browser shutdown` 会关掉浏览器并退出 daemon（已加固：即使 `browser.close()` 卡住也会超时兜底退出，不会再退不出/卡客户端）。\n- **空闲自动退出**：daemon 默认 **15 分钟无命令** 就自动关浏览器并退出（环境变量 `PW_BROWSER_IDLE_MS` 可改，设为 `0` 关闭该特性）。所以走开后不用手动 `shutdown`，它自己会清理，Chrome 不会一直挂着。\n- **监听端口可配 + 冲突避让**：默认 `127.0.0.1:19223`，可用环境变量 `PW_BROWSER_PORT` 覆盖。启动时若目标端口已有 daemon 存活（`/health` 返回 200），当前进程会直接退出（避免双开）；若被其它进程占用（`EADDRINUSE`），则自动递增端口直到可用，并把实际端口写回 `~/.pw-browser/daemon.json`。客户端读取该文件里的端口，无需手动指定。\n- **别直接杀进程**：用任务管理器 / `Stop-Process` 强杀 daemon 可能导致浏览器子进程残留（Windows 上 Playwright 的 job object 通常会回收，但不保证）。优先用 `shutdown` 或等空闲自动退出。\n\n---\n\n## 等待策略\n\n`wait-for` 支持多种目标格式：\n\n```bash\n# 等待 URL 匹配\npw-browser wait-for \"url:**/dashboard\"\n\n# 等待文本出现\npw-browser wait-for \"text=加载完成\"\n\n# 等待页面加载状态（load / domcontentloaded / networkidle）\npw-browser wait-for \"state:networkidle\"\n\n# 等待 CSS 选择器\npw-browser wait-for \".result-list\" --timeout 15000\n```\n\n---\n\n## 运行自定义代码（`run-code`）⚠️ 高级功能\n\n> ⚠️ **安全警告：** `run-code` 在 daemon 进程的 **受限沙箱（`vm` 模块）** 中执行 Playwright 代码。它**无法直接调用** Node.js 系统 API（`fs` / `child_process` / `process` / `require`）；但浏览器上下文本身可经下载（`download.saveAs`）或文件上传（`setInputFiles`）在本地磁盘读写文件、并能发起任意网络请求（沙箱不阻止），因此**仍能持久化数据到本地磁盘**。仅在用户明确指定的任务中使用，不要执行来源不明的代码片段。\n\n当内置命令不够用时，用 `run-code` 执行自定义 Playwright 代码：\n\n```bash\n# 获取页面标题\npw-browser run-code \"return await page.title();\"\n\n# 获取页面 HTML\npw-browser run-code \"return await page.content();\"\n\n# 在页面中执行 JS\npw-browser run-code \"return await page.evaluate(() => document.title);\"\n\n# 等待网络空闲\npw-browser run-code \"await page.waitForLoadState('networkidle');\"\n\n# 复杂场景：提取列表数据\npw-browser run-code \"\n  const items = await page.locator('.product-item').all();\n  const results = [];\n  for (const item of items) {\n    results.push({\n      title: await item.locator('.title').textContent(),\n      price: await item.locator('.price').textContent()\n    });\n  }\n  return JSON.stringify(results);\n\"\n```\n\n**注意：**\n- `run-code` 中直接使用 Playwright Page API\n- 代码在 async 函数中执行，`page` 对象已注入\n- 返回值自动序列化为字符串\n- ⚠️ 此命令可以触发实际的业务操作（提交订单、发送消息、删除数据等），执行前确认用户意图\n\n---\n\n## 分页策略\n\n### 步骤 1：识别分页类型\n\n> ⚠️ 下表关键词为**识别启发式**：中文 / 英文示例（\"下一页\"/\"Next\"、\"加载更多\"/\"Load more\"）并不穷举，非中文页面的文案会不同。实际定位请优先用 `snap` 的 `ref` 或 CSS 选择器，勿仅依赖可见文本。\n\n从 snap 判断：\n\n| 类型 | 识别信号 | 翻页方式 |\n|------|---------|---------|\n| **页码分页** | 底部有 1/2/3...页码、\"下一页\"/\"Next\"/\">\" | 点击页码或\"下一页\" |\n| **无限滚动** | 底部无分页控件，内容随滚动增加 | `mousewheel` 滚动 |\n| **加载更多** | 底部有\"加载更多\"/\"Load more\"/\"查看更多\" | 点击该按钮 |\n\n### 步骤 2：执行翻页\n\n**页码分页：**\n```bash\npw-browser snap                          # 找到\"下一页\"按钮的 ref\npw-browser click e42                     # 点击\npw-browser sleep 2 && pw-browser snap    # 验证\n```\n\n**无限滚动：**\n```bash\npw-browser mousewheel 0 800\npw-browser sleep 2 && pw-browser snap\n```\n\n**加载更多按钮：**\n```bash\npw-browser click <ref>\npw-browser sleep 2 && pw-browser snap\n```\n\n### 步骤 3：判断翻页成功\n\n| 方式 | 成功信号 | 失败/结束信号 |\n|------|---------|-------------|\n| 页码 | 内容更新，URL 变化 | \"下一页\"按钮 disabled 或消失 |\n| 滚动 | 内容增加，新元素出现 | 内容不变，\"没有更多了\" |\n| 按钮 | 新内容加载，按钮仍可点击 | \"已加载全部\"，按钮消失 |\n\n---\n\n## SPA 与富文本编辑器 ⚠️ 破坏性操作\n\n> ⚠️ **警告：** 以下操作会**真实修改**网页内容（知识库文档、CMS 页面等）。执行前确认当前处于编辑/草稿状态、修改内容已经用户确认。保存/发布操作不可逆。\n\n处理知识库、文档系统、CMS 等 SPA 页面的编辑操作：\n\n### 识别信号\n\n- 点击\"编辑\"后 URL 不变但按钮变化\n- snap 中出现 `contenteditable`、编辑器 toolbar\n- 不是普通 `input/textarea`，而是复杂编辑器\n\n### 编辑流程\n\n1. **进入编辑态：** `pw-browser click <编辑按钮的ref>`\n2. **验证进入：** `pw-browser snap` — 检查是否出现\"更新\"/\"保存\"按钮\n3. **写入内容（RTE）：**\n```bash\npw-browser run-code \"\n  const editor = page.locator('[contenteditable=\\\"true\\\"]').first();\n  await editor.click();\n  await page.keyboard.press('Control+A');\n  await page.keyboard.type('要写入的内容');\n  await page.waitForTimeout(500);\n\"\n```\n4. **保存：**\n```bash\npw-browser run-code \"\n  await page.evaluate(() => {\n    const btn = Array.from(document.querySelectorAll('button'))\n      .find(b => ['更新','保存','发布'].includes(b.textContent.trim()));\n    btn?.click();\n  });\n  await page.waitForTimeout(3000);\n\"\n```\n5. **验证：** `pw-browser snap` — 确认保存成功、内容正确\n\n> **不要**直接用 `innerText`/`textContent` 修改 RTE 内容。Playwright 的 `keyboard.type` 和 `fill` 是正确方式。\n\n---\n\n## 结构化输出\n\n所有命令在 daemon 端返回 JSON：\n\n```json\n{\"ok\": true, \"data\": {...}, \"elapsedMs\": 123}\n{\"ok\": false, \"error\": {\"kind\": \"ElementNotFound\", \"message\": \"...\"}, \"elapsedMs\": 50}\n```\n\nCLI 客户端默认以人类可读格式输出；加 `--json` 标志输出原始 JSON。\n\n## 错误处理\n\n| 错误类型 | 原因 | 处理 |\n|---------|------|------|\n| `ElementNotFound` | snap 后 ref 已失效 | 重新 snap 获取新 ref |\n| `NavigationTimeout` | 页面加载超时 | 先 snap 检查实际状态 |\n| 连接拒绝 | daemon 未运行 | 重新启动 daemon |\n| 空快照 | 页面未加载完成 | wait-for state:load 后重新 snap |\n\n---\n\n## 完整示例：百度搜索\n\n```bash\nSKILL_DIR=\"{SKILL_DIR}\"\nPW=\"NODE_PATH=${SKILL_DIR}/node_modules node ${SKILL_DIR}/pw-browser.js\"\n\n# 启动 daemon（首次）\n$PW daemon &\nsleep 4\n\n# 打开百度\n$PW open https://www.baidu.com\n\n# 快照 → 找到搜索框和按钮的 ref\n$PW snap\n# 例如：e12 = textarea（搜索框），e13 = button（百度一下）\n\n# 填搜索关键词\n$PW fill e12 \"天气预报\"\n\n# 点击搜索\n$PW click e13\nsleep 2\n\n# 检查搜索结果\n$PW snap | head -30\n\n# 清理\n$PW close --all\n```\n\n---\n\n## 深入参考\n\n| 场景 | 文件 |\n|------|------|\n| 端到端示例（表单/上传下载/shadow-iframe/cookie/act） | `QUICKSTART.md`（中文）/ `QUICKSTART.en.md`（英文） |\n| 翻页策略详解 | `references/pagination.md` |\n| 富文本编辑器策略 | `references/rich-text-editor.md` |\n| 运行自定义代码 | `references/running-code.md` |\n\nFile v1.0.15:README.md\n\n# pw-browser\n\n> 基于 Playwright 的浏览器自动化 CLI — daemon + client 架构，直接使用系统 Chrome/Edge，无需下载额外浏览器。\n\n> 📘 简体中文文档（本文件）。English documentation: [`README.en.md`](./README.en.md)。端到端示例见 [`QUICKSTART.md`](./QUICKSTART.md)（含 [`QUICKSTART.en.md`](./QUICKSTART.en.md)）。\n\n> ⚠️ **完整能力（含代码执行与凭证操控）**：本工具**不只是\"点开网页\"**——它内置 `eval`（页面上下文**任意 JavaScript** 执行）与 `run-code`（守护进程上下文执行 Playwright/Node 代码）两项**代码执行**能力，以及 `cookies` / `storage` **凭证读写原语**（可**无需代码执行**即提取/注入登录态与会话令牌）。全部经持久化本地守护进程（127.0.0.1:19223）控制，受 daemon token 认证保护；`PW_BROWSER_SAFE_MODE=1` 可**彻底禁用**代码执行与凭证读写（v1.3.2+）。请先阅读下方「🔐 权限模型与最低特权」与「⚠️ 安全边界与使用场景」了解完整攻击面与适用边界，再决定是否授权。\n\n## 特性\n\n- **系统浏览器复用**：通过 Playwright `channel: 'chrome'` 连接系统已安装的 Chrome 或 Edge，不下载 Chromium\n- **跨平台**：支持 Windows、macOS、Linux\n- **持久化会话**：daemon 保持浏览器状态跨命令存活，无需每次重启\n- **可访问性快照**：`snap` 命令生成页面元素树，带 ref 引用，无需写 CSS 选择器\n- **非 headless**：浏览器窗口始终可见，用户可实时监控所有操作\n- **人机协作**：遇到登录/验证码时自动交接给用户操作\n\n## 安装\n\n```bash\ngit clone <repo-url> pw-browser\ncd pw-browser\nnpm install\n```\n\n> 前提：系统已安装 [Node.js](https://nodejs.org/) 18+ 和 Chrome 或 Edge 浏览器。\n\n## 快速开始\n\n```bash\n# 1. 启动 daemon（后台运行）\nnode pw-browser.js daemon &\n\n# 2. 等待 daemon 就绪\nsleep 4\n\n# 3. 打开页面\nnode pw-browser.js open https://www.baidu.com\n\n# 4. 获取页面快照（必须！每次交互前都要 snap）\nnode pw-browser.js snap\n\n# 5. 交互 — 基于快照中的 e0, e1, e2... ref\nnode pw-browser.js fill e12 \"天气预报\"\nnode pw-browser.js click e13\n\n# 6. 查看结果\nsleep 2\nnode pw-browser.js snap\n\n# 7. 关闭\nnode pw-browser.js close --all\n```\n\n## 命令一览\n\n| 类别 | 命令 | 说明 |\n|------|------|------|\n| 生命周期 | `init` / `open <url>` / `close` / `close --all` / `recover` | 启动、导航、关闭、恢复 |\n| 状态感知 | `snap` / `wait-for <target>` | 快照（**ref 跨多次 snap 保持稳定**，同元素同 ref） |\n| 交互 | `click <ref>` / `fill <ref> \"text\"` / `type \"text\"` / `press <key>` / `hover <ref>` / `select <ref> <opt>` / `check <ref>` / `uncheck <ref>` / `upload <ref> <file...>` / `drag <ref源> <ref目标>` / `download <ref> [--path dir]` | 点击、填写、按键、上传、拖拽、下载等 |\n| Cookie/存储 | `cookies list` / `export [--path f]` / `import <file>` / `clear` / `set <name> <value> [--domain d]` · `storage get [key]` / `set <k> <v>` / `clear` / `export [--path f]` / `import <file>` | 读写 cookie 与 localStorage（需先 `open` 真实 http/https 页面） |\n| 导航 | `goto <url>` / `go-back` / `go-forward` / `reload` | 页面导航控制 |\n| Tab | `tab list` / `tab select <idx>` / `tab close <idx>` | 多标签页管理 |\n| 批量 | `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` | 一次性执行动作序列，中途检测 DOM 变化自动中断重规划 |\n| 历史 | `history [--limit N] [--clear]` | 查看/清空操作历史 |\n| 高级 | `screenshot [--annotate]` / `mousewheel <dx> <dy>` / `eval \"<expr>\" [ref]` ⚠️ / `run-code \"<code>\"` ⚠️ | 截图（`--annotate` 叠加与 ref 对应编号框）、滚动、代码执行。`eval` 不带 `ref` 时在页面全局求值（`scope: \"page\"`）；带 `ref` 时标识符 **`el`** 绑定到该 DOM 节点（`scope: \"element\"`），如 `eval \"el.textContent\" e3`。返回值须可 JSON 序列化 |\n| 延时 | `sleep <seconds>` | 等待 |\n| 对话框 | `dialog-accept [text]` / `dialog-dismiss` | 处理原生 alert/confirm/prompt 对话框 |\n| 守护进程 | `shutdown` | 关闭持久化 daemon，释放浏览器进程 |\n\n> ⚠️ `eval` 在**浏览器上下文**执行 JS（无 Node 权限）；`run-code` 在 daemon 的**受限沙箱（`vm`）**中执行 Playwright 代码（无直接 Node `fs`/`process`/`child_process` 权限，但可经浏览器下载/上传读写本地文件、发起任意网络请求）。二者均拥有完整浏览器控制权。仅在用户明确指定的任务中使用。\n\n## 架构\n\n```\n┌──────────────┐     HTTP (localhost:19223)     ┌──────────────┐\n│  pw-browser  │ ──────────────────────────────→│   Daemon     │\n│  (CLI 客户端) │                                │  (浏览器进程)  │\n└──────────────┘                                └──────┬───────┘\n                                                       │\n                                                       ├─ Playwright\n                                                       ├─ Chrome / Edge\n                                                       └─ 页面状态持久化\n```\n\ndaemon 启动后持续运行，浏览器和页面状态跨命令保持。CLI 每次通过 HTTP 调用 daemon。\n\n## 工作流核心规则\n\n1. **先观察再操作**：每次交互前必须 `snap`，基于快照 ref 操作\n2. **登录交接**：遇到登录/验证码页面时，告知用户手动操作，用户确认后继续\n3. **翻页前先识别类型**：页码分页 / 无限滚动 / 加载更多，对应不同翻页方式\n4. **daemon 故障恢复**：连接错误时杀端口 19223 进程 → 删 daemon.json → 重启\n\n详细规则和示例见 `SKILL.md`。\n\n## 增强能力（借鉴 browser-use）\n\n本工具的设计定位是「**安全可控的浏览器 CLI 基座 + 持久 daemon，规划交给外部 AI**」，而非 browser-use 那种「大脑+手一体」的 Agent 框架。我们从 browser-use 借鉴了四项对 CLI 基座同样有价值的能力：\n\n### A. 稳定的元素引用（stable ref）\n- `snap` 不再每次重排 ref，而是按元素的**语义身份**（有文本/placeholder/aria-label 用 `tag|text|placeholder|aria`；否则回退到 DOM 分支路径哈希）建立 `stableKey → ref` 持久映射。\n- **同一逻辑元素跨多次 snap 保持同一 ref**，外部 AI 不必每步重新 `snap`，可直接复用之前的 ref 继续交互（典型场景：填完表单再点提交，e13 还是那个提交按钮）。\n- `findElement` 新增 **Strategy 0：xpath 精确定位优先**——用快照里记录的精确 xpath 钉住元素，彻底解决纯语义查找「同名元素误点」的歧义。\n\n### B. 视觉辅助截图（`screenshot --annotate`）\n- `screenshot --annotate` 会在页面上注入覆盖层，为**每个与 snap ref 对应的元素**叠加红色边框 + 编号 label（用 `document.evaluate(xpath)` 精确定位）。\n- 截图后自动移除覆盖层。多模态 AI 可「看图读编号」直接定位 `e13` 在页面哪个位置，弥补纯文本快照缺少空间信息的短板。\n\n### C. 动作序列 + 自纠错（`act`）\n- `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` 一次性下发动作序列，复用 `executeSingle` 逐个执行。\n- 每个非首动作执行前重新 `snap` 并比对 `branchPathHash` 集合：若检测到页面出现**新元素**（DOM 变化），立即 `interrupted=true` 并返回最新快照，交由外部 AI 重新规划——这对应 browser-use 的 `multi_act` 中途中断重规划。\n- 失败的动作会附带 `diagnosis`（ref 是否仍在、相似 ref 建议），便于自愈。\n\n### D. 结构化输入 + 操作历史（`history`）\n- `act` 接受结构化 JSON 动作数组（而非零散子命令），降低外部 AI 拼 CLI 参数的出错率。\n- `history` 记录每次操作的 `{ ts, cmd, params（已过滤 token）, ok, elapsedMs }`，上限 500 条，可 `--clear`。外部 AI 可回溯「刚才点了什么、哪步失败」，实现上下文压缩与复盘——对应 browser-use 用 Mem0 压缩历史上下文的思路（此处用轻量本地历史替代向量库）。\n\n### E. Shadow DOM / iframe 穿透\n- 快照递归进入 **open shadow root** 与 **同源 iframe**，这些元素同样出现在 ref 表中并可直接 `click`/`fill`/`upload`/`drag`。\n- ref 信息带 `inShadow: true`（shadow 内）或 `frameChain`（iframe 链）标记；定位由 `css >>>` 穿透 + `frameLocator` 自动完成，外部 AI 无感知。\n- 跨域 iframe 不可访问，自动跳过；`--annotate` 截图仅标注主文档元素（shadow/iframe 元素无法用 xpath 定位标注，但文本快照中仍可点）。\n\n### F. 文件上传与拖拽\n- `upload <ref> <file1> [file2 ...]`：对 `<input type=\"file\">` 设置文件（支持多文件）。\n- `drag <ref源> <ref目标>`：基于 Playwright `dragTo` 实现拖拽。\n- `download <ref> [--path dir]`：对称于 `upload`——可选点击 `<ref>` 触发下载，保存到 `--path`（默认当前目录）；亦可作为 `act` 动作。\n  > 🔒 v1.3.10+：网页给出的建议文件名会被强制取 `basename` 并限制在 `--path` 目录内（越界返回 `PathTraversal`），恶意页面无法借 `../../` 之类的文件名写到目录之外。`--path` 若指向具体文件则按原样使用。\n\n### G. Cookie 与本地存储（一等命令）\n- `cookies list|export|import|clear|set`：读写当前上下文 cookie，不再需要靠 `eval` 曲线救国。\n- `storage get|set|clear|export|import`：基于 `localStorage` 的读写（依赖真实 http/https 页面源）。\n- 典型用法：登录后 `cookies export` 备份会话，下次 `cookies import` 直接恢复，免去重复登录。\n- 📁 **路径限制（操作者门禁）**：`export`/`import` 默认限制在 `~/.pw-browser/` 目录内（防凭证散落 `/tmp` 或加载外部攻击者构造文件；路径经 realpath 解析，软链无法逃逸）。自 v1.3.8 起**该限制不可由调用方自行解除**——越界需**同时**满足：① 操作者以 `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1` 启动 daemon；② 调用方显式传 `--unsafe`。仅传 `--unsafe` 会返回 `UnsafeOverrideNotPermitted`。命令返回带 `confined` 与 `warning` 字段，daemon stderr 记录每次凭证路径访问。\n- 🔒 **会话持久化风险（Rogue Agent / 中危）**：`export`/`import` 让登录态可落盘、跨运行恢复——便利，但失控/恶意 Agent 可用它做\"特权访问持久化\"。导出的会话文件等同密钥：用完即删；暂存留在 `~/.pw-browser/` 内即可——自 v1.3.9 起该目录 `0700`、凭证文件与 daemon token 均 `0600` 写入，无需手工 `chmod`（Windows 上依赖用户目录 ACL）。不想让凭证落盘的场景，用 `PW_BROWSER_CRED_PERSIST=off` 启动 daemon：**强制禁用** `export`/`import`（`CredentialPersistenceDisabled`），保留 `list`/`get`/`set`/`clear` 内存态操作——比安全模式粒度更细，适合需要完整自动化但禁止会话跨运行留存的长驻 agent。不用时 `cookies clear`/`storage clear` 并 `shutdown` 关 daemon（空闲默认 15min 自动退出）。对不完全可信的调用方，用 `PW_BROWSER_SAFE_MODE=1` 启动 daemon——自 v1.3.2 起它会**整体禁用**全部 cookies/storage 子命令（凭证原语与代码执行同级拦截），必要时再配合沙箱隔离。\n\n## 省 Token 与 Daemon 生命周期\n\n**没有「每步调 LLM」的 token 黑洞**：本 skill 只是确定性执行器 + 持久 daemon，大模型只在外部编排层、且仅在主动调用时产生 token。保持以下用法即持续省 token：\n\n- 用文本 `snap`（ref 表，几十 token）规划，而非每步 `screenshot`（数百 KB base64）喂视觉模型；\n- 复用跨 snap 稳定的 ref，不必每步重新 `snap`；\n- 把一连串操作攒成一次 `act '<json>'`，daemon 内部自纠错，中间不回模型。\n\n**daemon 是故意持久化的**（跨命令复用同一浏览器，避免每次重开 Chrome）。因此：\n\n- 任务结束后浏览器窗口还在是正常的——daemon 进程仍持有它；\n- 显式停止：`pw-browser shutdown`（已加固，`browser.close()` 卡住也会超时兜底退出，不会退不出/卡客户端）；\n- **空闲自动退出**：默认 **15 分钟无命令** 即自动关浏览器并退出，环境变量 `PW_BROWSER_IDLE_MS` 可调整（设为 `0` 关闭）。走开后无需手动清理；\n- **监听端口可配 + 冲突避让**：默认 `127.0.0.1:19223`，可用 `PW_BROWSER_PORT` 覆盖。启动时若目标端口已有 daemon 存活则直接退出（避免双开）；若被其它进程占用则自动递增端口，并把实际端口写回 `~/.pw-browser/daemon.json`，客户端自动读取。\n- 优先用 `shutdown` 或等空闲退出，避免直接强杀进程导致浏览器子进程残留。\n\n## 安全说明\n\n- daemon 监听 `127.0.0.1:19223`，仅本机可访问\n- **命令认证**：除 `/health` 存活探针外，所有 daemon 命令都要求随机 `token`。token 在 daemon 启动时生成，写入 `~/.pw-browser/daemon.json`（默认仅当前用户可读），CLI 自动携带。未读取该文件的本地进程无法调用——这关闭了\"无认证 HTTP 端点 = RCE 界面\"的缺口\n- **安全模式**：`PW_BROWSER_SAFE_MODE=1 node pw-browser.js daemon` 可彻底禁用 `run-code` / `eval` **以及全部 `cookies` / `storage` 凭证原语**（v1.3.2+），仅保留 snap/click/fill 等白名单命令——适合接入来源不完全可信的 agent\n- `eval` 运行在浏览器上下文；`run-code` 运行在 `vm` 受限沙箱（无直接 Node 系统 API，但可经浏览器下载/上传读写本地文件），仅限本地信任环境使用\n- 浏览器以非 headless 模式运行，用户可实时监控\n- 不推荐作为公开 API 服务暴露，如需请加认证和操作白名单\n- **权限模型说明**：frontmatter 的 `capabilities` / `allowed-tools` 是**自由文本元数据**，描述本技能的攻击面，**不是权限授予或沙箱边界**——拥有某项能力意味着技能*可以*做它，而非*只能*做它。本技能**没有**独立于 daemon token / 安全模式 / 路径限制之外的形式化权限系统；任何\"最小权限\"诉求须通过**部署模式选择（安全模式 + 沙箱）+ 操作范围约定**在编排层落实。详见 SKILL.md「🔐 权限模型与最低特权」。\n\n### 依赖安全说明（CVE-2025-59288）\n\n本项目直接依赖 `playwright-core@1.61.1`（Playwright 核心库，无浏览器下载逻辑，≥ 1.55.1）。CVE-2025-59288 影响的是完整 `playwright` 包**下载并安装浏览器**的安装脚本（`curl -k` 未校验证书）；本工具仅依赖 `playwright-core`，**根本不包含浏览器下载代码**，且运行时通过 `channel: 'chrome'` 复用系统已安装的 Chrome/Edge，因此该漏洞在本工具的使用路径上完全不可触发。\n\n## 作为 AI Skill 使用\n\n本工具可作为 AI 助手的 skill 使用。将整个目录放入 skill 安装路径，AI 助手通过 `SKILL.md` 中的规则指导自动化操作。\n\n## License\n\n[MIT](LICENSE)\n\nFile v1.0.15:_meta.json\n\n{\n  \"ownerId\": \"kn74y8h6hjpvfa40rgznyt8e3184yaqe\",\n  \"slug\": \"playwright-browser-use\",\n  \"version\": \"1.0.15\",\n  \"publishedAt\": 1785420897405\n}\n\nFile v1.0.15:references/pagination.md\n\n# 分页策略\n\n翻页前必须先**识别页面分页类型**，选对翻页方式。\n\n> 📝 **文档语言与本地化**：本文件为**简体中文**。识别信号表中的中文 / 英文关键词（\"下一页\"/\"Next\"、\"加载更多\"/\"Load more\"）仅为**启发式示例，并非穷举**——非中文页面的实际文案会不同。跨语言页面请优先用 `snap` 返回的 `ref` 或 CSS 选择器（`page.locator('.xxx')`）定位，避免依赖可见文本。完整语言说明见 `SKILL.md`「📝 文档语言与本地化说明」。\n\n## 步骤 1：识别分页类型\n\n从 `pw-browser snap` 的输出判断：\n\n| 类型 | 识别信号 | 翻页方式 |\n|------|---------|---------|\n| **页码分页** | 底部有页码（1,2,3...）、\"下一页\"/\"Next\"/\">\" | `click` 页码或\"下一页\" |\n| **无限滚动** | 底部无分页控件，内容随滚动增加 | `mousewheel` |\n| **加载更多** | 底部有\"加载更多\"/\"Load more\"/\"查看更多\" | `click` 该按钮 |\n\n### 常见网站参考\n\n| 网站 | 分页类型 | 翻页方式 |\n|------|---------|---------|\n| 百度搜索 | 页码分页 | 点击页码 |\n| 淘宝搜索 | 页码分页 | 点击页码 |\n| 京东搜索 | 页码分页 | 点击页码 |\n| 知乎 | 页码分页 | 点击页码 |\n| 小红书 | 无限滚动 | 滚动加载 |\n| 抖音 | 无限滚动 | 滚动加载 |\n\n## 步骤 2：执行翻页\n\n### A. 页码分页\n\n```bash\n# 从 snap 找到\"下一页\"按钮 ref\npw-browser snap\npw-browser click e42          # 点击\"下一页\"\npw-browser sleep 2\npw-browser snap               # 验证\n```\n\n备选方式 — 直接点页码：\n\n```bash\npw-browser snap\n# 找到页码数字（如 \"2\"）对应的 ref\npw-browser click e50\npw-browser sleep 2 && pw-browser snap\n```\n\n### B. 无限滚动\n\n```bash\npw-browser mousewheel 0 800\npw-browser sleep 2\npw-browser snap\n```\n\n连续多次滚动直到内容不再增加。\n\n### C. 加载更多按钮\n\n```bash\npw-browser snap\n# 找到按钮 ref\npw-browser click e30\npw-browser sleep 2\npw-browser snap\n```\n\n## 步骤 3：判断翻页成功\n\n| 分页方式 | 成功信号 | 结束信号 |\n|---------|---------|---------|\n| 页码 | snap 内容变化，URL 可能变化 | \"下一页\"按钮消失或 disabled |\n| 滚动 | snap 出现新元素 | 内容不变，出现\"没有更多了\" |\n| 按钮 | 新内容加载 | 出现\"已加载全部\"，按钮消失 |\n\n## 批量翻页提取\n\n```bash\n# 使用 run-code 批量翻页\npw-browser run-code \"\n  const allResults = [];\n  let hasNext = true;\n  while (hasNext) {\n    const items = await page.locator('.item').all();\n    for (const item of items) {\n      allResults.push(await item.textContent());\n    }\n    const nextBtn = page.locator('text=下一页');\n    if (await nextBtn.count() === 0 || await nextBtn.isDisabled()) {\n      hasNext = false;\n    } else {\n      await nextBtn.click();\n      await page.waitForTimeout(2000);\n    }\n  }\n  return JSON.stringify(allResults);\n\"\n```\n\n## 常见问题\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| 滚动后不加载新内容 | 实际上是页码分页 | 检查 snap 底部是否有页码，改为 click |\n| 点击页码没反应 | 按钮 disabled 或需要等待 | `sleep 1` 后再点击 |\n| 翻页后内容相同 | AJAX 加载，需要等待 | 延长 sleep 时间或用 `wait-for` |\n| 页码按钮被遮挡 | 需先滚动到底部 | `mousewheel 0 1000` 再 snap |\n\nFile v1.0.15:references/rich-text-editor.md\n\n# SPA 与富文本编辑器 ⚠️\n\n> ⚠️ **破坏性操作警告：** 以下操作会真实修改网页内容（知识库文档、CMS 页面等）。执行前确认：① 处于编辑/草稿状态而非已发布内容；② 修改内容已经用户确认；③ 保存/发布操作不可逆。\n\n处理知识库、文档系统、CMS（如 Notion/语雀/飞书类页面）中的 SPA 编辑态和富文本编辑器写入。\n\n## 何时使用\n\n满足**任一条件**时，遵循本指南：\n\n- URL/页面属于知识库、文档、笔记、CMS 类站点\n- 任务要求创建/编辑/保存文档正文\n- 点击\"编辑\"按钮后 URL 不变但页面状态变化\n- snap 中出现 `contenteditable`、编辑器 toolbar、\"插入\"/\"正文\"等\n- 表单不是普通 input/textarea，而是复杂编辑器\n\n## 核心原则\n\n- **不要**直接用 `innerText`/`textContent` 写 RTE——不会被编辑器状态机接受\n- 先确认进入编辑态，再用 Playwright 键盘输入\n- 保存后验证内容而不是只看按钮状态\n\n## 流程\n\n### 1. 进入编辑态\n\n```bash\n# 先 snap 找到编辑按钮\npw-browser snap\n\n# 点击编辑按钮\npw-browser click <编辑ref>\n\n# 验证进入编辑态\npw-browser run-code \"\n  return await page.evaluate(() => ({\n    hasUpdate: Array.from(document.querySelectorAll('button'))\n      .some(b => b.textContent.trim() === '更新'),\n    editableCount: document.querySelectorAll('[contenteditable]').length\n  }));\n\"\n```\n\n如果 `hasUpdate=true` 或 `editableCount > 0`，继续；否则尝试重试点击。\n\n### 2. 写入内容（RTE）\n\n```bash\npw-browser run-code \"\n  const editor = page.locator('[contenteditable=\\\"true\\\"], [contenteditable=\\\"plaintext-only\\\"]').first();\n  await editor.click();\n  await page.keyboard.press('Control+A');\n  await page.keyboard.type('要写入的文本内容');\n  await page.waitForTimeout(1000);\n  const text = await editor.textContent();\n  return text;\n\"\n```\n\n### 3. 保存 ⚠️\n\n> ⚠️ 保存/发布操作不可逆，确认内容无误后再执行。\n\n```bash\npw-browser run-code \"\n  await page.evaluate(() => {\n    const btn = Array.from(document.querySelectorAll('button'))\n      .find(b => ['更新','保存','发布','完成'].includes(b.textContent.trim()));\n    btn?.click();\n  });\n  await page.waitForTimeout(3000);\n\"\n```\n\n### 4. 验证\n\n```bash\npw-browser run-code \"\n  const title = document.querySelector('h1, [class*=title]')?.textContent || document.title;\n  const main = document.querySelector('main') || document.body;\n  const blocks = Array.from(main.querySelectorAll('p, h1, h2, h3, li'))\n    .map(b => b.textContent.trim().slice(0, 80))\n    .filter(Boolean);\n  return JSON.stringify({ title, sampleBlocks: blocks.slice(0, 10) });\n\"\n```\n\n成功标准：\n- `hasUpdate=false`（编辑态已退出）\n- 正文包含目标文本\n- 标题未被误改或清空\n- 没有重复写入的文本\n\n## 弹窗处理\n\n- DOM 浮层（弹窗、抽屉、popover）：通过 snap 识别并 click 关闭按钮\n- 原生 JS dialog（alert/confirm/prompt）：用 `pw-browser dialog-accept` / `dialog-dismiss`\n\n## 失败处理\n\n- 点击\"编辑\"超时后，先检查是否已进入编辑态，不要重复点击\n- 后续命令超时，执行 `pw-browser recover` 恢复 daemon\n- 恢复后如果已在编辑态，继续输入和保存\n\nFile v1.0.15:references/running-code.md\n\n# 运行自定义代码（`run-code`）⚠️\n\n> ⚠️ **安全警告：** `run-code` 在 daemon 进程的 **受限沙箱（`vm` 模块）** 中执行代码，**无法直接调用** Node.js 系统 API（`fs` / `child_process` / `process` / `require`）。但**浏览器上下文本身可经下载（`download.saveAs`）或文件上传（`setInputFiles`）在本地磁盘读写文件，并能发起任意网络请求**——这些由注入的 `page` 句柄提供，沙箱不阻止。因此它**仍能把数据持久化到本地磁盘**（见下方\"文件下载\"小节）。仅在用户明确指定任务中使用，不要执行来源不明的代码片段。\n\n当内置命令不够用时，用 `pw-browser run-code` 执行 Playwright 代码。\n\n## `eval`：页面上下文任意 JavaScript 执行 ⚠️\n\n> ⚠️ **安全警告：** `eval` 在**当前页面的 JavaScript 上下文**中执行你提供的任意代码（等价于在浏览器开发者工具控制台里直接输入并执行）。它能读取 `document.cookie`、`localStorage`/`sessionStorage`、发起**携带当前页面凭证**的 `fetch`/`XMLHttpRequest`，并直接操控 DOM、触发点击与表单提交。**它不等于\"执行一个无害的 JS 表达式\"——它是完整的页面级代码执行（page-context RCE）。**\n>\n> `eval` 与 `run-code` 是同一类\"代码执行\"能力，只是作用域不同：\n> - **`eval`** → 页面上下文，能触及页面里的所有数据与会话凭证\n> - **`run-code`** → Node 沙箱上下文，能驱动浏览器但拿不到宿主机 `fs`/`process`\n>\n> 两者都受 daemon **token 认证**保护（未持 token 的外部进程无法调用），且都在 `PW_BROWSER_SAFE_MODE=1` 启动时**被禁用**。仅在用户明确指定的任务、且目标页面可信时使用；不要对来源不明或高权限页面执行。\n\n```bash\n# 读取当前页面所有 cookie（含会话令牌）\npw-browser eval \"document.cookie\"\n\n# 读取 localStorage\npw-browser eval \"JSON.stringify(localStorage)\"\n\n# 在指定元素上下文执行（ref 来自 snap）\npw-browser eval \"el.innerText\" e5\n\n# 发起带页面凭证的请求（可被滥用于 CSRF / 数据外泄，慎用）\npw-browser eval \"await (await fetch('/api/me')).text()\"\n```\n\n## 语法\n\n```bash\npw-browser run-code \"<code>\"\n```\n\n代码在 daemon 的 `vm` 沙箱中执行，`page` 对象已注入（标准的 Playwright Page）。沙箱仅暴露 `page` 和安全 JS 全局，不提供 Node 系统模块。\n\n## 等待策略\n\n```bash\n# 等待网络空闲\npw-browser run-code \"await page.waitForLoadState('networkidle');\"\n\n# 等待元素出现\npw-browser run-code \"await page.locator('.loading').waitFor({ state: 'hidden' });\"\n\n# 等待自定义条件\npw-browser run-code \"await page.waitForFunction(() => window.appReady === true);\"\n\n# 带超时的等待\npw-browser run-code \"await page.locator('.result').waitFor({ timeout: 10000 });\"\n```\n\n## 页面信息\n\n```bash\n# 获取标题\npw-browser run-code \"return await page.title();\"\n\n# 获取 URL\npw-browser run-code \"return page.url();\"\n\n# 获取整个 HTML\npw-browser run-code \"return await page.content();\"\n\n# 视口大小\npw-browser run-code \"return JSON.stringify(page.viewportSize());\"\n```\n\n## 在页面中执行 JS（evaluate）\n\n```bash\n# 获取 userAgent\npw-browser run-code \"return await page.evaluate(() => navigator.userAgent);\"\n\n# 获取所有链接\npw-browser run-code \"\n  return await page.evaluate(() =>\n    [...document.querySelectorAll('a')].map(a => ({ text: a.textContent.trim(), href: a.href }))\n  );\n\"\n\n# 获取 localStorage\npw-browser run-code \"return await page.evaluate(() => JSON.stringify(localStorage));\"\n```\n\n## Iframe 操作\n\n```bash\npw-browser run-code \"\n  const frame = page.frameLocator('iframe#my-iframe');\n  await frame.locator('button.submit').click();\n\"\n```\n\n## 文件下载 ⚠️\n\n> ⚠️ 文件将写入本地磁盘，注意目标路径，避免覆盖已有文件。\n\n```bash\npw-browser run-code \"\n  const [download] = await Promise.all([\n    page.waitForEvent('download'),\n    page.locator('text=下载').click()\n  ]);\n  await download.saveAs('./downloaded-file.pdf');\n  return download.suggestedFilename();\n\"\n```\n\n## 错误处理\n\n```bash\npw-browser run-code \"\n  try {\n    await page.locator('button.submit').click({ timeout: 3000 });\n    return 'clicked';\n  } catch (e) {\n    return 'element not found: ' + e.message;\n  }\n\"\n```\n\n## 复杂场景：多页数据采集\n\n```bash\npw-browser run-code \"\n  const results = [];\n  for (let i = 1; i <= 5; i++) {\n    await page.goto('https://example.com/page/' + i);\n    const items = await page.locator('.item').allTextContents();\n    results.push(...items);\n  }\n  return JSON.stringify(results);\n\"\n```\n\n## 复杂场景：表单填写 ⚠️\n\n> ⚠️ 此操作会真实提交表单，可能触发实际的业务操作（注册账号、下单、发送消息等）。执行前确认目标页面和表单内容已经用户确认。\n\n```bash\npw-browser run-code \"\n  await page.fill('#name', '张三');\n  await page.fill('#email', 'test@example.com');\n  await page.selectOption('#city', '北京');\n  await page.check('#agree');\n  await page.locator('button[type=submit]').click();\n  await page.waitForURL('**/success');\n  return 'form submitted';\n\"\n```\n\n## 注意事项\n\n- `run-code` 中直接使用 Playwright API，无需额外的 `page.evaluate` 包装\n- 客户端请求超时为 120 秒，长耗时操作请合理拆分\n- 返回值自动序列化为字符串，复杂对象请用 `JSON.stringify()`\n- 如果代码中有引号冲突，优先用单引号包裹 JS 字符串\n\nFile v1.0.15:CHANGELOG.md\n\n# Changelog\n\nAll notable changes to `pw-browser` are documented here. This project follows semver-ish versioning (`MAJOR.MINOR.PATCH`).\n\n## [1.3.10] — 2026-07-30\n\n### Security — download path traversal (proactive full-source audit)\n\nEvery prior hardening round was driven by external audit findings, all of which\nconcentrated on the credential primitives. A full sweep of `pw-browser.js` turned\nup an unrelated and **more directly attacker-reachable** hole in `download`.\n\n- **`download` no longer trusts the page-supplied filename.** `dl.suggestedFilename()`\n  is chosen by the *web page being automated*, not by the operator. When the caller\n  passed a destination **directory** (the normal case, e.g. `~/Downloads`), the old\n  code did `path.join(dest, suggested)` — so a malicious page returning\n  `../../.bashrc` wrote **outside** that directory. This is an arbitrary-file-write\n  primitive triggered by merely visiting a hostile site, no operator mistake needed.\n  New module-scoped `safeDownloadTarget()` reduces the name to a bare `path.basename`\n  and then asserts the resolved target still sits inside the resolved destination,\n  returning error kind `PathTraversal` otherwise. Passing an explicit **file** path\n  as `dest` keeps its previous meaning (operator authority, used verbatim).\n- **Constant-time daemon token comparison.** `provided !== daemonToken` short-circuits\n  and leaks length/position through timing. Replaced with `timingSafeToken()` built on\n  `crypto.timingSafeEqual` (length-checked first). Negligible risk against a 24-byte\n  random token, but free to get right — defence in depth for the token that authorises\n  `eval` / `run-code` / credential access.\n\n### Internal\n\n- CLI entry is now guarded by `require.main === module`, and `safeDownloadTarget` /\n  `timingSafeToken` are exported when the file is `require`d — so pure logic can be\n  unit-tested without launching a browser or a daemon.\n\n### Tests\n\n- New `tests/d-download.js` (15 assertions, browser-free — no daemon, no Chrome):\n  traversal names (`../../.bashrc`, `../escape.png`, absolute paths, embedded\n  separators) are all neutralised into `dest`; bare `..`, `a/..`, `../../` and empty\n  names are rejected as `PathTraversal`; ordinary names pass through unchanged.\n  `timingSafeToken` accepts the exact token and rejects wrong-value, shorter, longer,\n  empty and `undefined` without throwing. Wired into `npm test` as the sixth group.\n\n## [1.3.9] — 2026-07-30\n\n### Security — session-persistence hardening (audit findings, 86–98% confidence)\n\nPrevious versions documented two credential-hygiene practices that were **advice only**, never enforced by code. This release turns both into real controls.\n\n- **`PW_BROWSER_CRED_PERSIST=off` — credential-persistence kill switch.** Blocks\n  `cookies export|import` and `storage export|import` (new error kind\n  `CredentialPersistenceDisabled`) while leaving `list` / `get` / `set` / `clear`\n  fully usable. This is deliberately **finer-grained than `PW_BROWSER_SAFE_MODE=1`**,\n  which disables the credential primitives wholesale: the rogue-agent risk is\n  dominated by session state *reaching the filesystem*, where it outlives the\n  daemon and becomes a reusable identity. Removing only that primitive keeps\n  normal automation working. Operator-controlled process env — a caller cannot\n  lift it.\n- **Secret files are now owner-only.** `~/.pw-browser/` is created `0700`; the\n  daemon auth token (`daemon.json`) and every exported cookie/localStorage dump\n  are written `0600`, with an explicit `chmod` so files left `0644` by older\n  versions get tightened on rewrite. Previously `daemon.json` — which holds the\n  token authorising `eval` / `run-code` / credential access — was written with\n  default `0644`, letting **any other local user take over the browser session**.\n  (POSIX only; on Windows mode bits are not OS-enforced and user-profile ACLs apply.)\n\n### Docs\n\n- Mitigation guidance rewritten across SKILL.md / README (zh+en) / QUICKSTART (zh+en):\n  \"remember to `chmod 600`\" → \"already enforced\"; \"don't persist credentials in\n  automated flows\" → \"start the daemon with `PW_BROWSER_CRED_PERSIST=off`\".\n- Deployment matrix gains a **禁凭证落盘 / cred-persist-off** tier between \"full\" and\n  \"safe mode\"; frontmatter `permissions.enforcement` gains `cred-persist-killswitch`\n  and `secret-file-permissions`; execution-boundary list grows from 4 to 6 items.\n\n### Tests\n\n- `b-negative.js` phase 4: `CRED_PERSIST=off` blocks export/import yet in-memory\n  `cookies list` / `storage set` still succeed.\n- `b-negative.js` phase 5: exported credential file and `daemon.json` are `0600`,\n  state dir is `0700` (assertions skipped with an existence check on win32).\n\n## [1.3.8] — 2026-07-30\n\n### Security — credential path confinement is no longer caller-waivable (privilege escalation, 97%)\n\n- **Fixed a privilege-escalation path**: `--unsafe` alone used to lift the credential\n  path confinement. Since the flag is supplied by the *caller* (possibly an untrusted or\n  compromised agent), the guard could be disabled by the very actor it was meant to\n  contain — a guard whose key is held by the constrained party is not a guard.\n- **Privilege separation**: escaping confinement now requires **both**\n  (1) the operator starting the daemon with `PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1`\n  (a process env var a caller sending HTTP commands cannot set), and\n  (2) the caller passing `--unsafe` (intent, not authority).\n  Without (1), `--unsafe` is rejected with `UnsafeOverrideNotPermitted` and nothing is written.\n- **Symlink escape closed**: paths are now resolved via `realpath` on the deepest existing\n  ancestor, so a symlink planted inside `~/.pw-browser` can no longer redirect credential\n  files outside the confinement.\n- **Structured errors**: refusals now carry error kinds `CredentialPathConfined` /\n  `UnsafeOverrideNotPermitted` instead of a bare 400 message.\n- **Audit trail**: every credential-path access (allowed, denied, or operator-escaped) is\n  logged to the daemon's stderr; the daemon also prints a startup warning when the\n  override env is enabled. Responses gain a `confined` boolean.\n- Docs aligned across SKILL.md (frontmatter `permissions.enforcement` +\n  `operator-gated-override`, enforcement-boundary section, Rogue Agent mitigation ⑤),\n  README zh/en, QUICKSTART zh/en, and the `--help` banner (now lists daemon env vars).\n- Tests: `b-negative.js` gains a phase 3 daemon with the operator opt-in; asserts\n  `--unsafe` alone is refused and writes nothing, that the env alone does not weaken the\n  default, and that env + flag yields `confined:false` + warning. `a-group.js` exports now\n  stay inside `STATE_DIR`.\n\n## [1.3.7] — 2026-07-30\n\n### Fixed\n\n- **`eval \"<expr>\" <ref>` (element-scoped) was broken** — the expression was closed over from Node scope\n  (`el.evaluate(el => eval(expr))`). Playwright serializes the callback to a string and runs it in the\n  browser, so `expr` was never defined there and the call failed with a `ReferenceError`. The expression\n  is now passed explicitly as an evaluate argument. This path had **no test coverage**, which is why the\n  defect went unnoticed.\n\n### Changed\n\n- `eval` responses now carry an explicit `scope` field (`\"page\"` | `\"element\"`) so callers can verify\n  which semantics were applied instead of guessing.\n- Result serialization hardened: `JSON.stringify` failures (circular structures, `undefined`) no longer\n  produce a broken payload; they fall back to `String(v)`.\n- Rewrote the `eval` code comment to state the real execution semantics (Playwright does **not** capture\n  Node closures; `el` is the documented binding for the element-scoped form) alongside the security\n  boundary, removing an ambiguity that invited unsafe assumptions about where code actually runs.\n\n### Documentation\n\n- `SKILL.md` / `README.md` / `README.en.md` now document the scope semantics of `eval`: without `ref` the\n  expression evaluates against the page global; with `ref` the identifier `el` is bound to the DOM node.\n  Return values must be JSON-serializable (DOM nodes and circular structures cannot cross the bridge).\n\n### Tests\n\n- `tests/features.js`: +5 assertions covering the previously untested element-scoped `eval` path\n  (`el.textContent`, `el.id.toUpperCase()`, `el.getAttribute(...)`, `scope` field for both forms).\n\n## [1.3.6] — 2026-07-30\n\n### Docs / Security (intent-code divergence, 93% confidence)\n- **修正 `run-code` 沙箱注释夸大安全边界**：原注释称\"user code can drive the browser but cannot reach the host\"，暗示已被锁死，但忽略注入的 `page` 对象本身就是完整浏览器权威（可访问任意站点、读认证会话、触发下载/上传、以用户身份发带凭证请求）。新注释明确：VM 隔离的是**宿主机**（OS/文件系统/进程），**不是浏览器会话/网络**；`page` 等同把用户的浏览器交给代码；必须仅在可信、用户可见的本地环境启用，不可信输入绝不能到达；并澄清 token 认证只限网络可达性、非能力沙箱。SKILL.md/README 既有 `run-code` 描述本已准确，无需改动。\n\n## [1.3.5] — 2026-07-30\n\n### Docs / Security (description-behavior mismatch, 95% confidence)\n- **包描述不再轻描淡写**：`package.json` 的 `description` 原为\"普通浏览器自动化 CLI\"，漏掉了实质高危能力。现改为如实声明 `INCLUDES HIGH-RISK PRIMITIVES`：cookie/localStorage 凭证提取/注入（会话令牌）、页面任意 JS（`eval`）、守护进程代码执行（`run-code`）及持久化凭证会话，并标注 daemon token 认证与 `PW_BROWSER_SAFE_MODE=1` 可禁用全部代码执行与凭证原语。\n- README 中英文开头同步新增 ⚠️ 醒目警示，明确\"本工具不只是点开网页\"，列出代码执行与凭证操控能力，并指向「🔐 权限模型与最低特权」「⚠️ 安全边界」两节，消除与 SKILL.md 已有准确描述的失配。三处（package.json / README.zh / README.en）现已一致。\n\n## [1.3.4] — 2026-07-30\n\n### Docs / Security (LP3 — MCP least-privilege, 90% confidence)\n- **形式化权限模型**：此前 SKILL.md 仅有自由文本 `capabilities` / `allowed-tools`，缺乏可被编排者/审核者解析的权限模型，容易把\"能力清单\"误读成\"授权边界\"从而低估攻击面。本版在 SKILL.md frontmatter 新增**机器可读的 `permissions` 块**（`model: none-formal`、`enforcement`、`least-privilege`、`metadata-is`），并新增章节「🔐 权限模型与最低特权」：\n  - 明确 `capabilities` 是**攻击面描述，不是权限授予、也不是沙箱边界**；\n  - 给出**三级部署矩阵**（全能力 / 安全模式 / 沙箱+安全模式）及每级残余风险；\n  - 列出**实际执行边界**（token 仅限网络可达性、安全模式是唯一内建能力开关、路径限制、宿主工具策略局限）；\n  - 声明本技能**没有**独立于上述四点的形式化权限系统，最低特权须在编排层落实。\n- README 中英文「安全说明」同步补充权限模型提示，与 SKILL.md 三处一致。\n\n## [1.3.3] — 2026-07-30\n\n### Fixed\n- **技能名不合规导致安装失败**：SKILL.md frontmatter 的 `name` 由 `Playwright-browser-use` 改为全小写 `playwright-browser-use`。宿主平台校验规则要求技能名只能包含小写字母、数字与连字符，原大写首字母会在添加技能时报 `Invalid skill name`。\n- 同步将发布镜像目录改名为 `playwright-browser-use`，并更新 `sync.sh` / `make-release.sh` 的默认路径，保证目录名与技能名一致。\n\n## [1.3.2] — 2026-07-27\n\n### Security\n- **安全模式覆盖凭证原语（闭环 Rogue Agent / 会话持久化 97% 置信度发现）**：`PW_BROWSER_SAFE_MODE=1` 现在**整体禁用**全部 `cookies` / `storage` 子命令（`list/get/set/clear/export/import`），与 `eval` / `run-code` 同级返回 `Disabled`。此前安全模式只拦代码执行、拦不住凭证读写——不可信 agent 仍可无代码执行地导出/注入会话凭证做特权持久化；本版从代码层关闭该空档。\n- **文档一致性**：SKILL.md（frontmatter description / capabilities / 核心能力声明 / Rogue Agent 缓解块）、README 中英文、QUICKSTART 中英文的\"safe mode 拦不住 cookies/storage\"表述全部改为新行为，消除描述-行为失配。\n\n### Added\n- `tests/b-negative.js` safe-mode 阶段新增 4 断言：`cookies list/export`、`storage get/import` 在安全模式下均返回 `Disabled`。\n\n## [1.3.1] — 2026-07-27\n\n### Security\n- **会话凭证落盘路径限制（缓解 Rogue Agent / 令牌盗窃）**：`cookies` / `storage` 的 `export` / `import` 默认**限制在 `~/.pw-browser/` 目录内**，禁止静默把凭证写入 `/tmp` 或加载来自任意路径的攻击者构造文件；写到/读自该目录之外必须显式 `--unsafe`（不推荐，命令会拒绝否则）。新增 `CRED_WARNING` 常量，所有 `export`/`import` 的 JSON 响应均附带 `warning` 字段提示文件含实时会话凭证。\n- **审计发现的文档闭环**：针对连续多条安全审计发现（导出未警告、导入+Agent 常态化、会话持久化 Rogue Agent、描述-行为失配、无警告/确认/路径限制），在 QUICKSTART（中英文）示例 4、SKILL.md 核心能力声明与「会话持久化风险」块、README（中英文）命令参考与「Cookie 与本地存储」节统一补齐能力声明、路径限制说明与缓解措施。\n\n### Changed\n- 版本号升至 `1.3.1`，并同步 `package-lock.json`。\n\n## [1.3.0] — 2026-07-27\n\n### Added\n- **超大 DOM 快照保护**：`buildSnapshot` 新增 `SNAP_LIMIT`（默认 3000，环境变量 `PW_BROWSER_SNAP_LIMIT` 可配，设为 `0` 关闭上限）。遍历超上限即停止收集并标记 `truncated`；`snap` 的 JSON 返回 `truncated`、文本输出加截断提示，避免超大型页面拖慢/撑爆快照。\n- **负向/错误路径测试 `tests/b-negative.js`**：点击/填充不存在 ref → `ElementNotFound`；`act` 未知 ref/未知动作 → 对应 `kind`；`storage`/`cookies set` 缺参 → `ok:false`；`open` 缺 url → `ok:false`；`wait-for` 超时 → `WaitTimeout`；安全模式 `eval`/`run-code` → `Disabled`。\n- **快照上限测试 `tests/c-snapcap.js`**：验证 `PW_BROWSER_SNAP_LIMIT` 生效与早期元素仍被捕获。\n- **端到端 Cookbook**：`QUICKSTART.md`（中文）/ `QUICKSTART.en.md`（英文），含表单提交、上传下载、shadow/iframe 穿透、cookie/storage 会话恢复、`act` 多步自纠错五个真实示例。\n- **双语文档**：新增 `README.en.md`（英文 README），与 `README.md` 互相交叉链接；`SKILL.md` 顶部加英文文档指针。\n\n### Changed\n- `resolveRef` 返回结构统一带上 `ok: false`（此前缺省 `ok` 字段），使 `act` 结果条目契约一致。\n- 快照 `refMap` 元素信息补充 `id` 字段（便于识别元素）。\n- `sync.sh` 白名单新增 `README.en.md` / `QUICKSTART.md` / `QUICKSTART.en.md`。\n\n### Security\n- （沿用）全部命令除 `/health` 外均要求随机 token 认证；可经 `PW_BROWSER_SAFE_MODE=1` 彻底禁用代码执行。\n\n## [1.2.0] — 2026-07-27\n\n### Added\n- **`download` 命令**：对称于 `upload`。可选 `<ref>` 触发下载，监听页面 `download` 事件并 `saveAs` 到 `--path`（默认当前目录）；亦可作为 `act` 动作 `{action:\"download\", ref:\"eN\", path:\"/tmp/x.csv\"}` 使用。\n- **`cookies` 一等命令**：`cookies list` / `export [--path file]` / `import <file>` / `clear` / `set <name> <value> [--domain --path]`。不再需要靠 `eval` 曲线救国。\n- **`storage` 一等命令**：`storage get [key]` / `set <key> <value>` / `clear` / `export [--path file]` / `import <file>`，基于 `localStorage`。\n- **Daemon 端口可配 + 冲突避让**：\n  - 环境变量 `PW_BROWSER_PORT` 覆盖默认 `19223`。\n  - 启动时探测端口：若已有 daemon 存活则直接退出（避免双开），否则 `EADDRINUSE` 时自动递增端口直到可用，并把实际端口写回 `daemon.json`。客户端已读 `daemon.json` 端口，无需改动。\n\n### Changed\n- 成熟度工程化：上一轮已加入测试骨架（`tests/smoke.js`）、一键同步脚本（`sync.sh`）、git 初始化。\n\n### Security\n- （沿用 1.1.0）全部命令除 `/health` 外均要求随机 token 认证（优先 `Authorization: Bearer` 头，避免泄漏到 URL/日志）；`/health` 不回传 `pageUrl`；`run-code` 走 `Object.create(null)` VM 沙箱；可用 `PW_BROWSER_SAFE_MODE=1` 彻底禁用代码执行。\n\n## [1.1.0] — 2026-07-27\n\n### Added\n- 借鉴 browser-use 的四项能力：A 稳定元素 ref（跨 snap 保持一致）、B `screenshot --annotate` 视觉标注、C `act` 动作序列 + 自纠错中断、D `history` 操作历史。\n- iframe / open shadow DOM 穿透（`buildSnapshot` 递归 shadow root 与同源 iframe，`cssPath`/`frameChain` 定位）。\n- `upload`、`drag` 命令；`resolveRef` 去重；daemon 空闲自动退出（`PW_BROWSER_IDLE_MS`）+ `shutdown` 加固（`safeCloseBrowser`/`gracefulExit`）。\n- 测试骨架 + `sync.sh` + git 初始化。\n\n### Security\n- token 改为 header 传输；`/health` 去 `pageUrl` 泄漏。\n\nFile v1.0.15:QUICKSTART.en.md\n\n# pw-browser Quickstart Cookbook\n\n> End-to-end examples for AI agents / scripts. Every example can be copied and run as-is.\n> Convention: `pw-browser` stands for the full command\n> `NODE_PATH=\"<SKILL_DIR>/node_modules\" node \"<SKILL_DIR>/pw-browser.js\"`.\n> All examples assume a daemon is already running in the background (`pw-browser daemon &`).\n\n---\n\n## Example 1: Fill a form and submit (most common)\n\nGoal: open a page, locate the input and button, type text and submit.\n\n```bash\n# Open the page\npw-browser open https://example.com/login\n\n# You MUST snap first to get the ref table\npw-browser snap\n# → e1 input placeholder=\"Username\"\n# → e2 input placeholder=\"Password\"\n# → e5 button \"Sign in\"\n\n# Fill and click (using the refs from snap)\npw-browser fill e1 \"alice\"\npw-browser fill e2 \"s3cret\"\npw-browser click e5\n\n# Verify\npw-browser wait-for \"text=Welcome\" --timeout 8000\npw-browser snap\n```\n\nKey point: `snap` must come before `click`/`fill`, and refs stay stable across snaps (e1 this time is still e1 next time).\n\n---\n\n## Example 2: File upload + download round-trip\n\nGoal: upload a local file, then trigger a download and confirm it landed on disk.\n\n```bash\npw-browser open https://example.com/upload\n\npw-browser snap\n# → e3 input type=file \"Choose file\"\n# → e4 button \"Upload\"\n\n# Upload (multiple files supported, space- or comma-separated)\npw-browser upload e3 /tmp/report.pdf /tmp/appendix.xlsx\npw-browser click e4\n\n# Wait for upload to finish\npw-browser wait-for \"text=Upload complete\" --timeout 10000\n\n# Trigger download: click a link/button that downloads, save to a dir\npw-browser snap\n# → e9 a \"Export CSV\"\npw-browser download e9 --path /tmp/downloads --timeout 30000\n# → { \"ok\": true, \"savedPath\": \"/tmp/downloads/export.csv\", \"suggestedFilename\": \"export.csv\" }\n```\n\nNote: `download` is symmetric to `upload`; without `--path` it saves to the current working directory. It can also be an `act` action:\n`{\"action\":\"download\",\"ref\":\"e9\",\"path\":\"/tmp/downloads\"}`.\n\n> 🔒 v1.3.10+: `suggestedFilename` is chosen by the web page and is untrusted input. When `--path` is a directory, the name is reduced to a `basename` and confined inside it; anything escaping returns `PathTraversal`.\n\n---\n\n## Example 3: Shadow DOM / iframe piercing\n\nGoal: some elements live inside a shadow root or a same-origin iframe where normal selectors can't reach — the tool pierces them automatically.\n\n```bash\npw-browser open https://example.com/widget\n\npw-browser snap\n# → e2 button \"Inner button\"  inShadow:true\n# → e7 button \"iframe submit\"  frameChain:[{sel:\"iframe#frame1\"}]\n\n# Just click! Location uses css >>> piercing + frameLocator automatically, invisib\n\nArchive v1.0.14: 19 files, 65819 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), CHANGELOG.md (5771b), LICENSE (1062b), MANIFEST.txt (453b), package-lock.json (747b), package.json (720b), pw-browser (266b), pw-browser.js (64514b), QUICKSTART.en.md (7810b), QUICKSTART.md (7415b), README.en.md (13822b), README.md (12937b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5570b), skill-card.md (2812b), SKILL.md (32150b), _meta.json (142b)\n\nArchive v1.0.13: 19 files, 62256 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), CHANGELOG.md (3801b), LICENSE (1062b), MANIFEST.txt (495b), package-lock.json (747b), package.json (720b), pw-browser (266b), pw-browser.js (62307b), QUICKSTART.en.md (6784b), QUICKSTART.md (6404b), README.en.md (13157b), README.md (12357b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5570b), skill-card.md (3116b), SKILL.md (30062b), _meta.json (142b)\n\nArchive v1.0.12: 19 files, 59837 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), CHANGELOG.md (3801b), LICENSE (1062b), MANIFEST.txt (453b), package-lock.json (747b), package.json (720b), pw-browser (266b), pw-browser.js (62307b), QUICKSTART.en.md (5510b), QUICKSTART.md (5170b), README.en.md (12562b), README.md (11844b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5570b), skill-card.md (2920b), SKILL.md (29063b), _meta.json (142b)\n\nArchive v1.0.11: 14 files, 31583 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), LICENSE (1062b), package-lock.json (1670b), package.json (527b), pw-browser (266b), pw-browser.js (32892b), README.md (5445b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5570b), skill-card.md (2698b), SKILL.md (21737b), _meta.json (142b)\n\nArchive v1.0.10: 13 files, 30395 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), LICENSE (1062b), package.json (527b), pw-browser (266b), pw-browser.js (32522b), README.md (5297b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5459b), skill-card.md (2957b), SKILL.md (21166b), _meta.json (142b)\n\nArchive v1.0.9: 13 files, 30387 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), LICENSE (1062b), package.json (527b), pw-browser (266b), pw-browser.js (32522b), README.md (5297b), references/pagination.md (3377b), references/rich-text-editor.md (3295b), references/running-code.md (5459b), skill-card.md (2851b), SKILL.md (21166b), _meta.json (141b)\n\nArchive v1.0.8: 13 files, 25859 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), LICENSE (1062b), package.json (527b), pw-browser (266b), pw-browser.js (29022b), README.md (4595b), references/pagination.md (2912b), references/rich-text-editor.md (3295b), references/running-code.md (3700b), skill-card.md (2657b), SKILL.md (17000b), _meta.json (141b)\n\nArchive v1.0.7: 13 files, 25746 bytes\n\nFiles: .clawhubignore (44b), .gitignore (39b), LICENSE (1062b), package.json (527b), pw-browser (266b), pw-browser.js (29022b), README.md (4166b), references/pagination.md (2912b), references/rich-text-editor.md (3295b), references/running-code.md (3700b), skill-card.md (2932b), SKILL.md (17000b), _meta.json (141b)","readmeExcerpt":"Skill: playwright-browser-use Owner: yicko Summary: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— cookies / storage 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 PW_BROWSER_SAFE_MODE=1 会将其与代码执行一并禁用）；(2) eval 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) run Tags: latest:1.0.16 Version history: v1.0.16 | 2026-08-01T16:07:30.572Z | us","codeSnippets":[],"executableExamples":[{"language":"text","snippet":"┌──────────────┐     HTTP (localhost:19223)     ┌──────────────┐\n│  pw-browser  │ ──────────────────────────────→│   Daemon     │\n│  (CLI 客户端) │                                │  (浏览器进程)  │\n└──────────────┘                                └──────┬───────┘\n                                                       │\n                                                       ├─ Playwright\n                                                       ├─ Chromium 浏览器\n                                                       └─ 页面状态持久化"},{"language":"bash","snippet":"SKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" daemon &\nsleep 4"},{"language":"bash","snippet":"SKILL_DIR=\"{SKILL_DIR}\"\nNODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" init"},{"language":"bash","snippet":"pw-browser close --all"},{"language":"bash","snippet":"> SKILL_DIR=\"{SKILL_DIR}\"\n> NODE_PATH=\"${SKILL_DIR}/node_modules\" node \"${SKILL_DIR}/pw-browser.js\" <cmd> [args] [--json]\n>"},{"language":"bash","snippet":"# 1. 启动 daemon（会话开始一次）\npw-browser daemon &\n\n# 2. 打开页面\npw-browser open https://www.baidu.com\n\n# 3. 获取页面快照（必须！每次交互前都要 snap）\npw-browser snap\n\n# 4. 交互 — 基于快照中的 e0, e1, e2... ref 引用\npw-browser click e8          # 点击 ref=e8 的元素\npw-browser fill e5 \"hello\"   # 在 ref=e5 的输入框填入文本\npw-browser press Enter       # 键盘按键\n\n# 5. 等待\npw-browser wait-for \"text=加载完成\" --timeout 8000\npw-browser wait-for \"url:https://example.com/*\"\npw-browser wait-for \"state:networkidle\"\n\n# 6. Tab 管理\npw-browser tab list\npw-browser tab select 1\npw-browser tab close 0\n\n# 7. 关闭\npw-browser close             # 关闭当前页面\npw-browser close --all       # 关闭浏览器 + daemon"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: playwright-browser-use\ndescription: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会将其与代码执行一并禁用）；(2) `eval` 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) `run-code` 在守护进程上下文执行 Playwright/Node 代码（vm 沙箱隔离）。全部经持久化本地守护进程（127.0.0.1:19223，浏览器状态跨命令保持）控制，受随机 token 认证保护；`PW_BROWSER_SAFE_MODE=1` 可彻底禁用代码执行与 cookies/storage 凭证读写（v1.3.2+）。仅在可信、用户可见的本地环境中授权使用；会话凭证落盘须遵循后文安全警告。\nallowed-tools: Bash(node:*), Bash(pw-browser:*), Bash(curl:*)\ncapabilities:\n  - \"code-execution: page-context (eval — arbitrary JS in current page)\"\n  - \"code-execution: daemon-vm (run-code — null-prototype VM sandbox, no host fs/process)\"\n  - \"network: arbitrary via browser/page context (credentialed requests possible)\"\n  - \"browser-state: persistent credentialed session across commands\"\n  - \"credential-access: direct read/export/import/clear/set of cookies & localStorage (session tokens) — no code execution needed; gated by daemon token, and fully DISABLED by PW_BROWSER_SAFE_MODE=1 (v1.3.2+)\"\n  - \"file: write to local disk (run-code can trigger downloads; cookies/storage export writes credential files)\"\n# --- Formal permission model (machine-readable) -------------------------\n# NOTE: This skill has NO built-in fine-grained permission system. The\n# fields below are declarative metadata for orchestrators/reviewers, NOT\n# an enforcement boundary. Real constraints come only from:\n#   - daemon token auth (gates NETWORK reachability, not per-capability)\n#   - PW_BROWSER_SAFE_MODE=1 (disables code-exec + credential primitives)\n#   - export/import path confinement to ~/.pw-browser (v1.3.1+); since v1.3.8\n#     the caller-supplied --unsafe flag alone can NOT lift it — the operator\n#     must also start the daemon with PW_BROWSER_ALLOW_UNSAFE_CRED_PATH=1\n#   - PW_BROWSER_CRED_PERSIST=off (v1.3.9) removes credential persistence only\n#     (export/import), keeping in-memory cookie/storage automation usable\n# \"capabilities\" lists the MINIMUM attack surface, not a maximum; the skill\n# can additionally drive any site the browser can reach, including ones where\n# the user is already authenticated. Least privilege is achieved at the\n# orchestration layer via deploy mode (safe mode + sandbox) + scoped ops.\npermissions:\n  model: \"none-formal\"          # no built-in RBAC / capability-dropping\n  enforcement:\n    - \"daemon-token\"            # gates network reachability to 127.0.0.1:19223\n    - \"safe-mode-env\"           # PW_BROWSER_SAFE_MODE=1 disables code-exec + creds\n    - \"path-confinement\"        # cookies/storage IO limited to ~/.pw-browser\n    - \"operator-gated-override\" # lifting confinement needs daemon env, not a caller flag\n    - \"cred-persist-killswitch\" # PW_BROWSER_CRED_PERSIST=off blocks export/import only\n    - \"secret-file-permissions\" # daemon token + credential dumps written 0600, dir 0700\n    - \"downloa"},{"path":"README.md","content":"# pw-browser\n\n> 基于 Playwright 的浏览器自动化 CLI — daemon + client 架构，直接使用系统 Chrome/Edge，无需下载额外浏览器。\n\n> 📘 简体中文文档（本文件）。English documentation: [`README.en.md`](./README.en.md)。端到端示例见 [`QUICKSTART.md`](./QUICKSTART.md)（含 [`QUICKSTART.en.md`](./QUICKSTART.en.md)）。\n\n> ⚠️ **完整能力（含代码执行与凭证操控）**：本工具**不只是\"点开网页\"**——它内置 `eval`（页面上下文**任意 JavaScript** 执行）与 `run-code`（守护进程上下文执行 Playwright/Node 代码）两项**代码执行**能力，以及 `cookies` / `storage` **凭证读写原语**（可**无需代码执行**即提取/注入登录态与会话令牌）。全部经持久化本地守护进程（127.0.0.1:19223）控制，受 daemon token 认证保护；`PW_BROWSER_SAFE_MODE=1` 可**彻底禁用**代码执行与凭证读写（v1.3.2+）。请先阅读下方「🔐 权限模型与最低特权」与「⚠️ 安全边界与使用场景」了解完整攻击面与适用边界，再决定是否授权。\n\n## 特性\n\n- **系统浏览器复用**：通过 Playwright `channel: 'chrome'` 连接系统已安装的 Chrome 或 Edge，不下载 Chromium\n- **跨平台**：支持 Windows、macOS、Linux\n- **持久化会话**：daemon 保持浏览器状态跨命令存活，无需每次重启\n- **可访问性快照**：`snap` 命令生成页面元素树，带 ref 引用，无需写 CSS 选择器\n- **非 headless**：浏览器窗口始终可见，用户可实时监控所有操作\n- **人机协作**：遇到登录/验证码时自动交接给用户操作\n\n## 安装\n\n```bash\ngit clone <repo-url> pw-browser\ncd pw-browser\nnpm install\n```\n\n> 前提：系统已安装 [Node.js](https://nodejs.org/) 18+ 和 Chrome 或 Edge 浏览器。\n\n## 快速开始\n\n```bash\n# 1. 启动 daemon（后台运行）\nnode pw-browser.js daemon &\n\n# 2. 等待 daemon 就绪\nsleep 4\n\n# 3. 打开页面\nnode pw-browser.js open https://www.baidu.com\n\n# 4. 获取页面快照（必须！每次交互前都要 snap）\nnode pw-browser.js snap\n\n# 5. 交互 — 基于快照中的 e0, e1, e2... ref\nnode pw-browser.js fill e12 \"天气预报\"\nnode pw-browser.js click e13\n\n# 6. 查看结果\nsleep 2\nnode pw-browser.js snap\n\n# 7. 关闭\nnode pw-browser.js close --all\n```\n\n## 命令一览\n\n| 类别 | 命令 | 说明 |\n|------|------|------|\n| 生命周期 | `init` / `open <url>` / `close` / `close --all` / `recover` | 启动、导航、关闭、恢复 |\n| 状态感知 | `snap` / `wait-for <target>` | 快照（**ref 跨多次 snap 保持稳定**，同元素同 ref） |\n| 交互 | `click <ref>` / `fill <ref> \"text\"` / `type \"text\"` / `press <key>` / `hover <ref>` / `select <ref> <opt>` / `check <ref>` / `uncheck <ref>` / `upload <ref> <file...>` / `drag <ref源> <ref目标>` / `download <ref> [--path dir]` | 点击、填写、按键、上传、拖拽、下载等 |\n| Cookie/存储 | `cookies list` / `export [--path f]` / `import <file>` / `clear` / `set <name> <value> [--domain d]` · `storage get [key]` / `set <k> <v>` / `clear` / `export [--path f]` / `import <file>` | 读写 cookie 与 localStorage（需先 `open` 真实 http/https 页面） |\n| 导航 | `goto <url>` / `go-back` / `go-forward` / `reload` | 页面导航控制 |\n| Tab | `tab list` / `tab select <idx>` / `tab close <idx>` | 多标签页管理 |\n| 批量 | `act '[{\"action\":\"click\",\"ref\":\"e13\"}, ...]'` | 一次性执行动作序列，中途检测 DOM 变化自动中断重规划 |\n| 历史 | `history [--limit N] [--clear]` | 查看/清空操作历史 |\n| 高级 | `screenshot [--annotate]` / `mousewheel <dx> <dy>` / `eval \"<expr>\" [ref]` ⚠️ / `run-code \"<code>\"` ⚠️ | 截图（`--annotate` 叠加与 ref 对应编号框）、滚动、代码执行。`eval` 不带 `ref` 时在页面全局求值（`scope: \"page\"`）；带 `ref` 时标识符 **`el`** 绑定到该 DOM 节点（`scope: \"element\"`），如 `eval \"el.textContent\" e3`。返回值须可 JSON 序列化 |\n| 延时 | `sleep <seconds>` | 等待 |\n| 对话框 | `dialog-accept [text]` / `dialog-dismiss` | 处理原生 alert/confirm/prompt 对话框 |\n| 守护进程 | `shutdown` | 关闭持久化 daemon，释放浏览器进程 |\n\n> ⚠️ `eval` 在**浏览器上下文**执行 JS（无 Node 权限）；`run-code` 在 daemon 的**受限沙箱（`vm`）**中执行 Playwright 代码（无直接 Node "},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn74y8h6hjpvfa40rgznyt8e3184yaqe\",\n  \"slug\": \"playwright-browser-use\",\n  \"version\": \"1.0.16\",\n  \"publishedAt\": 1785600450572\n}"},{"path":"references/pagination.md","content":"# 分页策略\n\n翻页前必须先**识别页面分页类型**，选对翻页方式。\n\n> 📝 **文档语言与本地化**：本文件为**简体中文**。识别信号表中的中文 / 英文关键词（\"下一页\"/\"Next\"、\"加载更多\"/\"Load more\"）仅为**启发式示例，并非穷举**——非中文页面的实际文案会不同。跨语言页面请优先用 `snap` 返回的 `ref` 或 CSS 选择器（`page.locator('.xxx')`）定位，避免依赖可见文本。完整语言说明见 `SKILL.md`「📝 文档语言与本地化说明」。\n\n## 步骤 1：识别分页类型\n\n从 `pw-browser snap` 的输出判断：\n\n| 类型 | 识别信号 | 翻页方式 |\n|------|---------|---------|\n| **页码分页** | 底部有页码（1,2,3...）、\"下一页\"/\"Next\"/\">\" | `click` 页码或\"下一页\" |\n| **无限滚动** | 底部无分页控件，内容随滚动增加 | `mousewheel` |\n| **加载更多** | 底部有\"加载更多\"/\"Load more\"/\"查看更多\" | `click` 该按钮 |\n\n### 常见网站参考\n\n| 网站 | 分页类型 | 翻页方式 |\n|------|---------|---------|\n| 百度搜索 | 页码分页 | 点击页码 |\n| 淘宝搜索 | 页码分页 | 点击页码 |\n| 京东搜索 | 页码分页 | 点击页码 |\n| 知乎 | 页码分页 | 点击页码 |\n| 小红书 | 无限滚动 | 滚动加载 |\n| 抖音 | 无限滚动 | 滚动加载 |\n\n## 步骤 2：执行翻页\n\n### A. 页码分页\n\n```bash\n# 从 snap 找到\"下一页\"按钮 ref\npw-browser snap\npw-browser click e42          # 点击\"下一页\"\npw-browser sleep 2\npw-browser snap               # 验证\n```\n\n备选方式 — 直接点页码：\n\n```bash\npw-browser snap\n# 找到页码数字（如 \"2\"）对应的 ref\npw-browser click e50\npw-browser sleep 2 && pw-browser snap\n```\n\n### B. 无限滚动\n\n```bash\npw-browser mousewheel 0 800\npw-browser sleep 2\npw-browser snap\n```\n\n连续多次滚动直到内容不再增加。\n\n### C. 加载更多按钮\n\n```bash\npw-browser snap\n# 找到按钮 ref\npw-browser click e30\npw-browser sleep 2\npw-browser snap\n```\n\n## 步骤 3：判断翻页成功\n\n| 分页方式 | 成功信号 | 结束信号 |\n|---------|---------|---------|\n| 页码 | snap 内容变化，URL 可能变化 | \"下一页\"按钮消失或 disabled |\n| 滚动 | snap 出现新元素 | 内容不变，出现\"没有更多了\" |\n| 按钮 | 新内容加载 | 出现\"已加载全部\"，按钮消失 |\n\n## 批量翻页提取\n\n```bash\n# 使用 run-code 批量翻页\npw-browser run-code \"\n  const allResults = [];\n  let hasNext = true;\n  while (hasNext) {\n    const items = await page.locator('.item').all();\n    for (const item of items) {\n      allResults.push(await item.textContent());\n    }\n    const nextBtn = page.locator('text=下一页');\n    if (await nextBtn.count() === 0 || await nextBtn.isDisabled()) {\n      hasNext = false;\n    } else {\n      await nextBtn.click();\n      await page.waitForTimeout(2000);\n    }\n  }\n  return JSON.stringify(allResults);\n\"\n```\n\n## 常见问题\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| 滚动后不加载新内容 | 实际上是页码分页 | 检查 snap 底部是否有页码，改为 click |\n| 点击页码没反应 | 按钮 disabled 或需要等待 | `sleep 1` 后再点击 |\n| 翻页后内容相同 | AJAX 加载，需要等待 | 延长 sleep 时间或用 `wait-for` |\n| 页码按钮被遮挡 | 需先滚动到底部 | `mousewheel 0 1000` 再 snap |"},{"path":"references/rich-text-editor.md","content":"# SPA 与富文本编辑器 ⚠️\n\n> ⚠️ **破坏性操作警告：** 以下操作会真实修改网页内容（知识库文档、CMS 页面等）。执行前确认：① 处于编辑/草稿状态而非已发布内容；② 修改内容已经用户确认；③ 保存/发布操作不可逆。\n\n处理知识库、文档系统、CMS（如 Notion/语雀/飞书类页面）中的 SPA 编辑态和富文本编辑器写入。\n\n## 何时使用\n\n满足**任一条件**时，遵循本指南：\n\n- URL/页面属于知识库、文档、笔记、CMS 类站点\n- 任务要求创建/编辑/保存文档正文\n- 点击\"编辑\"按钮后 URL 不变但页面状态变化\n- snap 中出现 `contenteditable`、编辑器 toolbar、\"插入\"/\"正文\"等\n- 表单不是普通 input/textarea，而是复杂编辑器\n\n## 核心原则\n\n- **不要**直接用 `innerText`/`textContent` 写 RTE——不会被编辑器状态机接受\n- 先确认进入编辑态，再用 Playwright 键盘输入\n- 保存后验证内容而不是只看按钮状态\n\n## 流程\n\n### 1. 进入编辑态\n\n```bash\n# 先 snap 找到编辑按钮\npw-browser snap\n\n# 点击编辑按钮\npw-browser click <编辑ref>\n\n# 验证进入编辑态\npw-browser run-code \"\n  return await page.evaluate(() => ({\n    hasUpdate: Array.from(document.querySelectorAll('button'))\n      .some(b => b.textContent.trim() === '更新'),\n    editableCount: document.querySelectorAll('[contenteditable]').length\n  }));\n\"\n```\n\n如果 `hasUpdate=true` 或 `editableCount > 0`，继续；否则尝试重试点击。\n\n### 2. 写入内容（RTE）\n\n```bash\npw-browser run-code \"\n  const editor = page.locator('[contenteditable=\\\"true\\\"], [contenteditable=\\\"plaintext-only\\\"]').first();\n  await editor.click();\n  await page.keyboard.press('Control+A');\n  await page.keyboard.type('要写入的文本内容');\n  await page.waitForTimeout(1000);\n  const text = await editor.textContent();\n  return text;\n\"\n```\n\n### 3. 保存 ⚠️\n\n> ⚠️ 保存/发布操作不可逆，确认内容无误后再执行。\n\n```bash\npw-browser run-code \"\n  await page.evaluate(() => {\n    const btn = Array.from(document.querySelectorAll('button'))\n      .find(b => ['更新','保存','发布','完成'].includes(b.textContent.trim()));\n    btn?.click();\n  });\n  await page.waitForTimeout(3000);\n\"\n```\n\n### 4. 验证\n\n```bash\npw-browser run-code \"\n  const title = document.querySelector('h1, [class*=title]')?.textContent || document.title;\n  const main = document.querySelector('main') || document.body;\n  const blocks = Array.from(main.querySelectorAll('p, h1, h2, h3, li'))\n    .map(b => b.textContent.trim().slice(0, 80))\n    .filter(Boolean);\n  return JSON.stringify({ title, sampleBlocks: blocks.slice(0, 10) });\n\"\n```\n\n成功标准：\n- `hasUpdate=false`（编辑态已退出）\n- 正文包含目标文本\n- 标题未被误改或清空\n- 没有重复写入的文本\n\n## 弹窗处理\n\n- DOM 浮层（弹窗、抽屉、popover）：通过 snap 识别并 click 关闭按钮\n- 原生 JS dialog（alert/confirm/prompt）：用 `pw-browser dialog-accept` / `dialog-dismiss`\n\n## 失败处理\n\n- 点击\"编辑\"超时后，先检查是否已进入编辑态，不要重复点击\n- 后续命令超时，执行 `pw-browser recover` 恢复 daemon\n- 恢复后如果已在编辑态，继续输入和保存"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— `cookies` / `storage` 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 `PW_BROWSER_SAFE_MODE=1` 会将其与代码执行一并禁用）；(2) `eval` 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) `run Skill: playwright-browser-use Owner: yicko Summary: 浏览器自动化 CLI（Playwright 版，纯 Node.js 实现）。除常规自动化（打开网页/截图/点击/填表/翻页）外，提供三类能力：(1) 会话凭证读写原语 —— cookies / storage 命令可**无需代码执行**即列出/导出/导入/清除/设置 cookie 与 localStorage，直接提取或注入登录态与会话令牌（此路径独立于代码执行；自 v1.3.2 起 PW_BROWSER_SAFE_MODE=1 会将其与代码执行一并禁用）；(2) eval 在页面上下文执行任意 JavaScript（可读 cookie/存储、发起带凭证请求）；(3) run Tags: latest:1.0.16 Version history: v1.0.16 | 2026-08-01T16:07:30.572Z | us","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1208,"uniquenessScore":50,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T12:29:31.512Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T12:29:31.512Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T14:43:24.597Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}