$npx -y skills add getpaperclipai/paperclip --skill agent-browserDrive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.
| 1 | # Agent Browser |
| 2 | |
| 3 | Use a controlled browser to verify behavior, capture evidence, or extract information from web pages that a static fetch cannot reach (SPAs, login-gated pages, dynamic content). This skill is about supervised verification, not unattended scraping. |
| 4 | |
| 5 | ## When to use |
| 6 | |
| 7 | - You need a screenshot of a deployed page or a local dev server to confirm a UI change. |
| 8 | - You need to read JavaScript-rendered content that `curl`/`wget` will not see. |
| 9 | - A user reports a UI bug and you need to reproduce it interactively to capture console errors, network requests, or layout state. |
| 10 | - You need to walk through a short flow (load page, click, observe) to verify acceptance criteria. |
| 11 | |
| 12 | ## When not to use |
| 13 | |
| 14 | - The page is reachable as static HTML. Use `curl`/HTTP fetch — it is cheaper, faster, and more reliable. |
| 15 | - The task is unattended large-scale scraping. That belongs to a dedicated scraper with rate limits, robots.txt handling, and a real user agent policy — not this skill. |
| 16 | - The site is behind authentication you do not own credentials for, or whose terms of service prohibit automation. |
| 17 | - The site involves sensitive accounts (banking, healthcare, government) where automation risks lockout or compliance issues. |
| 18 | |
| 19 | ## Before launching the browser |
| 20 | |
| 21 | - Confirm the URL and what state should be true after navigation. |
| 22 | - Decide what evidence is needed: full-page screenshot, viewport screenshot, console log, network trace, HTML snapshot, extracted text. |
| 23 | - Decide the viewport size that matters for the task (mobile vs desktop). Default to a desktop size unless the task is mobile-specific. |
| 24 | - For local dev servers, confirm the server is running and the port is what you expect. |
| 25 | |
| 26 | ## Driving the browser |
| 27 | |
| 28 | A typical verification session: |
| 29 | |
| 30 | 1. **Launch with a real-looking user agent** when the target is the public internet; an unrealistic UA flags automation traffic. |
| 31 | 2. **Set a sane viewport** (e.g., 1366×768 desktop, 390×844 iPhone-ish). |
| 32 | 3. **Navigate and wait for the right signal.** Prefer waiting for a specific selector or network-idle over arbitrary sleeps. |
| 33 | 4. **Capture evidence immediately** after the wait condition succeeds, before any interaction perturbs the state. |
| 34 | 5. **Interact deliberately.** One click at a time, with a wait between actions; re-screenshot after each meaningful state change. |
| 35 | 6. **Read the console and network panels** for unexpected errors, 4xx/5xx responses, or slow requests. |
| 36 | 7. **Close the browser cleanly** when done. Long-running browser sessions leak memory and hold ports. |
| 37 | |
| 38 | ## What evidence to record |
| 39 | |
| 40 | For a verification task, deliver: |
| 41 | |
| 42 | - A full-page or viewport screenshot of each meaningful state. |
| 43 | - The console log, filtered to warnings/errors. |
| 44 | - Any non-2xx network response with the URL, status, and a short response body excerpt. |
| 45 | - A short narration: "Navigated to X, observed Y, clicked Z, observed W." |
| 46 | |
| 47 | For a UI bug repro, also record: |
| 48 | |
| 49 | - The exact reproduction steps the user can follow. |
| 50 | - Viewport size and (where relevant) device pixel ratio. |
| 51 | - Whether the bug reproduces on first load vs after interaction. |
| 52 | |
| 53 | ## Login-gated pages |
| 54 | |
| 55 | - Prefer programmatic auth (API token, magic link) over UI login. |
| 56 | - If UI login is the only path, the user must provide credentials explicitly for this run. Never reuse credentials outside the session. |
| 57 | - Do not store credentials in the session log, screenshot, or returned output. |
| 58 | |
| 59 | ## Performance and politeness |
| 60 | |
| 61 | - Throttle to one navigation per few seconds when touching shared infra. |
| 62 | - Respect `robots.txt` for public sites you are inspecting at any volume. |
| 63 | - Cancel navigations if a page exceeds a reasonable timeout (e.g., 30s); the page is broken or rate-limiting you. |
| 64 | - Do not retry forever on failure. Retry once with a longer timeout, then escalate. |
| 65 | |
| 66 | ## Common failure modes |
| 67 | |
| 68 | - **Selector not found.** Page changed, or you are waiting before render. Take a screenshot to see actual state; adjust the selector. |
| 69 | - **Click does nothing.** The element is offscreen, covered by a modal, or in a shadow DOM. Scroll into view or pierce the shadow root. |
| 70 | - **Headless detection.** Some sites detect headless Chrome and serve a different page. Use a non-headless mode or a fingerprint-realistic configuration only when authorized. |
| 71 | - **Cross-origin iframe blocking.** Iframes you do not own cannot be inspected; the page must offer the data outside the iframe or the task is infeasible. |
| 72 | |
| 73 | ## Anti-patterns |
| 74 | |
| 75 | - Long unsupervised browser sessions that drift from the original task. |
| 76 | - Scraping behind authentication you do not own. |
| 77 | - Captioning a screenshot with "looks good" without saying what state was loaded and what s |