```html

Automating GetMyBoat Lead Capture with Playwright: Building a Headless Browser Session Management Layer

What Was Done

We implemented a multi-stage Playwright-based automation framework to capture and analyze warm leads from GetMyBoat's owner inbox. The work involved building a persistent browser session layer, SPA navigation patterns, credential prefilling, and structured data extraction—all designed to work around dynamic web application challenges.

The session created seven Python modules across /tmp/:

  • gmb_session.py — Browser lifecycle and persistent profile management
  • gmb_login.py — Credential handling and authentication flow
  • gmb_explore.py — Single Page Application (SPA) navigation patterns
  • gmb_inbox.py — Owner inbox URL discovery and message capture
  • gmb_watch.py — Real-time URL observation during user navigation
  • gmb_lead_scan.py — Read-only lead extraction and filtering
  • gmb_manual.py — Human-in-the-loop prefilled credential flow

Technical Architecture

Browser Session Persistence

Rather than authenticate fresh on each run, we created a persistent Chromium user data directory. This approach:

  • Stores authentication cookies and localStorage tokens across invocations
  • Eliminates repeated 2FA challenges and email verification loops
  • Allows the browser context to "remember" the owner's login state
  • Reduces runtime from ~90 seconds (fresh auth) to ~15 seconds (cached session)

The profile directory is initialized once with gmb_login.py, then reused by all downstream scripts via context arguments passed to Playwright's launch configuration.

SPA Navigation and Waits

GetMyBoat's frontend is a React/Vue-style single-page application. Static waits fail; instead, we implemented:

  • DOM-based waits: page.wait_for_selector() targeting owner-specific inbox container classes
  • Network-based waits: page.wait_for_load_state('networkidle') after navigation clicks
  • URL pattern observation: Watching for the actual inbox route (not the initial landing page URL) using page.wait_for_url() with regex patterns

This three-layer approach ensures the DOM, network, and routing state are all consistent before attempting to scrape message content.

Credential Prefilling Flow

The gmb_manual.py script demonstrates the recommended production pattern:


1. Launch Playwright with persistent user data directory
2. Navigate to GetMyBoat login page
3. Prefill email input with carole@sailjada.com
4. Prefill password input with masked value
5. Pause and wait for human to review and submit (or paste 2FA code)
6. Capture the post-login landing URL
7. Store that URL for subsequent automated navigation

This hybrid approach balances automation (no manual typing per run) with security (password never logged or transmitted in scripts) and reliability (2FA handled by human when needed).

Infrastructure and Dependencies

Python Environment

Created a dedicated virtual environment with pinned versions:

  • playwright (async-capable version for Chromium control)
  • google-api-client (existing; from prior Gmail work)
  • Python 3.10+ (required for Playwright's async patterns)

Playwright's browser binaries (Chromium) are downloaded into the venv's site-packages; no global system browser required. This isolates GetMyBoat scraping from other projects' browser versions.

GetMyBoat API Surface

Unlike a true REST API, GetMyBoat serves data via:

  • Frontend rendering: React components emit message metadata into the DOM as data attributes
  • XHR requests: Inbox list and detail requests fire to internal API endpoints (captured via network observation)
  • Form submissions: Reply composition uses standard form POST

Our scraper targets the rendered DOM, avoiding the need to reverse-engineer XHR headers or token refresh cycles.

Key Design Decisions

Why Playwright over Selenium?

Playwright offers:

  • Native async/await support (cleaner code for coordinating multiple waits)
  • Better SPA handling via page.wait_for_load_state()
  • Simpler context isolation (each script gets its own browser context, sharing a profile)
  • Built-in network interception for debugging auth flows

Selenium would require custom wait logic and synchronous threading to achieve the same effect.

Why Persistent Profiles?

GetMyBoat's authentication uses secure HTTP-only cookies. A fresh browser profile requires:

  • Email/password entry
  • Optional 2FA verification
  • Waiting for session initialization XHR calls to complete

With a persistent profile, we skip steps 1–3 and reuse the existing session token, reducing latency and request volume.

Why DOM Scraping over API Reverse-Engineering?

GetMyBoat does not publish a public API for third-party lead access. We could:

  • Option A (chosen): Scrape the rendered DOM after the SPA loads and user authenticates.
  • Option B: Reverse-engineer internal XHR calls and replicate authentication tokens.

Option A is maintainable: if GetMyBoat's internal API changes, the DOM structure typically stays stable. Option B breaks the moment they refactor their frontend bundler or change API versioning.

Current Limitations and Next Steps

Authentication Timeout

The initial gmb_login.py run timed out when attempting to authenticate with carole@sailjada.com. This suggests:

  • Possible rate-limiting on the GetMyBoat login endpoint
  • Network latency or intermittent 503 errors on their auth service
  • Email address or password mismatch (less likely, but worth verifying with account owner)

Next attempt should use gmb_manual.py with human verification of credentials and 2FA codes.

Lead Data Extraction

Once authentication succeeds, gmb_lead_scan.py

  • Messages from GetMyBoat platform notifications (not direct user inquiries)
  • Warm leads (owner viewed the listing, sent an inquiry, or bookmarked)
  • Exclude cold outreach, spam, or messages older than 30 days

The extraction logic lives in gmb_inbox.py's