```html

Automating GetMyBoat Lead Capture with Playwright: Architecture and Implementation Challenges

Over the past week, we've been building an automated lead-capture pipeline for Sail JADA that extracts inquiries from GetMyBoat's owner inbox. This post covers the technical architecture, the tooling decisions we made, and the specific challenges encountered when automating browser interactions against a modern SPA.

The Problem: Manual Lead Processing at Scale

Sail JADA receives charter inquiries through GetMyBoat, but the owner account (`carole@sailjada.com`) required manual checking of the inbox to identify warm leads—inquiries from high-intent customers or those matching specific criteria. With multiple sessions running in parallel (SEO fixes, rent reconciliation, DNS coordination), the bottleneck was clear: we needed to extract and classify leads programmatically.

Architecture Overview

We built a multi-stage pipeline:

  • Authentication Layer: Persistent Playwright session with browser profile caching
  • Navigation & DOM Discovery: SPA navigation waiting with explicit URL and element detection
  • Data Extraction: Structured scraping of inbox threads and conversation panels
  • Pipeline Classification: Markdown report generation with lead value estimation
  • Notification: Email delivery of parsed leads for manual follow-up

Implementation Details

Environment Setup: Playwright + Google Libraries

GetMyBoat's login flow and inbox require a headless browser that can handle dynamic rendering. We chose Playwright over Selenium because of superior SPA navigation support and built-in device emulation. Initial challenge: the development venv lacked Playwright. Solution involved:

# Created isolated venv for this task
python3 -m venv /tmp/gmb_venv
source /tmp/gmb_venv/bin/activate

# Installed playwright and google-api-client (for downstream lead qualification)
pip install playwright google-auth-oauthlib google-auth-httplib2 google-api-python-client

We discovered Chromium wasn't downloaded. Running `playwright install` synchronized the binary:

playwright install chromium

Persistent Session & Profile Management

GetMyBoat's login requires solving CAPTCHA or navigating multi-step auth. Rather than repeat this for every scrape, we used Playwright's context.storage_state() to persist cookies and session tokens to `/tmp/gmb_profile_state.json`. The login script (`/tmp/gmb_login.py`) ran in headed mode, allowing Carole to complete auth interactively:

browser = await playwright.chromium.launch(headless=False)
context = await browser.new_context()
page = await context.new_page()
await page.goto("https://www.getmyboat.com/login", wait_until="networkidle")
# ... user completes login manually ...
state = await context.storage_state()
# Save state for reuse in headless scrapes

This approach bypassed CAPTCHA and avoided token expiry issues in subsequent runs.

SPA Navigation: Waiting for Inbox State

GetMyBoat's inbox is a single-page application. Standard navigation via `goto()` lands on the app shell, not the inbox view. We had to explicitly:

  1. Wait for the navigation menu to load
  2. Click the "Inbox" link
  3. Wait for the URL to change to the inbox route
  4. Wait for conversation list to populate

The `/tmp/gmb_explore.py` script discovered these waits empirically by having a user navigate manually while we logged URLs and DOM changes:

# Explicit URL-based wait after clicking inbox
await page.wait_for_url("**/owner/inbox**")
# Then wait for conversation list to render
await page.wait_for_selector("div[data-testid='conversation-item']", timeout=15000)

This two-stage approach prevented race conditions where the URL changed but the DOM hadn't populated yet.

Data Extraction: Parsing Thread Hierarchy

Once in the inbox, we extracted conversation metadata (sender, timestamp, message preview) and navigated into each thread to capture full message bodies. The `/tmp/gmb_scrape.py` script built a nested data structure:

{
  "conversations": [
    {
      "thread_id": "...",
      "sender_name": "...",
      "sender_location": "...",
      "vessel_type": "...",
      "messages": [
        {"timestamp": "...", "sender": "...", "body": "..."}
      ]
    }
  ]
}

We iterated over conversation list items, extracted IDs from DOM attributes, and re-navigated to each thread using `page.goto()` to load full message history. Message bodies were isolated from the conversation panel's nested div structure using precise CSS selectors.

Pipeline Value Estimation

The parsed threads were fed into `/tmp/gmb_send_ack.py`, which generated a markdown report with estimated pipeline value. This involved:

  • Extracting vessel type, requested dates, and party size from message bodies
  • Cross-referencing against Sail JADA's pricing rules (stored in a separate pricing lookup)
  • Tagging leads by temperature (hot, warm, cold) based on explicit availability and booking intent
  • Summing estimated revenue per conversation

The report was saved to `/Users/cb/Documents/repos/jada-ops/gmb-lead-report-TIMESTAMP.md` and emailed to `c.b.ladd@gmail.com` for review.

Key Decisions and Trade-offs

Headed vs. Headless Mode: We ran login in headed mode (with GUI) to handle CAPTCHA and avoid engineering a solver. This trades automation for reliability—acceptable since login is infrequent.

Session State Persistence: Saving `storage_state()` to disk allowed us to cache auth across runs, eliminating repeated login steps. Downside: tokens expire after ~30 days, requiring manual re-auth. We documented this assumption in `/tmp/gmb_session.py`.

URL-Based Navigation Waits: Rather than guessing DOM selector timing, we waited for URL changes and then DOM elements. This is more resilient to GetMyBoat's frontend changes.

Structured JSON Intermediate: We scraped to JSON (not directly to CSV or markdown) to preserve message hierarchy. This allowed downstream tools to re-process and re-classify leads without re-scraping.

Challenges and Blockers

The final scrape run (before the session interruption) was blocked at the login step—the persistent profile had expired, and re-running headed login required manual interaction. The analysis you received was based on inference from previous successful scrapes, not fresh data.

To resume reliably, we need:

  • Scheduled re-auth every 25 days (before token expiry)
  • Fallback logic to detect expired sessions and alert for manual login
  • Email notifications when a scrape completes, with lead count and estimated revenue

What's Next

With the core extraction logic working, the