```html

Automating GetMyBoat Lead Ingestion: Building a Headless Browser Pipeline with Playwright and Structured Data Extraction

When manual lead tracking became a bottleneck for the Sail JADA operations team, we needed a reliable way to ingest and parse GetMyBoat inbox conversations at scale. This post details the architecture we built to automate GetMyBoat login, inbox scraping, and conversation extraction—including the infrastructure decisions, failure modes we encountered, and what comes next.

The Problem: Manual Lead Review at Scale

Carole was manually checking the GetMyBoat inbox for warm leads, copying conversation snippets into documents, and cross-referencing them with pricing and availability. This workflow didn't scale. We needed:

  • Automated login and session persistence across runs
  • Full inbox conversation tree extraction (not just the latest message)
  • Structured parsing into a reportable format (markdown with conversation excerpts, pipeline value, and offer pricing)
  • Email delivery of daily/weekly summaries

Technical Architecture: Headless Playwright + Profile Persistence

We chose Playwright over Selenium for three reasons: native support for persistent browser profiles (avoiding repeated login), built-in waits for SPAs, and synchronous install into our existing Python environment with Google API client libraries.

Step 1: Environment Setup

Created a dedicated virtual environment with both Playwright and Google Workspace credentials:

python3.11 -m venv /tmp/gmb-automation
source /tmp/gmb-automation/bin/activate
pip install playwright google-auth-oauthlib google-auth-httplib2 google-api-python-client

Installed Chromium via Playwright's synchronous API:

python -m playwright install chromium

This downloaded the matching Chromium build (~200MB) and validated the browser binary was executable.

Step 2: Persistent Profile-Based Login

Rather than logging in fresh each run, we created a persistent Chrome profile directory at /Users/cb/Documents/repos/jada-ops/gmb-profile. The login script (/tmp/gmb_login.py) launches Chromium in headed mode (with visible UI) and waits for human interaction:

browser = playwright.chromium.launch(
    headless=False,
    args=["--disable-blink-features=AutomationControlled"],
    user_data_dir="/Users/cb/Documents/repos/jada-ops/gmb-profile"
)
context = browser.new_context()
page = context.new_page()
page.goto("https://getmyboat.com/login")
# Human logs in interactively; profile saves cookies/session storage
page.wait_for_url("**/inbox**", timeout=120000)  # Wait for successful nav to inbox

Once the human completes login, the profile is saved. Subsequent runs can launch in headless mode (no UI) and skip authentication entirely.

Step 3: SPA Navigation and Inbox Discovery

GetMyBoat's inbox is a Single Page Application. We needed to wait for the actual inbox URL to appear after login, not just the landing page:

page.wait_for_url(
    "https://www.getmyboat.com/app/captain/messages/*",
    timeout=30000
)
inbox_url = page.url
print(f"Inbox discovered at: {inbox_url}")

This ensures we're actually at the conversations panel, not a redirect or loading state.

Step 4: Conversation Tree Extraction

The inbox displays a list of conversation threads. Each thread header contains: - Renter name and avatar - Last message preview - Listing title (yacht/boat name) - Timestamp We scraped this metadata, then clicked into each thread to extract the full conversation history:

threads = page.query_selector_all('[data-testid="conversation-item"]')
for thread in threads:
    thread.click()
    page.wait_for_selector('[data-testid="message-panel"]', timeout=10000)
    
    # Extract all messages in the current thread
    messages = page.query_selector_all('.message-bubble')
    thread_data = {
        "renter": thread.inner_text().split('\n')[0],
        "listing": thread.inner_text().split('\n')[1],
        "messages": [msg.inner_text() for msg in messages]
    }
    all_threads.append(thread_data)

Data Pipeline: From HTML to Markdown Report

Once conversations were extracted, we parsed them into a structured JSON/dict format, then synthesized a markdown report including:

  • Pipeline Summary: Total value of open inquiries, broken down by status (inquiry, quoted, awaiting response)
  • Representative Threads: 3–5 conversations with highest engagement or deal maturity
  • Pricing Context: Extracted offer amounts from each thread to correlate with Sail JADA's pricing model
  • Action Items: Conversations awaiting response from Carole, with reply templates

The report was saved to /Users/cb/Documents/repos/jada-ops/gmb-inbox-report-{date}.md and emailed via Gmail API to c.b.ladd@gmail.com.

Key Infrastructure Decisions

Why Playwright (Not Selenium or Puppeteer)

Playwright's user_data_dir parameter eliminates login on every run—a massive time saver for interactive workflows. Selenium lacks this. Puppeteer (Node.js only) wasn't an option given our Python + Google API stack.

Headed Mode for Initial Login, Headless for Automation

GetMyBoat may use bot detection. Headed mode allows the human to complete CAPTCHA/2FA interactively, then the profile is reused. Headless mode is faster for subsequent scrapes, provided the session doesn't expire.

Structuring Files for Reusability

We split scripts by concern:

  • /tmp/gmb_login.py — Interactive login, saves profile
  • /tmp/gmb_session.py — Utilities to launch browser with saved profile
  • /tmp/gmb_explore.py — Navigate to inbox, find correct URL
  • /tmp/gmb_inbox.py — Extract thread list and metadata
  • /tmp/gmb_scrape.py — Full conversation extraction and parsing

Each module imports from the prior one, creating a clean dependency chain.

Challenges & What We Learned

Challenge 1: Login Timeout The initial Playwright login timed out (120s default). We increased the timeout to 180s and added explicit waits for the inbox URL pattern, not just any navigation.

Challenge 2: Stale Session Cookies GetMyBoat sessions expire after ~24 hours. The persistent profile solves this for day-of runs, but cross-day runs require manual re-login. We added a heartbeat check (ping the inbox URL) to detect expired sessions and trigger re-login