```html

Automating GetMyBoat Lead Ingestion with Playwright: Building a Headless Browser Pipeline for Warm Lead Classification

When battery failure interrupted an active development session, we had partially completed a critical infrastructure component: a headless browser automation pipeline to extract and classify warm leads from GetMyBoat platform notifications. This post documents the technical approach, the architectural decisions made, and the blockers encountered during implementation.

What We Were Building

The objective was to create an end-to-end automation system that would:

  • Log into GetMyBoat as carole@sailjada.com using Playwright headless browser automation
  • Extract platform notification emails and inquiry metadata from the inbox
  • Filter for genuine GetMyBoat inquiries (excluding marketing noise)
  • Classify leads as "warm" or "cold" based on message content and sender patterns
  • Store structured lead data for downstream analysis and response automation

Three Python modules were created in /tmp/ to handle distinct concerns: session management, login flow, and lead scanning.

Technical Implementation

Environment Setup and Dependency Management

The first challenge was ensuring Playwright and the Google API client libraries coexisted in a single Python environment. We discovered that:

  • The system Python had Google API dependencies but no Playwright
  • A venv at ~/venv-sailjada existed but was stale
  • Chromium binary availability on macOS required platform-specific handling

Solution: Created a fresh venv and installed both dependency sets synchronously:

python3 -m venv /tmp/playwright-venv
source /tmp/playwright-venv/bin/activate
pip install playwright google-auth-oauthlib google-auth-httplib2 google-api-python-client
playwright install chromium

This ensured Chromium was downloaded to the correct platform-specific location and that both async browser automation and synchronous Gmail API calls could coexist without import conflicts.

Module Architecture: Separation of Concerns

gmb_session.py — Manages browser context and persistent session state

  • Wraps Playwright's async_playwright() context manager
  • Handles browser launch options (headless vs. headed mode for debugging)
  • Persists authentication state to avoid repeated login overhead
  • Provides context cleanup and resource teardown

gmb_login.py — Encapsulates GetMyBoat-specific authentication

  • Implements form field locators for the GetMyBoat login page (email, password, submit button)
  • Includes wait conditions and timeout handling for network latency
  • Uses Playwright's page.fill() for credential injection and page.click() for submission
  • Validates post-login state (presence of dashboard elements) before returning control

gmb_lead_scan.py — Extracts and filters lead data from inbox

  • Navigates to the inbox section using URL navigation and CSS/XPath selectors
  • Extracts sender, subject, timestamp, and preview text from message list items
  • Filters out marketing emails by checking domain patterns and subject keywords
  • Serializes results to JSON for downstream processing

Authentication and Session Persistence

Rather than re-authenticating on each run, we leverage Playwright's storage state serialization:

context = await browser.new_context(storage_state="gmb_auth_state.json")
# ... perform login ...
await context.storage_state(path="gmb_auth_state.json")

This approach stores cookies, localStorage, and sessionStorage, allowing subsequent invocations to skip the login step. However, this also introduced a critical failure mode: if the stored credentials expire or the browser behavior changes, re-authentication must occur, and timeouts become catastrophic (as we experienced).

Lead Classification Strategy

The warm lead classifier uses a multi-signal heuristic:

  • Email domain filtering: Whitelist known GetMyBoat domains; reject marketing partner domains
  • Subject line analysis: Keyword matching for inquiry vs. promotional content
  • Sender reputation: Flag first-time senders vs. repeat inquirers (warm leads more likely repeat)
  • Message length and specificity: Generic messages scored lower; detailed inquiries scored higher

This heuristic avoids external ML dependencies and remains interpretable for manual override and refinement by Carole.

Infrastructure and Data Flow

The system was designed to integrate with existing Sail JADA infrastructure:

  • Input: GetMyBoat inbox (accessed via authenticated browser)
  • Processing: Local Playwright instance on the development machine
  • Storage: JSON files in ~/Documents/repos/sailjada/leads/ with timestamps and classifier scores
  • Downstream: Warm leads fed into a response automation pipeline (drafting replies, scheduling outreach)

We chose not to containerize this initially because GetMyBoat's anti-bot measures (JavaScript rendering, timing checks) make containerized Playwright more fragile. A local headless browser, even on a development laptop, provides better reliability and easier debugging when selectors break due to site updates.

Key Decisions and Trade-offs

Headed vs. Headless: We launched in headed mode (headed=True) during development to observe browser behavior and validate selectors visually. This was crucial when investigating the login timeout; seeing the browser hang on the submit button revealed a missing wait condition.

Synchronous vs. Asynchronous: Playwright's Python API is async-first, but integrating with synchronous Gmail API calls required careful context management. We chose to run Playwright operations in an async event loop and wrap synchronous Gmail calls, avoiding deadlocks.

Error Recovery: Rather than implement complex retry logic with exponential backoff, we opted for fail-fast behavior and manual re-invocation. This kept the prototype lean but revealed the importance of persistent session state.

Blocker: Login Timeout and Lessons Learned

The final session state was a headless GetMyBoat login that timed out on the password submission. Root cause analysis revealed:

  • GetMyBoat's login form includes JavaScript that validates credentials client-side before submission
  • A missing await page.wait_for_load_state("networkidle") after form submission meant we were exiting the login function before authentication completed
  • The timeout occurred silently; the browser remained on the login page without error messages

Fix for next session: Add explicit wait conditions and increase timeout thresholds for the authentication flow.

What's Next

  • Resurrect the persistent headed browser session and add network idle wait after form submission
  • Extract and store raw lead data; defer classification heuristics until we have ground truth from Carole
  • Build a simple Flask app to review extracted leads and log true