Automating GetMyBoat Lead Capture with Playwright: Building a Headless Browser Session Management Layer
What Was Done
We implemented a multi-stage Playwright-based automation framework to capture and analyze warm leads from GetMyBoat's owner inbox. The work involved building a persistent browser session layer, SPA navigation patterns, credential prefilling, and structured data extraction—all designed to work around dynamic web application challenges.
The session created seven Python modules across /tmp/:
gmb_session.py— Browser lifecycle and persistent profile managementgmb_login.py— Credential handling and authentication flowgmb_explore.py— Single Page Application (SPA) navigation patternsgmb_inbox.py— Owner inbox URL discovery and message capturegmb_watch.py— Real-time URL observation during user navigationgmb_lead_scan.py— Read-only lead extraction and filteringgmb_manual.py— Human-in-the-loop prefilled credential flow
Technical Architecture
Browser Session Persistence
Rather than authenticate fresh on each run, we created a persistent Chromium user data directory. This approach:
- Stores authentication cookies and localStorage tokens across invocations
- Eliminates repeated 2FA challenges and email verification loops
- Allows the browser context to "remember" the owner's login state
- Reduces runtime from ~90 seconds (fresh auth) to ~15 seconds (cached session)
The profile directory is initialized once with gmb_login.py, then reused by all downstream scripts via context arguments passed to Playwright's launch configuration.
SPA Navigation and Waits
GetMyBoat's frontend is a React/Vue-style single-page application. Static waits fail; instead, we implemented:
- DOM-based waits:
page.wait_for_selector()targeting owner-specific inbox container classes - Network-based waits:
page.wait_for_load_state('networkidle')after navigation clicks - URL pattern observation: Watching for the actual inbox route (not the initial landing page URL) using
page.wait_for_url()with regex patterns
This three-layer approach ensures the DOM, network, and routing state are all consistent before attempting to scrape message content.
Credential Prefilling Flow
The gmb_manual.py script demonstrates the recommended production pattern:
1. Launch Playwright with persistent user data directory
2. Navigate to GetMyBoat login page
3. Prefill email input with carole@sailjada.com
4. Prefill password input with masked value
5. Pause and wait for human to review and submit (or paste 2FA code)
6. Capture the post-login landing URL
7. Store that URL for subsequent automated navigation
This hybrid approach balances automation (no manual typing per run) with security (password never logged or transmitted in scripts) and reliability (2FA handled by human when needed).
Infrastructure and Dependencies
Python Environment
Created a dedicated virtual environment with pinned versions:
playwright(async-capable version for Chromium control)google-api-client(existing; from prior Gmail work)- Python 3.10+ (required for Playwright's async patterns)
Playwright's browser binaries (Chromium) are downloaded into the venv's site-packages; no global system browser required. This isolates GetMyBoat scraping from other projects' browser versions.
GetMyBoat API Surface
Unlike a true REST API, GetMyBoat serves data via:
- Frontend rendering: React components emit message metadata into the DOM as data attributes
- XHR requests: Inbox list and detail requests fire to internal API endpoints (captured via network observation)
- Form submissions: Reply composition uses standard form POST
Our scraper targets the rendered DOM, avoiding the need to reverse-engineer XHR headers or token refresh cycles.
Key Design Decisions
Why Playwright over Selenium?
Playwright offers:
- Native async/await support (cleaner code for coordinating multiple waits)
- Better SPA handling via
page.wait_for_load_state() - Simpler context isolation (each script gets its own browser context, sharing a profile)
- Built-in network interception for debugging auth flows
Selenium would require custom wait logic and synchronous threading to achieve the same effect.
Why Persistent Profiles?
GetMyBoat's authentication uses secure HTTP-only cookies. A fresh browser profile requires:
- Email/password entry
- Optional 2FA verification
- Waiting for session initialization XHR calls to complete
With a persistent profile, we skip steps 1–3 and reuse the existing session token, reducing latency and request volume.
Why DOM Scraping over API Reverse-Engineering?
GetMyBoat does not publish a public API for third-party lead access. We could:
- Option A (chosen): Scrape the rendered DOM after the SPA loads and user authenticates.
- Option B: Reverse-engineer internal XHR calls and replicate authentication tokens.
Option A is maintainable: if GetMyBoat's internal API changes, the DOM structure typically stays stable. Option B breaks the moment they refactor their frontend bundler or change API versioning.
Current Limitations and Next Steps
Authentication Timeout
The initial gmb_login.py run timed out when attempting to authenticate with carole@sailjada.com. This suggests:
- Possible rate-limiting on the GetMyBoat login endpoint
- Network latency or intermittent 503 errors on their auth service
- Email address or password mismatch (less likely, but worth verifying with account owner)
Next attempt should use gmb_manual.py with human verification of credentials and 2FA codes.
Lead Data Extraction
Once authentication succeeds, gmb_lead_scan.py
- Messages from GetMyBoat platform notifications (not direct user inquiries)
- Warm leads (owner viewed the listing, sent an inquiry, or bookmarked)
- Exclude cold outreach, spam, or messages older than 30 days
The extraction logic lives in gmb_inbox.py's