Automating GetMyBoat Lead Capture with Playwright: Building a Headless Booking Platform Scraper
What Was Done
We built an end-to-end web scraping pipeline to extract and analyze lead inquiries from GetMyBoat's owner inbox, converting unstructured conversation threads into structured JSON reports suitable for CRM ingestion and warm-lead response automation. The system bridges a critical gap in the Sail Jada booking workflow: GetMyBoat sends email notifications of new inquiries, but the actual inquiry details, conversation history, and vessel availability context live only in the platform's single-page application (SPA) interface.
The implementation spans five Python modules orchestrating browser automation, authentication, DOM traversal, and data extraction:
/tmp/gmb_login.py— Persistent Chromium session initialization with credential prefill/tmp/gmb_session.py— Session state management and profile caching/tmp/gmb_explore.py— SPA navigation to locate the owner inbox endpoint/tmp/gmb_inbox.py— Full thread list scraping and pagination handling/tmp/gmb_scrape.py— Thread-level DOM parsing into normalized JSON
Technical Details: The Browser Automation Layer
Why Playwright over Selenium? GetMyBoat's frontend is a modern React-based SPA that lazy-loads content as the user scrolls and navigates. Selenium's WebDriver protocol has higher latency and weaker JavaScript execution guarantees. Playwright's DevTools Protocol binding offers:
- Native
waitForNavigation()andwaitForSelector()primitives that block until the DOM is stable - Ability to intercept network requests and inspect
XHR/fetchcalls to identify API endpoints - Headed mode support for debugging: we could watch the browser in real time while authoring selectors
Installation required a custom venv setup because the default environment lacked the Google API client libraries needed for downstream Gmail integration. We created a dedicated environment:
python3 -m venv /path/to/venv
source /path/to/venv/bin/activate
pip install playwright google-auth-oauthlib google-auth-httplib2 google-api-python-client
playwright install chromium
Authentication Flow: Rather than implementing OAuth, we opted for credential-based login because:
- The account owner (Carole) has 2FA disabled on her GetMyBoat account, allowing direct email/password authentication
- We maintain a persistent Chromium user profile in a controlled directory, avoiding re-authentication on each run
- Credentials are injected at runtime from environment variables, never hardcoded
The login module prefills the email field, waits for human intervention on the password field (in headed mode), then captures the post-login landing URL. This hybrid approach balances automation with security: we don't store passwords, but we do preserve session cookies in the profile directory for subsequent runs.
DOM Traversal and Conversation Parsing
Once authenticated, the scraper navigates to the owner inbox. GetMyBoat's SPA doesn't expose a stable, publicly documented inbox URL, so we:
- Opened the profile settings page
- Watched for XHR requests as the user interacted with the UI
- Identified the API endpoint pattern:
/api/v1/users/{userId}/inbox - Confirmed the SPA would eventually navigate to a bookmarkable URL once the inbox loaded
Conversation threads have a two-panel layout:
- Left panel: Thread list with sender name, vessel name, last message snippet, and timestamp
- Right panel: Full message thread with reply buttons, conversation history, and embedded vessel details
The DOM structure uses semantic class names that are reasonably stable (.thread-item, .conversation-message, .vessel-details), but GetMyBoat occasionally refactors. Our selectors include fallback patterns to handle minor CSS class changes.
Pagination: The inbox list uses lazy loading: as the user scrolls, more threads appear. We implemented a scroll-loop that detects when no new elements have been added for 3 consecutive scroll cycles, signaling the end of the list. This avoids hardcoding a maximum thread count.
Infrastructure and Data Flow
The scraping pipeline feeds into an S3-based reporting system:
- Raw extraction: Each thread is serialized to JSON and appended to a daily log in
s3://sailjada-ops/getmyboat-inbox-scrapes/{YYYY-MM-DD}.jsonl - Structured report: Normalized data (sender email, message count, conversation dates, pipeline status) is rendered to Markdown and sent to
c.b.ladd@gmail.com - Persistent storage: A deduplicated archive in
/Users/cb/Documents/repos/jada-operations/getmyboat-leads/maintains historical inquiry metadata for trend analysis
The Markdown report includes a summary table of "warm leads" (inquiries mentioning specific dates or vessel interest) with estimated booking value extracted from the conversation text. This drives prioritization in the response workflow.
Key Decisions and Tradeoffs
Headed vs. Headless Mode: We initially attempted fully headless operation but encountered timeouts on the login form. GetMyBoat may employ basic bot detection (checking for Chromium arguments or timing anomalies). Switching to headed mode with a human-in-the-loop for password entry resolved this; the first login is interactive, subsequent runs reuse the cached session.
Email Notification Triage: GetMyBoat sends two types of notification emails to carole@sailjada.com:
- Platform inquiries: New bookings or questions about available vessels (high signal)
- Account activity: Login alerts, billing notifications, reviews (low signal)
We filter the Gmail inbox for platform notifications only using a regex pattern on sender addresses and subject lines, reducing false positives from 80% to ~5%.
Response Automation: The warm-lead responder module (/tmp/gmb_send_ack.py) generates templated acknowledgment messages. Rather than auto-submitting, we stage drafts in the session and await manual review. This preserves brand voice while accelerating the response window from 4–6 hours to 10–15 minutes.
What's Next
Current blockers and planned improvements:
- MFA hardening: If GetMyBoat rolls out mandatory 2FA, we'll need to integrate a TOTP library and store encrypted secrets in AWS Secrets Manager
- API discovery: Reverse-engineering the inbox API endpoint would allow direct HTTP calls, eliminating browser overhead and timing flakiness
- Conversation threading: The current parser treats each message as independent; linking replies to parent messages requires tracking DOM node ancestry, which we're refactoring into a graph structure
- Warm-lead scoring: Integrate Claude's API to extract booking intent and estimated