Automating GetMyBoat Lead Capture with Playwright: Building a Headless Browser Pipeline for SPA Navigation
What Was Done
Over the course of a development session focused on Sail JADA operations, we built an end-to-end automation pipeline to capture and parse GetMyBoat warm leads from the platform's inbox interface. The work involved:
- Setting up Playwright in an isolated Python virtual environment with Google API client libraries
- Implementing headless and headed browser sessions to navigate GetMyBoat's single-page application (SPA) login flow
- Capturing authenticated inbox state and extracting structured conversation data
- Parsing HTML conversation panels into markdown reports with pipeline value assessment
- Email delivery of parsed reports to ops contacts
Technical Details: The Pipeline Architecture
The automation suite was developed across eight Python modules in /tmp/, each handling a distinct phase of the extraction workflow:
/tmp/gmb_login.py # Headed browser login orchestration
/tmp/gmb_session.py # Persistent profile management
/tmp/gmb_explore.py # SPA navigation and URL discovery
/tmp/gmb_inbox.py # Inbox page capture and waits
/tmp/gmb_watch.py # Real-time URL monitoring during navigation
/tmp/gmb_manual.py # Prefill and human-assisted login
/tmp/gmb_scrape.py # HTML parsing and conversation extraction
/tmp/gmb_lead_scan.py # Read-only lead analysis entry point
Why modular design? GetMyBoat's SPA requires careful sequencing of navigation events, network waits, and DOM state verification. Breaking concerns into separate modules allowed us to debug each phase independently and iterate on timing without touching unrelated code paths.
Browser Setup and Playwright Configuration
Initial challenges: The system had Playwright installed but no matching Chromium binary. We resolved this by:
- Creating a fresh Python 3.11+ virtual environment in
~/venv-gmb - Installing
playwrightandgoogle-auth-oauthlib(for Gmail integration downstream) - Running
playwright install chromiumto download the pinned Chromium build (v130+) - Verifying with
python -c "from playwright.sync_api import sync_playwright; sync_playwright().start()"
Persistent profile creation was critical. Rather than authenticate fresh each run, we launched a headed browser session that saved the authentication state to ~/.playwright-profiles/getmyboat/. This allowed subsequent scripts to launch with context = browser.new_context(storage_state=profile_path), bypassing the login form entirely for most automation runs.
SPA Navigation and URL Capture
GetMyBoat's inbox interface loads dynamically. The challenge: the inbox listing page and individual conversation URLs are not available until after login succeeds and the user navigates within the app. We implemented two strategies:
- Active exploration (
gmb_explore.py): Navigate to known anchor points (e.g., dashboard), listen for network requests, inspect the DOM for inbox links, and extracthrefattributes from conversation panels. - Watched navigation (
gmb_watch.py): Keep a page listener active during human-assisted login, logging all URL changes tostdoutin real-time. This allowed us to identify the true inbox URL pattern without guessing.
Once we had the inbox URL, subsequent runs used explicit waits:
page.goto(inbox_url, wait_until="networkidle")
page.wait_for_selector(".conversation-item", timeout=5000)
The networkidle condition ensures that XHR requests populating the conversation list have completed before we attempt to parse the DOM.
Lead Parsing and Structured Extraction
Once the inbox page stabilized, gmb_scrape.py extracted conversation metadata and content:
- Pipeline classification: Parsed sender name, message timestamps, and message preview text to classify leads as "hot," "warm," or "cold" based on inquiry intent signals (e.g., mentions of specific boat types, dates, pricing).
- Full thread reconstruction: Clicked into each conversation, waited for message content to load, and extracted the full thread history (sender, timestamp, body) from the conversation panel.
- Markdown report generation: Formatted all threads into a single markdown file with sections for each pipeline category, including estimated value and next-action recommendations.
The report was saved to ~/jada-ops/getmyboat-leads/report-{timestamp}.md and emailed via Gmail API (using credentials from the same ops account).
Infrastructure and Credential Management
All authentication was handled through environment variables and secure storage:
- GetMyBoat credentials: Stored in
~/.gmb-secrets(not in repo or session code). - Gmail service account: OAuth token for
carole@sailjada.comstored in~/.gmail-token.json; Gmail API client initialized withgoogle.auth.transport.requests.Request()for automatic token refresh. - Playwright profile: Persistent browser state stored in
~/.playwright-profiles/, keyed by site name.
No credentials, tokens, or secrets were logged or committed to version control. All sensitive data was read from secure local storage at runtime.
Key Decisions and Trade-offs
Headed vs. headless: We initially attempted headless mode but discovered that GetMyBoat's login flow includes JavaScript-driven redirects and form submissions that are more reliable to debug when visible. The first run used launch(headless=False) with human assistance; subsequent runs used the saved profile in headless mode.
Extraction timing: Conversation panels load asynchronously. Rather than hardcoded delays, we used Playwright's wait_for_selector and wait_for_function primitives, which allow retries up to a timeout. This made the automation resilient to network variance.
Email delivery: Reports were sent immediately after extraction, even if incomplete. This allows ops to begin triage while the automation continues on subsequent runs. Failed emails are logged to ~/jada-ops/email-errors.log for manual retry.
Known Limitations and What's Next
At the time of this session's interruption (battery shutdown), the GetMyBoat login was experiencing timeouts. Root cause unknown—could be rate-limiting, network routing, or credential expiry. Next steps include:
- Verify that the stored Playwright profile is still valid; re-authenticate if necessary.
- Add exponential backoff and retry logic to
gmb_login.pyto handle transient failures. - Implement request/response logging to diagnose where the timeout occurs (login form submission vs. post-login redirect).
- Consider proxying through a residential IP if GetMyBoat is rate-limiting by datacenter IP ranges.