Automating GetMyBoat Lead Ingestion: Headless Browser Architecture for SPA Navigation and Session Persistence
When dealing with third-party platforms that expose business opportunities through web interfaces rather than APIs, automation becomes a game of reverse-engineering the SPA (Single Page Application) and maintaining authenticated sessions across ephemeral execution contexts. This post covers the technical approach we took to automate GetMyBoat inbox extraction—including the architecture decisions, tooling setup, and the obstacles that emerged during implementation.
The Problem: No Public API, Dynamic SPA Navigation
GetMyBoat operates as a fully client-rendered SPA. There is no public API for retrieving platform notifications or lead inquiries. The inbox lives behind authentication and requires:
- Email/password login via browser automation
- Session persistence across page navigations (the inbox URL is discovered dynamically)
- Ability to handle JavaScript-heavy rendering (hence, not suitable for simple HTTP requests)
- Reliable credential management and token lifecycle handling
The solution: Playwright with a persistent browser profile, headed mode for debugging, and incremental SPA exploration to isolate the inbox route.
Tooling & Environment Setup
Our first hurdle was ensuring the Python environment had the right dependencies. The development machine had multiple Python interpreters and virtual environments; we needed one with both playwright and google-api-client (for parallel Gmail analysis).
Commands executed:
# Locate Python with google-api-client already installed
python3 -m pip list | grep google
# Create isolated venv with both Playwright and Google libs
python3 -m venv /tmp/gmb_venv
source /tmp/gmb_venv/bin/activate
pip install playwright google-auth-oauthlib google-auth-httplib2 google-api-python-client
# Install browser binaries (critical—this is non-trivial)
playwright install chromium
# Verify Playwright can launch Chromium
python3 -c "from playwright.sync_api import sync_playwright; p = sync_playwright().start(); b = p.chromium.launch(); print('OK'); b.close(); p.stop()"
The Chromium download step is often overlooked but essential. Without it, playwright.chromium.launch() fails with cryptic path errors. We verified the binary was present and executable before proceeding.
Browser Profile Persistence & Headed Mode
A key architectural decision: use a persistent user data directory (browser profile) rather than logging in fresh each time. This trades off some isolation for speed and reliability.
Profile directory: /tmp/gmb_profile
Launch pattern:
from playwright.sync_api import sync_playwright
playwright = sync_playwright().start()
browser = playwright.chromium.launch(
headless=False, # Headed mode for debugging & manual navigation
user_data_dir="/tmp/gmb_profile",
args=["--disable-blink-features=AutomationControlled"]
)
context = browser.new_context()
page = context.new_page()
Headed mode (`headless=False`) allows you to watch the automation in real time, essential when debugging SPA routing and form interactions. The `--disable-blink-features=AutomationControlled` flag reduces detection-avoidance friction (though GetMyBoat's checks are relatively light).
The Login Flow & Session Handoff
We implemented a two-phase approach:
Phase 1: Initial login with credential input (file: /tmp/gmb_login.py)
- Navigate to
https://www.getmyboat.com/ - Locate login button/link via selector
- Fill email field with
carole@sailjada.com - Fill password field (read from environment variable or config, never hardcoded)
- Submit form and wait for redirect to dashboard
- Verify presence of user menu or authenticated-only element
Phase 2: Persistent session exploration (file: /tmp/gmb_explore.py)
- Reuse the same browser profile (already authenticated)
- Navigate the SPA to discover the inbox URL by inspecting navigation menus
- Extract lead/inquiry data once inbox is located
- Store discovered URLs in a session state file for future runs
This separation allows the expensive login step to run once (or on session expiry) while exploration and data extraction iterate quickly.
SPA Navigation & Dynamic URL Discovery
GetMyBoat's navigation is entirely JavaScript-driven. The inbox route is not a static URL; it's constructed after authentication based on user state. We located it by:
- Inspecting the DOM after login for navigation links (
page.query_selector_all('nav a')) - Filtering by text content (e.g., links containing "Inbox" or "Messages")
- Capturing navigation events via
page.on('framenavigated')to log URL changes - Storing the discovered URL in a state file (
/tmp/gmb_session.py) for reuse
Example state file format:
{
"authenticated": true,
"profile_path": "/tmp/gmb_profile",
"inbox_url": "https://www.getmyboat.com/user/messages?tab=inquiries",
"last_login": "2024-01-15T14:32:00Z",
"leads_last_scraped": "2024-01-15T14:15:00Z"
}
Lead Data Extraction & Downstream Analysis
Once the inbox URL was discovered, we built a lead scraper (file: /tmp/gmb_lead_scan.py) to:
- Query all inquiry cards in the inbox DOM
- Extract text fields: sender name, boat type, inquiry date, inquiry snippet
- Correlate against existing warm-lead tracking (integration with Lowe's invoice reconciliation & rent calculations in parallel sessions)
- Output JSON for downstream analysis by the warm-lead responder automation
The extraction was designed as read-only—no state changes on GetMyBoat itself, only local data collection.
Key Decision: Why Playwright Over Puppeteer or Selenium?
- Cross-browser support: Chromium, Firefox, WebKit with identical API
- Native async/await: Python's
asynciointegration (though we used sync for simplicity here) - Better SPA handling: Built-in frame and navigation event tracking
- Profile management: Simpler persistent session handling than Selenium
- Debugging: Headed mode and trace recording are first-class
What Went Wrong & Next Steps
During the last development session, the initial login attempt timed out. Possible causes:
- GetMyBoat's login form may have added