Automating GetMyBoat Lead Capture with Playwright: Building a Headless Browser Session Manager for SPA Navigation
Over the past development session, we tackled a critical infrastructure challenge: automating lead extraction from GetMyBoat's single-page application (SPA) without hitting authentication timeouts or losing session state. This post details the technical approach, the pitfalls we encountered, and the architecture we built to handle persistent browser sessions across navigation events.
The Problem: SPA Navigation and Session Fragility
GetMyBoat uses a modern SPA built on client-side routing. Traditional HTTP-based scraping fails here because:
- The inbox URL doesn't load until after login completes and the client-side router initializes
- Navigation between pages doesn't trigger full page reloads—only JavaScript state updates
- Session cookies are validated at login, but the inbox endpoint may redirect or require additional client-side handshakes
- Timing is critical: wait too short and the page hasn't rendered; wait too long and we timeout
Our initial approach—using Selenium or basic Playwright scripts with fixed waits—failed because we couldn't reliably detect when the inbox had loaded. We needed a persistent session manager that could observe navigation and react dynamically.
Architecture: Persistent Headed Browser Sessions
Rather than spinning up a fresh browser for each scrape, we implemented a persistent headed browser session architecture:
- Launch a persistent Chromium instance with a user data directory that survives script restarts
- Maintain session state across multiple script invocations by reusing the browser profile
- Listen for navigation events in real-time and capture the actual inbox URL as the SPA router changes it
- Decouple authentication from data extraction—login once, then query inbox repeatedly without re-authenticating
Implementation: Core Scripts
We created a modular suite of Python scripts in /tmp to isolate concerns:
/tmp/gmb_session.py— Session manager; handles browser lifecycle and profile persistence/tmp/gmb_login.py— Headless login orchestration; waits for post-login redirects and cookies/tmp/gmb_explore.py— SPA navigation explorer; clicks through nav elements to discover URLs/tmp/gmb_inbox.py— Inbox state extractor; queries inbox and formats lead data/tmp/gmb_watch.py— Navigation event listener; logs all page transitions in real-time/tmp/gmb_lead_scan.py— Read-only lead scanner; extracts inbox without authentication
Technical Details: The Persistent Profile Approach
The key innovation was using Chromium's --user-data-dir flag to maintain a persistent profile directory. Here's the pattern:
# Launch with persistent user data directory
browser = await chromium.launch(
headless=False,
args=[
f"--user-data-dir=/path/to/persistent/profile",
"--disable-blink-features=AutomationControlled",
"--no-first-run",
"--no-default-browser-check"
]
)
context = await browser.new_context()
page = await context.new_page()
# Login once
await page.goto("https://getmyboat.com/login")
await page.fill("input[type='email']", "carole@sailjada.com")
await page.fill("input[type='password']", "***")
await page.click("button[type='submit']")
await page.wait_for_url("**/dashboard**", timeout=30000)
# Profile now contains valid cookies and local storage
# Subsequent script runs can reuse this profile without re-authenticating
Why this matters: GetMyBoat's session validation happens at login time. Once authenticated, subsequent requests to the inbox endpoint carry the session cookie. By persisting the profile, we avoid re-authentication overhead and reduce the risk of triggering rate limits or CAPTCHAs.
Handling SPA Navigation: The Watch Pattern
SPAs don't trigger traditional page navigation events. Instead, we listen to the browser's request/response lifecycle and URL change events:
page.on("response", lambda response: handle_response(response))
page.on("framenavigated", lambda frame: log_navigation(frame))
# When user clicks "Inbox" in the SPA nav, the router changes location.href
# We intercept this and wait for the corresponding API call
await page.wait_for_url("**/inbox**", timeout=5000)
# At this point, the inbox HTML is rendered and ready to query
inbox_leads = await page.query_selector_all("div[data-lead-item]")
This approach is more reliable than waiting for specific DOM elements to appear, because network requests are deterministic—the browser will issue them before rendering.
Lead Extraction: From HTML to Structured Data
Once the inbox page stabilizes, we extract lead data using CSS selectors:
leads = []
for lead_element in inbox_elements:
lead = {
"id": lead_element.evaluate("el => el.dataset.leadId"),
"name": lead_element.query_selector("span.lead-name").text_content(),
"message": lead_element.query_selector("p.lead-message").text_content(),
"timestamp": lead_element.query_selector("time").get_attribute("datetime"),
"status": lead_element.query_selector("span.status").text_content()
}
leads.append(lead)
# Write to JSON for downstream processing
with open("/tmp/getmyboat_leads.json", "w") as f:
json.dump(leads, f, indent=2)
Infrastructure: Environment and Dependencies
We isolated the Playwright + Google API client stack into a dedicated Python virtual environment to avoid conflicts with the main Sail JADA application:
- venv location:
/Users/cb/Documents/repos/.venv-gmb - Python version: 3.11+ (required for async/await syntax)
- Key packages:
playwright>=1.40,google-auth>=2.0,google-api-python-client>=2.100 - Chromium binary: Downloaded and cached by Playwright during first run (~500 MB)
# Set up venv and install dependencies
python3.11 -m venv /Users/cb/Documents/repos/.venv-gmb
source /Users/cb/Documents/repos/.venv-gmb/bin/activate
pip install playwright google-auth google-api-python-client
# Download Playwright browsers (includes Chromium)
playwright install chromium
# Verify Playwright can launch Chromium
python -c "from playwright.sync_api import sync_playwright; p = sync_playwright().start(); print(p.chromium.launch())"
Key Decisions and Rationale
- Headed browser (not headless):