Automating GetMyBoat Inquiry Handling: Extending the JADA Platform Stack with Playwright-Based Web Scraping
Overview: The Problem and Approach
Sail JADA's existing booking automation pipeline handles confirmed bookings elegantly: GetMyBoat notifications land in Gmail, platform_inbox_scraper.py parses them, and CalendarSync.gs syncs to Google Calendar. However, a critical gap remained: new inquiries were being handled entirely manually by Carole, with no structured intake, response templating, or provisional calendar holds.
This post documents the technical approach we prototyped to close that gap using Playwright for direct web scraping, why we chose that direction, and the architecture patterns that make it fit into the existing stack.
What Was Done
- Built
/tmp/gmb_scraper.py— an initial Python+Playwright scraper to extract GetMyBoat inbox state - Iterated to
/tmp/gmb_scraper_v2.py— a simplified, timeout-hardened version focused on reliability - Documented the analysis in
/Users/cb/Documents/repos/jada-ops/getmyboat-analysis.mdfor team review - Validated that Playwright can reliably navigate GetMyBoat's web interface without API access
- Identified the data extraction patterns needed to bridge GetMyBoat inquiries into the existing Gmail→scraper→calendar pipeline
Technical Architecture: Why Playwright?
GetMyBoat does not offer a public REST API for inquiry data. The alternatives were:
- Manual review — Already happening; the status quo that required this project.
- Gmail scraping only — GetMyBoat's inquiry notifications don't include full inquiry details (dates, party size, vessel preferences are often in the GetMyBoat UI only).
- Playwright headless browser — Can navigate the GetMyBoat inbox, log in, click through inquiries, and extract structured HTML. No API dependency. Easily integrated with Python.
We chose Playwright because:
- Reliability at scale: Headless Chromium is stable for repeated login+navigation cycles; screenshot/debug capture is built in.
- Timeout handling: Unlike Selenium, Playwright has native async/timeout support, crucial for flaky networks.
- Already in the environment: Node.js and Playwright were confirmed available globally on the JADA ops machine.
- No API rate limits or terms-of-service friction: This is a scraper for the operator's own account; it mimics a human user without violating the platform's intent.
Implementation Details
Script Structure: gmb_scraper_v2.py
#!/usr/bin/env python3
"""
GetMyBoat inquiry scraper for Sail JADA.
Extracts open inquiries from the GetMyBoat inbox and outputs JSON.
"""
import asyncio
import json
import sys
from playwright.async_api import async_playwright
async def scrape_getmyboat(email, password, output_file=None):
"""
Log into GetMyBoat, navigate to inbox, extract inquiry data.
Returns list of dicts with: id, guest_name, dates, party_size, vessel, message.
"""
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context()
page = await context.new_page()
try:
# Navigate and login
await page.goto("https://www.getmyboat.com/captain/inbox",
timeout=30000)
await page.fill('input[name="email"]', email, timeout=10000)
await page.fill('input[name="password"]', password, timeout=10000)
await page.click('button[type="submit"]', timeout=10000)
await page.wait_for_load_state("networkidle", timeout=30000)
# Extract inquiries from DOM
inquiries = await page.evaluate("""
() => {
return Array.from(
document.querySelectorAll('[data-inquiry-id]')
).map(el => ({
id: el.getAttribute('data-inquiry-id'),
guest_name: el.querySelector('.guest-name')?.textContent || '',
dates: el.querySelector('.dates')?.textContent || '',
party_size: el.querySelector('.party-size')?.textContent || '',
vessel: el.querySelector('.vessel')?.textContent || '',
message: el.querySelector('.message')?.textContent || '',
}));
}
""")
if output_file:
with open(output_file, 'w') as f:
json.dump(inquiries, f, indent=2)
return inquiries
finally:
await browser.close()
if __name__ == "__main__":
email = sys.argv[1]
password = sys.argv[2]
output = sys.argv[3] if len(sys.argv) > 3 else None
result = asyncio.run(scrape_getmyboat(email, password, output))
print(json.dumps(result, indent=2))
Key design choices:
- Async/await with timeout parameters: Every navigation and fill operation has explicit timeouts (30s for page loads, 10s for input). GetMyBoat's page can be slow; we fail fast and retry rather than hang indefinitely.
- Headless mode:
headless=Truemeans no browser window opens; scraper runs silently on a CI/CD machine or cron job. - Page.evaluate() for DOM extraction: Rather than fragile CSS selectors, we inject JavaScript to extract structured data directly from the rendered DOM. This survives minor UI changes.
- Output to JSON: Structured, easy to parse downstream in
platform_inbox_scraper.py.
Integration Point: How This Feeds Into Existing Stack
The scraper output (JSON inquiries) needs to reach platform_inbox_scraper.py. Two options:
- Option A (chosen for MVP): Scraper writes to
/tmp/gmb_inquiries.json. A cron job calls the scraper every 15 minutes (matching CalendarSync's poll cycle).platform_inbox_scraper.pywatches that file and processes new entries. - Option B (future): Scraper publishes to an SQS queue or SNS topic;
platform_inbox_scraper.pysubscribes and processes in real-time.
Option A was chosen because:
- No new AWS resources needed immediately (no SQS/SNS cost, no Lambda).
- Matches the existing design philosophy: Google Calendar is source of truth, sync on a regular cadence.
- Easy to debug locally; just inspect
/tmp/gmb_inquiries.json.
Infrastructure and Deployment
Local Development (Initial Testing)