Automating GetMyBoat Inquiry Handling: Extending the JADA Booking Pipeline with Playwright Web Scraping
What Was Done
We extended the existing JADA booking automation infrastructure to handle GetMyBoat inquiries—the gap between new customer interest and confirmed calendar entries. Previously, incoming inquiry emails were processed manually by crew leadership. We built a new scraper capability using Node.js and Playwright to directly extract structured inquiry data from GetMyBoat's web interface, bridging the inbox monitoring layer (platform_inbox_scraper.py) with the calendar sync system (CalendarSync.gs).
Two Python prototypes were developed and tested in the session:
/tmp/gmb_scraper.py— Initial Selenium-based scraper with full browser automation/tmp/gmb_scraper_v2.py— Playwright-based scraper with optimized timeouts and error handling
Both scripts target the same objective: authenticate to GetMyBoat, navigate to the inbox, parse inquiry metadata (customer name, requested dates, party size, vessel type, message body), and output structured JSON for downstream processing.
Technical Details: Scraper Architecture
Why Playwright over Selenium?
GetMyBoat's UI is JavaScript-heavy with asynchronous content loading. Selenium's synchronous model introduces flakiness; Playwright's native async/await pattern and built-in network idle waiting reduce false timeouts. For low-frequency inquiry polling (every 15–30 minutes), startup overhead is negligible; reliability matters more.
Authentication Flow
The scraper authenticates using GetMyBoat credentials stored as environment variables:
GMB_EMAIL=<operator-email>
GMB_PASSWORD=<operator-password>
On launch, the script:
- Creates a Playwright browser context with persistent storage
- Navigates to
https://www.getmyboat.com/owner/inbox - Detects login form and submits credentials programmatically
- Waits for navigation to complete and DOM to stabilize (30-second timeout default)
- Extracts inquiry cards from the inbox list
Data Extraction Pattern
Once authenticated, the scraper targets inquiry cards via CSS selectors:
.inquiry-card, [data-testid="inquiry-item"], .message-thread
From each card, it extracts:
customer_name— Displayed sender nameinquiry_date— When the message arrivedrequested_dates— Check-in and check-out from inquiry text or booking formparty_size— Number of guestsvessel_requested— Boat name or typemessage_body— Full inquiry textinquiry_id— GetMyBoat's internal ID for tracking
Output is serialized to JSON and written to /tmp/gmb_inquiries_<timestamp>.json.
Infrastructure Integration Points
Credential Management
Credentials are stored in a .env file (excluded from version control via .gitignore) and loaded at runtime by the Python script using the python-dotenv library. This avoids hardcoding secrets in repository or CloudFormation templates. For production deployment on AWS Lambda or EC2, credentials would be stored in AWS Secrets Manager with IAM policy enforcement.
Existing Platform Integration
The scraper output feeds into two existing systems:
platform_inbox_scraper.py— Already monitors Gmail for booking notifications. A new conditional branch will check if the Gmail message is a GetMyBoat inquiry (not a confirmed booking) and route it to a queue for response drafting and calendar hold creation.CalendarSync.gs— The Google Apps Script running on a 15-minute cycle that syncs Google Calendar as the source of truth. When an inquiry is received, a new "Provisional Hold" calendar entry (with a distinct color, e.g., yellow vs. confirmed green) would be created for the requested dates, preventing double-booking while negotiation is ongoing.
Output Staging
Scraped inquiry JSON is written to a local directory (default: /tmp/gmb_inquiries/) with automatic directory creation. In production, this would be replaced with an S3 bucket configured specifically for inquiry data:
s3://jada-getmyboat-inquiries/raw/<YYYY>/<MM>/<DD>/inquiries_<timestamp>.json
An S3 event notification (via SNS or Lambda trigger) would automatically invoke downstream processing Lambda functions when new inquiry files arrive.
Key Decisions & Trade-offs
Polling vs. Webhooks
GetMyBoat does not expose inquiry webhooks in their public API. Polling is the only viable option. A 15–30 minute cadence is cost-effective (minimal Lambda invocations) and acceptable for inquiry response SLAs (same-day turnaround). If real-time response is required, this could be upgraded to 5-minute polling with conditional cost optimization (only scrape if new messages are detected via a cheaper metadata API).
Browser Automation vs. REST API
GetMyBoat's public API lacks inquiry endpoints. Web scraping is the only way to extract unread inquiry data. The trade-off: scraping is slower and more fragile than API calls (UI changes break selectors), but it's the only current option. A formal API partnership or webhook integration would eliminate this dependency.
Timeout Tuning
Initial versions used aggressive timeouts (5–10 seconds), leading to timeout errors when GetMyBoat's servers were slow. Version 2 raised the timeout to 30 seconds and added explicit network idle waits. The balance: faster detection (30 seconds per poll) vs. reliability (no false failures).
What's Next
The scraper is functional but prototype-stage. Next steps:
- Response Drafting — Auto-generate inquiry response templates using operator preferences (availability rules, pricing, vessel descriptions) and send via GetMyBoat messaging or email.
- Provisional Calendar Hold — Automatically create tentative Google Calendar entries for inquired dates, preventing crew from overcommitting.
- Payment Routing — When an inquiry converts to a booking, extract payment details and route to QuickBooks or the operator's accounting system.
- AWS Lambda Deployment — Move scraper from local/manual to Lambda with CloudWatch scheduled triggers (every 15 minutes) and error alerting via SNS.
- Monitoring & Observability — Add CloudWatch Logs and X-Ray tracing to detect scraper failures (GetMyBoat UI changes, auth failures, network issues).
Once this pipeline is live, JADA's entire inquiry-to-booking workflow will be automated, eliminating manual