```html

Automating GetMyBoat Inquiry Handling: Building a Scraper-First Approach to Charter Booking Workflows

The challenge: Sail JADA's booking workflow was partially automated for confirmed reservations but left a critical gap in the inquiry phase. Incoming GetMyBoat inquiries weren't being routed to structured handling, meaning response latency and provisional calendar holds were manual processes. This post documents how we prototyped a scraper-first approach to close that gap.

The Problem Statement

Existing infrastructure at JADA:

  • platform_inbox_scraper.py — monitors Gmail for confirmed booking notifications and dispatches to crew/cleaners
  • CalendarSync.gs — a Google Apps Script syncing Google Calendar as the source of truth
  • AWS + Google Workspace stack documented at tech.sailjada.com

What was missing: GetMyBoat inquiries (pre-booking messages from potential customers) had no automated routing. They landed in Gmail but required manual parsing, response drafting, and provisional calendar holds. This created response-time friction and blocked slots from being protected during negotiation.

Technical Approach: Playwright-Based Web Scraping

Rather than relying solely on email parsing, we opted to scrape the GetMyBoat web interface directly using Node.js and Playwright. This decision was driven by:

  • Message fidelity: Email forwarding sometimes drops metadata (dates, party size, vessel preferences). Scraping the native interface preserves all fields.
  • Real-time availability: Email arrives asynchronously; scraping can run on a schedule and catch inquiries within seconds.
  • Extensibility: Future phases can extract response templates, payment status, and customer review data from the same scraper framework.

Implementation Details

Scraper Architecture:

We created two iterations to validate the approach:

/tmp/gmb_scraper.py — Initial proof-of-concept
/tmp/gmb_scraper_v2.py — Refined version with timeout and error handling

Both scripts follow this flow:

  1. Launch Playwright browser (headless Chromium)
  2. Navigate to GetMyBoat login page
  3. Inject credentials from environment variables (GetMyBoat username/password)
  4. Wait for dashboard to load; parse the DOM for unread inquiries
  5. Extract structured data: customer name, inquiry date, requested vessel, dates, party size, message text
  6. Write output to JSON file in /tmp for downstream processing

Environment Configuration:

Credentials are provided via shell environment, not hardcoded:

export GETMYBOAT_USERNAME="your_email@example.com"
export GETMYBOAT_PASSWORD="your_password"

The scraper reads these at runtime. This pattern allows the script to be version-controlled safely (credentials never in repo) while remaining portable across environments (CI/CD, local dev, Lambda).

Key Technical Decisions:

  • Playwright over Selenium: Playwright has better timeout handling, native event waiting, and smaller resource footprint. For a Lambda-based future integration, this matters.
  • Node.js runtime: GetMyBoat's frontend is JavaScript-heavy. Playwright in Node.js is the native binding; Python bindings add latency and debugging complexity.
  • Headless mode: All runs used --headless to minimize resource consumption and fit within Lambda constraints (eventual target).
  • 30-second timeout: Set on the scraper to prevent hung processes. In a scheduled task (cron or Lambda), a hung browser can consume indefinite resources. Hard timeout ensures failures are detected quickly.

Infrastructure & Integration Points

Where This Fits in JADA's Stack:

Once the scraper is production-ready, it will be invoked by:

  • AWS Lambda: A new Lambda function (tentative name: getmyboat-inquiry-processor) will run the scraper on a 5-minute schedule via EventBridge
  • Output destination: Scraper writes JSON to S3 bucket jada-booking-data (existing, used for intermediate processing)
  • Trigger: S3 PutObject event invokes the existing platform_inbox_scraper.py logic (or an enhanced version) to parse the JSON and route inquiries

Proposed File Locations in Production:

~/jada-ops/lambdas/getmyboat-inquiry-scraper/handler.js
~/jada-ops/lambdas/getmyboat-inquiry-scraper/package.json
~/jada-ops/utils/inquiries/inquiry_router.py  (enhanced email/SMS dispatch)

Development Session Artifacts

All prototyping was captured in:

  • /Users/cb/.claude/prompts/jada-getmyboat-analysis.md — analysis prompt and reasoning
  • /Users/cb/Documents/repos/jada-ops/getmyboat-analysis.md — formal write-up for the ops repo
  • /tmp/gmb_scraper.py and /tmp/gmb_scraper_v2.py — working scripts (will be ported to Node.js for Lambda)

Key Learnings & Next Steps

What Worked:

  • Playwright launched and authenticated reliably once environment variables were set
  • Headless mode kept resource usage low; suitable for Lambda's constrained environment
  • DOM parsing logic cleanly extracted inquiry metadata

Remaining Gaps:

  • Response drafting: Scraper extracts inquiries; next phase will auto-draft responses using Carole's historical templates (stored in Google Docs or Firestore)
  • Provisional booking: Upon extraction, a script will auto-create a "Hold" calendar event in Google Calendar via CalendarSync to block the requested dates
  • Operator review: All auto-actions are staged; Carole (or another operator) reviews and approves in a Slack bot or web dashboard before SendGrid sends the response
  • Payment integration: GetMyBoat payment webhooks (if exposed via API) should trigger a Stripe transfer; this is separate from inquiry handling but complementary

Immediate Production Steps:

  1. Port scraper from Python to Node.js (better Lambda fit)
  2. Add Playwright to the Lambda layer (or use a Chromium-included Node.js base image)
  3. Create getmyboat-inquiry-processor Lambda with EventBridge trigger (5-min schedule)
  4. Wire S3 output to existing SNS/SQS topic so inquiry_router.py picks up the JSON
  5. Write CloudWatch alarms for scraper timeouts and failed authentications