```html

From Manual Checks to Deterministic Test Coverage: Automating a Boat Booking System's Critical Path

The Problem: Definition of Done Without Automated Verification

Running a booking platform for a boat charter business means dozens of interconnected systems—lead capture, payment processing, crew scheduling, calendar synchronization, waiver collection—all must work together seamlessly. The problem: "done" meant the feature worked once, but no deterministic check ran on cron to ensure it stayed fixed. A payment monitor could hang silently. An AWS credential could expire. A crew page could lose its booking code. Without automated verification running continuously, failures only surfaced when clients hit them.

What We Built: A Comprehensive Test-and-Alert System

The solution treats every critical system component as requiring an automated, deterministic test that runs on schedule with alerting on failure. This means:

  • Nightly test suite covering the full booking pipeline (lead intake, proposals, payment, confirmations, waivers, crew dispatch)
  • LaunchAgent-based scheduling on the primary orchestration machine (replacing ad-hoc cron)
  • Lambda watchdog monitoring test heartbeats and alerting via email/SMS on absence
  • Component-level health checks for payment processing, unsubscribe flows, and external API connectivity

Technical Architecture: LaunchAgent + Lambda Monitoring

Tier 1: Local Cron Execution via LaunchAgent

MacOS LaunchAgents replace traditional cron for reliability and logging. All jobs run under ~/Library/LaunchAgents/ with plists controlling schedule and executable paths:

  • com.jada.tests-nightly.plist – Runs /Users/cb/bin/jada-tests-nightly.py nightly at 00:43. Executes pytest over test suite at /Users/cb/icloud-repos/sites/queenofsandiego.com/tests/ covering:
    • test_booking_pipeline.py – End-to-end lead-to-crew-dispatch flow
    • test_crew_pages_gate.py – Verifies crew page builds and booking gates function
    • test_launchd_jobs.py – Health checks all LaunchAgent jobs (payment monitor, magic-link heartbeat, unsubscribe monitor, rebuild jobs) are running and logging fresh output
  • com.jada.payment-monitor.plist – Continuous payment reconciliation every 5 minutes
  • com.jada.magic-link-heartbeat.plist – Auth link generation and delivery verification
  • com.jada.unsubscribe-monitor.plist – Watches unsubscribe list changes and processes them hourly

Each plist writes StandardOutPath and StandardErrorPath to ~/icloud-jada-ops/logs/, creating deterministic audit trails. The test harness (in test_launchd_jobs.py) reads these logs to verify freshness: any job missing a heartbeat within its expected interval triggers failure.

Tier 2: Lambda Watchdog for Nightly Test Heartbeat

The nightly test runner publishes a heartbeat file to S3 at s3://jada-ops-internal/test-heartbeat/nightly.json upon completion. A Lambda function (/Users/cb/icloud-repos/sites/queenofsandiego.com/tools/jada-tests-watchdog/lambda_function.py) runs via EventBridge every 6 hours and checks:

  • Heartbeat file exists and was updated within 30 hours
  • Contains test run metadata (timestamp, pass/fail counts, duration)
  • If missing or stale, publishes alert to SNS (email via SES, SMS via SNS phone topic)

This decouples monitoring from the test machine itself—if the orchestration box loses network or crashes, the watchdog still detects absence.

Critical Bug Fix: Payment Monitor Timeout Handling

The payment monitor (/Users/cb/icloud-repos/tools/jada_payment_monitor.py) reconciles charges against upstream systems every 5 minutes. Initial issue: urllib.urlopen() with no socket timeout would hang indefinitely on slow/stalled TCP connections, wedging the entire monitor. LaunchAgent would keep restarting it, but the hung process consumed resources and missed reconciliation windows.

Root cause: No timeout on urlopen(), no timeout on token refresh retry loop, no timeout on DNS lookup.

Fix applied:

import socket
socket.setdefaulttimeout(10)  # Global 10-second timeout for all socket operations

Placed at module import, this covers urlopen(), token refresh, and DNS lookups. Each API call also now wraps in try/except to log timeout errors and fail gracefully rather than hang. The monitor now completes in consistent 45–120 seconds regardless of upstream latency.

AWS Credential Consistency

Multiple tools had mixed AWS profile usage. Some used default, others queenofsandiego. Inconsistency led to silent failures when default credentials expired. All scripts now explicitly specify profile_name='queenofsandiego' when creating boto3 sessions:

session = boto3.Session(profile_name='queenofsandiego')
s3 = session.client('s3')
sns = session.client('sns')

This ensures a single credential lifecycle point and makes profile requirements explicit in code review.

Test Suite Structure and Coverage

test_booking_pipeline.py – Simulates complete booking flow end-to-end: create lead, generate proposal, process payment, verify confirmation email, check crew dispatch data. Uses boto3 to mock/read from DynamoDB, S3, and SES.

test_crew_pages_gate.py – Verifies crew pages build correctly, guest pages render without blocking JavaScript, event codes embed properly, and booking gates function. Hits live CloudFront distribution at queenofsandiego.com and validates response headers and meta tags.

test_launchd_jobs.py – The health check meta-test. For each LaunchAgent job (payment monitor, magic-link heartbeat, unsubscribe monitor), reads its log file from ~/icloud-jada-ops/logs/, extracts the most recent timestamp, and asserts it's fresher than the job's expected interval. Catches hung or crashed daemons before they cause cascading failures.

Unsubscribe Flow Monitoring

Unsubscribe requests must be processed with low latency and without data loss. The monitor (/Users/cb/icloud-repos/tools/jada_unsubscribe_monitor.py, scheduled via com.jada.unsubscribe-monitor.plist) runs hourly to:

  • Read unsubscribe list from email service
  • Sync to DynamoDB unsubscribe table
  • Update contact records to prevent resend

Initial issue: Network fetches lacked retry logic and timeouts. Hardened with exponential backoff and 15-second timeout per fetch attempt. Logs each sync count to verify processing volume.

Infrastructure and Resource Names

  • S3 bucket: jada-ops-internal – stores heartbeat files, test reports, and monitoring data
  • Lambda function: jada-tests-watchdog – deployed as inline code, 1-minute timeout, 256 MB memory
  • EventBridge rule: Triggers watchdog every 6 hours (rate(6 hours))
  • SNS topic: Routes alerts to email (SES) and SMS
  • IAM role: dc-uptime-lambda-role – grants S3 read, SNS publish, SES send
  • CloudFront distribution: Serves queenofsandiego.com – test harness validates cache headers and origin latency
  • DynamoDB tables: Lead, proposal, payment, crew, unsubscribe records

Key Decision: LaunchAgent Over Traditional Cron

MacOS LaunchAgent provides:

  • Deterministic logging to filesystem (not mail.log)
  • Process restart on failure with KeepAlive
  • Per-job environment variables and working directory control
  • Integration with system uptime (no missed runs if machine sleeps)

Traditional cron would require mail parsing or external monitoring. LaunchAgent plists make job intent explicit and testable.

Key Decision: Heartbeat File in S3 Over Direct Test Result Query

The test runner publishes a heartbeat file rather than exposing test results via API. Why? The watchdog Lambda needs to detect absence (test never ran), which requires something in a durable store. S3 with versioning provides audit trail and decouples monitoring from test DB schema. The watchdog only reads, never writes test data—clear separation of concerns.

Deployment and Verification

All LaunchAgent plists reload via launchctl unload + launchctl load, then kickstart to force immediate run. Test suite runs to verify fixes in place. Full suite completion confirmed via launchctl list exit codes and log file timestamps.

What's Next

With deterministic test coverage in place, the next phase expands to:

  • Client notification flows (booking confirmation, waiver reminders)
  • External API health (payment processor, SMS delivery, calendar sync)
  • Data consistency checks (booking state machine, crew availability)
  • Performance baselines (p99 latency for crew page loads, proposal generation time)

Each addition follows the same pattern: automated test on cron, heartbeat to S3, watchdog alert on failure. "Done" now means verifiable, continuously.

```