```html

Building a Three-Layer Deterministic Test System for a Booking Pipeline

What Was Done

Rebuilt the automated test infrastructure for a charter booking system that had silently degraded over months. The nightly test suite (dead since June 12), payment monitoring (hung for 4.5 days), SMS alerting, and user cron jobs (dead since ~April 10) were all non-functional. The rebuild establishes three independent verification layers, each watching the one below:

  • Nightly suite: 279 deterministic tests running at 07:30 via launchd, checking every load-bearing system
  • Immediate SMS alerts: Text on any failure, not just dashboard review
  • Cloud watchdog: Lambda function running every 6 hours, independent of local infrastructure

This enforces a single definition of "done": an automated, deterministic test on cron that keeps the thing fixed.

The Discovery: Eight Silent Failures

The existing test runner had stopped producing results without raising alarms. Root causes:

  • jada-tests-nightly.py and magic-link heartbeat: macOS TCC (Transparency, Consent, and Control) blocks /bin/bash from accessing iCloud under launchd, silently failing without error output
  • Payment monitor: urllib.request.urlopen() called without a timeout parameter; one stalled TCP connection from a flaky API hung the entire process for 4.5 days
  • sms-flush (sends SMS via AWS): crashed every 2 minutes due to expired default AWS credentials; should have been using AWS_PROFILE=queenofsandiego
  • Rebuild-crew-pages: missing aws CLI in PATH due to shell initialization not being invoked by launchd
  • User crontab: dead since ~April 10 (macOS upgrade or daemon reset)

Each was verified fixed by triggering a real run and checking logs.

Technical Details: The Three-Layer Architecture

Layer 1: Nightly Suite (Local, Fast Feedback)

Runs every morning at 07:30 via /Users/cb/Library/LaunchAgents/com.jada.tests-nightly.plist. The LaunchAgent points to a Python wrapper script, /Users/cb/bin/jada-tests-nightly.py, which:

  • Uses Homebrew's python3 (bypasses TCC issues)
  • Explicitly sets AWS_PROFILE=queenofsandiego
  • Sets PATH to include /usr/local/bin for AWS CLI and other tools
  • Runs 279 tests across 11 test files in /Users/cb/icloud-repos/sites/queenofsandiego.com/tests/
  • Publishes a heartbeat to S3 with timestamp, test count, and pass/fail status
  • On failure, sends SMS to alerting number with test count and first failure name

New test files added this session:

  • test_launchd_jobs.py: Verifies every scheduled job (8 total) is loaded in launchd, exits with 0, and has a fresh logfile. A broken cron will turn the dashboard red the next morning.
  • test_booking_pipeline.py: Checks Stripe key is in live mode and authenticates, charter product exists in Stripe, SES senders verified, DynamoDB tables ACTIVE, schemas for ledger/contacts/suppression exist, waiver template has ≥3 columns and ≥30 rows, manifest padding logic, all 9 booking scripts compile without syntax errors
  • test_crew_pages_gate.py: Crew templates load, magic-link authentication works end-to-end against live infrastructure

Test results populate a dashboard at https://queenofsandiego.com/g/tests/dashboard.html. All 279 tests ran clean on the first nightly pass after fixes.

Layer 2: Immediate SMS Alerting

The wrapper script sends SMS to a configured alert number on any failure. You received a real alert at 02:08 when the initial test run found 14 failures out of 207 tests. After fixes, the 07:30 run passed all 279. This inverts the alerting model from "you check the dashboard" to "the system texts you."

Layer 3: Cloud Watchdog (Lambda + EventBridge)

A Lambda function jada-tests-watchdog runs every 6 hours via EventBridge, triggered by a cron expression. It:

  • Fetches the heartbeat file from S3 (updated by the nightly suite and other jobs)
  • Checks timestamp freshness (alerts if heartbeat is >6 hours old)
  • Makes HTTPS requests to all 4 booking-related domains and checks HTTP status
  • Emails jadasailing@gmail.com with status on failure
  • Runs independently of the local machine, surviving the Mac being dead or unreachable

Verified live by invoking the Lambda manually and confirming email delivery.

Infrastructure Changes

  • LaunchAgents: Updated 4 plists to use full Python path, set AWS_PROFILE, and prepend /usr/local/bin to PATH: com.jada.tests-nightly.plist, com.jada.magic-link-heartbeat.plist, com.jada.rebuild-crew-pages.plist, com.jada.unsubscribe-monitor.plist
  • Lambda: Created function code at /Users/cb/icloud-repos/sites/queenofsandiego.com/tools/jada-tests-watchdog/lambda_function.py (Python 3.11)
  • EventBridge: Rule scheduled to invoke watchdog every 6 hours
  • S3: Heartbeat file published to a path accessible by both the local suite and Lambda
  • Socket timeout: Payment monitor patched to set a global socket timeout in addition to request timeouts, preventing indefinite hangs

Key Decisions

Why three layers, not one? The local nightly suite is fast (5–10 minutes) and catches issues first, but it depends on a single machine staying healthy. The Lambda watchdog costs almost nothing (EventBridge + 6h cadence) and survives infrastructure failure, but runs less frequently and can't run the full test suite in 30 seconds. Together: fast local feedback + slow cloud safety net.

Why SMS, not just email or dashboard? Email is async; dashboard requires active monitoring. A text is immediate and creates accountability. You got the first real alert at 02:08 because a test failed; the system woke you up instead of letting it rot until morning.

Why test the launchd jobs themselves? The root cause of most failures was invisible cron death. By treating "launchd job is loaded, exit 0, log-fresh" as a first-class test, any future cron breakage is detected within 24 hours, not months.

Why explicit AWS_PROFILE in the LaunchAgent? Launchd does not inherit a shell's environment initialization (`.zshrc`, `.bashrc` are not sourced). Hardcoding the profile avoids credential fallback chains that fail silently.

Standing Rule for Future Jobs

Any new scheduled task must:

  • Have an entry script in ~/bin (or ~/icloud-repos/tools for shared logic)
  • Use Homebrew python3 (not system Python)
  • Have a LaunchAgent plist in ~/Library/LaunchAgents that explicitly sets AWS_PROFILE and PATH
  • Have a row in test_launchd_jobs.py that verifies it is loaded, exits 0, and has a fresh logfile

This is enforced before the job is considered "done."

What's Next

A private guest page was discovered missing noindex (security risk; should not be crawlable). Fixed locally; awaiting deploy approval. Full follow-up list in /Users/cb/icloud-jada-ops/HANDOFF-2026-07-03-test-system.md.

The 07:30 nightly pass is now the first fully-armed scheduled check. If anything breaks, your phone texts you.

```