Building a Three-Layer Deterministic Test System for a Booking Pipeline
What Was Done
Rebuilt the automated test infrastructure for a charter booking system that had silently degraded over months. The nightly test suite (dead since June 12), payment monitoring (hung for 4.5 days), SMS alerting, and user cron jobs (dead since ~April 10) were all non-functional. The rebuild establishes three independent verification layers, each watching the one below:
- Nightly suite: 279 deterministic tests running at 07:30 via launchd, checking every load-bearing system
- Immediate SMS alerts: Text on any failure, not just dashboard review
- Cloud watchdog: Lambda function running every 6 hours, independent of local infrastructure
This enforces a single definition of "done": an automated, deterministic test on cron that keeps the thing fixed.
The Discovery: Eight Silent Failures
The existing test runner had stopped producing results without raising alarms. Root causes:
jada-tests-nightly.pyand magic-link heartbeat: macOS TCC (Transparency, Consent, and Control) blocks/bin/bashfrom accessing iCloud under launchd, silently failing without error output- Payment monitor:
urllib.request.urlopen()called without a timeout parameter; one stalled TCP connection from a flaky API hung the entire process for 4.5 days sms-flush(sends SMS via AWS): crashed every 2 minutes due to expired default AWS credentials; should have been usingAWS_PROFILE=queenofsandiego- Rebuild-crew-pages: missing
awsCLI in PATH due to shell initialization not being invoked by launchd - User crontab: dead since ~April 10 (macOS upgrade or daemon reset)
Each was verified fixed by triggering a real run and checking logs.
Technical Details: The Three-Layer Architecture
Layer 1: Nightly Suite (Local, Fast Feedback)
Runs every morning at 07:30 via /Users/cb/Library/LaunchAgents/com.jada.tests-nightly.plist. The LaunchAgent points to a Python wrapper script, /Users/cb/bin/jada-tests-nightly.py, which:
- Uses Homebrew's
python3(bypasses TCC issues) - Explicitly sets
AWS_PROFILE=queenofsandiego - Sets
PATHto include/usr/local/binfor AWS CLI and other tools - Runs 279 tests across 11 test files in
/Users/cb/icloud-repos/sites/queenofsandiego.com/tests/ - Publishes a heartbeat to S3 with timestamp, test count, and pass/fail status
- On failure, sends SMS to alerting number with test count and first failure name
New test files added this session:
test_launchd_jobs.py: Verifies every scheduled job (8 total) is loaded in launchd, exits with 0, and has a fresh logfile. A broken cron will turn the dashboard red the next morning.test_booking_pipeline.py: Checks Stripe key is in live mode and authenticates, charter product exists in Stripe, SES senders verified, DynamoDB tables ACTIVE, schemas for ledger/contacts/suppression exist, waiver template has ≥3 columns and ≥30 rows, manifest padding logic, all 9 booking scripts compile without syntax errorstest_crew_pages_gate.py: Crew templates load, magic-link authentication works end-to-end against live infrastructure
Test results populate a dashboard at https://queenofsandiego.com/g/tests/dashboard.html. All 279 tests ran clean on the first nightly pass after fixes.
Layer 2: Immediate SMS Alerting
The wrapper script sends SMS to a configured alert number on any failure. You received a real alert at 02:08 when the initial test run found 14 failures out of 207 tests. After fixes, the 07:30 run passed all 279. This inverts the alerting model from "you check the dashboard" to "the system texts you."
Layer 3: Cloud Watchdog (Lambda + EventBridge)
A Lambda function jada-tests-watchdog runs every 6 hours via EventBridge, triggered by a cron expression. It:
- Fetches the heartbeat file from S3 (updated by the nightly suite and other jobs)
- Checks timestamp freshness (alerts if heartbeat is >6 hours old)
- Makes HTTPS requests to all 4 booking-related domains and checks HTTP status
- Emails
jadasailing@gmail.comwith status on failure - Runs independently of the local machine, surviving the Mac being dead or unreachable
Verified live by invoking the Lambda manually and confirming email delivery.
Infrastructure Changes
- LaunchAgents: Updated 4 plists to use full Python path, set
AWS_PROFILE, and prepend/usr/local/binto PATH:com.jada.tests-nightly.plist,com.jada.magic-link-heartbeat.plist,com.jada.rebuild-crew-pages.plist,com.jada.unsubscribe-monitor.plist - Lambda: Created function code at
/Users/cb/icloud-repos/sites/queenofsandiego.com/tools/jada-tests-watchdog/lambda_function.py(Python 3.11) - EventBridge: Rule scheduled to invoke watchdog every 6 hours
- S3: Heartbeat file published to a path accessible by both the local suite and Lambda
- Socket timeout: Payment monitor patched to set a global socket timeout in addition to request timeouts, preventing indefinite hangs
Key Decisions
Why three layers, not one? The local nightly suite is fast (5–10 minutes) and catches issues first, but it depends on a single machine staying healthy. The Lambda watchdog costs almost nothing (EventBridge + 6h cadence) and survives infrastructure failure, but runs less frequently and can't run the full test suite in 30 seconds. Together: fast local feedback + slow cloud safety net.
Why SMS, not just email or dashboard? Email is async; dashboard requires active monitoring. A text is immediate and creates accountability. You got the first real alert at 02:08 because a test failed; the system woke you up instead of letting it rot until morning.
Why test the launchd jobs themselves? The root cause of most failures was invisible cron death. By treating "launchd job is loaded, exit 0, log-fresh" as a first-class test, any future cron breakage is detected within 24 hours, not months.
Why explicit AWS_PROFILE in the LaunchAgent? Launchd does not inherit a shell's environment initialization (`.zshrc`, `.bashrc` are not sourced). Hardcoding the profile avoids credential fallback chains that fail silently.
Standing Rule for Future Jobs
Any new scheduled task must:
- Have an entry script in
~/bin(or~/icloud-repos/toolsfor shared logic) - Use Homebrew
python3(not system Python) - Have a LaunchAgent plist in
~/Library/LaunchAgentsthat explicitly setsAWS_PROFILEandPATH - Have a row in
test_launchd_jobs.pythat verifies it is loaded, exits 0, and has a fresh logfile
This is enforced before the job is considered "done."
What's Next
A private guest page was discovered missing noindex (security risk; should not be crawlable). Fixed locally; awaiting deploy approval. Full follow-up list in /Users/cb/icloud-jada-ops/HANDOFF-2026-07-03-test-system.md.
The 07:30 nightly pass is now the first fully-armed scheduled check. If anything breaks, your phone texts you.
```