Building Deterministic Cron-Based Test Coverage for a Live Booking Pipeline
The Problem: "Done" Without Verification
When your system is live and processing real bookings for paying clients, "fixed" doesn't mean "checked in." It means "has an automated deterministic test running on cron that proves it stays fixed." This session rebuilt the entire health-check infrastructure for a multi-service booking pipeline (lead intake → payments → calendar sync → crew dispatch → live sites) to enforce that standard.
What Was Built
- A comprehensive test suite spanning launchd job health, crew page gates, and end-to-end booking pipeline flow
- A Lambda-based watchdog that polls the test system every 6 hours and alerts via SMS on failure
- Fixed three critical timeout and credential bugs that were silently breaking automated checks
- Integrated S3-based heartbeat signaling so distributed components can detect each other's health
Architecture: Layered Test Verification
Layer 1: LaunchAgent Job Health
The foundation is a test_launchd_jobs.py that verifies the five critical background jobs actually run and complete daily:
com.jada.tests-nightly— the test suite itselfcom.jada.unsubscribe-monitor— watches for email unsubscribe bounces and processes themcom.jada.rebuild-crew-pages— rebuilds static HTML crew pages from database statecom.jada.magic-link-heartbeat— validates magic-link auth token generation pipelinecom.jada.morning-digest— sends daily digest to booking coordinators
Each LaunchAgent plist in ~/Library/LaunchAgents/ writes to a log file in ~/icloud-jada-ops/logs/. The test reads the log's modification time, exit code from launchctl list, and validates recent output format. If any job hasn't run in 26 hours or exited non-zero, the test fails.
Layer 2: Crew Page Gate Verification
The crew pages (served from CloudFront, sourced from S3, proxied through Route53) must have:
- Valid HTML structure and event metadata (date, time, crew names)
<meta name="robots" content="noindex">— these are time-limited landing pages, not SEO- Google Analytics tag injection (for conversion tracking)
- No stale data from previous event bookings
test_crew_pages_gate.py fetches two live guest pages from the CloudFront distribution, validates the noindex tag, checks that GA tracking is present, and verifies the HTML contains this week's actual event data. This catches cases where the rebuild job crashes silently or AWS credentials rotate without being rotated in the automation.
Layer 3: End-to-End Booking Pipeline
test_booking_pipeline.py exercises the full flow:
- Creates a test lead via the intake form (validates form processor and database write)
- Retrieves the generated proposal PDF (validates proposal templating and S3 upload)
- Submits a payment via Stripe (validates payment processor and webhook routing)
- Verifies the booking appears in DynamoDB with correct confirmation status
- Checks that a confirmation email was queued (validates email service integration)
The test uses fixtures to create deterministic, idempotent test records with names like TEST-PIPELINE-{TIMESTAMP} so they don't collide with real bookings. Assertions check for exact data shapes: row padding, field ordering, event code presence — the kind of subtle bugs that slip through manual testing.
Bug Fixes: Why Tests Failed Initially
Payment Monitor Timeout
The jada_payment_monitor.py script was making unbound HTTP requests via urlopen() with no socket timeout. When AWS token refresh stalled mid-request, the entire monitor process would hang indefinitely. Solution: set a global socket timeout at module import:
import socket
socket.setdefaulttimeout(30)
This blanket timeout covers both API calls and token refresh, making the monitor hang-proof. The payment monitor now completes or fails fast, every time.
AWS Credential Profile Mismatch
The rebuild-crew-pages and unsubscribe-monitor scripts were using the default AWS profile, but the LaunchAgent jobs needed the queenofsandiego profile (where the S3 buckets and DynamoDB tables live). Each script was updated to explicitly set AWS_PROFILE=queenofsandiego before invoking boto3. This also fixed the SMS alerting path — the alert script needed to assume the right role to access SNS.
SMS Flush Script Missing Executable Bit
The sms-flush wrapper in ~/bin/ wasn't marked executable, so LaunchAgent failures weren't triggering alerts. Fixed by adding the executable bit and validating the shebang and import paths.
The Watchdog: Lambda + EventBridge + S3 Heartbeat
Because launchd jobs can fail silently (especially if you don't actively watch logs), a second-order health check watches the watcher. jada-tests-watchdog is a Lambda function that runs every 6 hours on EventBridge schedule. It:
- Reads
s3://jada-test-heartbeat/heartbeat.json(published by the nightly test suite after completing) - Checks that the heartbeat timestamp is less than 26 hours old
- Verifies all test names in the heartbeat have status
passed - If either check fails, publishes an SMS alert to the ops phone via SNS
The watchdog is deployed from /Users/cb/icloud-repos/sites/queenofsandiego.com/tools/jada-tests-watchdog/lambda_function.py with an IAM role that permits s3:GetObject on the heartbeat bucket and sns:Publish to the alert topic.
Key Decisions
LaunchAgent over Cron
macOS LaunchAgent is more reliable than cron for long-running tasks because it respects system sleep, can be introspected via launchctl list, and logs to syslog. The test runner checks LaunchAgent exit codes directly rather than parsing log files, making health detection deterministic.
S3 Heartbeat as Service Discovery
Rather than hardcode IP addresses or check ports directly, tests and the watchdog communicate through a shared S3 object. This decouples components: the Lambda doesn't need to know where the test runner lives, only that it periodically publishes a heartbeat. If infrastructure changes (different Mac, CloudFront endpoints, etc.), the heartbeat protocol stays the same.
Test Names as Test Data
Test records in the pipeline use deterministic naming (TEST-PIPELINE-{TIMESTAMP}) and are created fresh on each run. This eliminates test isolation bugs: no shared state, no teardown errors, each run is a clean test.
Fixture-Based Assertions Over Mock-Based
The crew pages test fetches real HTML from CloudFront and checks for real HTML structure (noindex meta tag, GA script), not mocked responses. This catches cases where the templating is right but the S3 path is wrong or CloudFront is serving stale content. Mocks would have passed when the real system was broken.
Infrastructure Details
- S3 Buckets:
jada-test-heartbeat(heartbeat signal), per-site buckets for crew pages - CloudFront Distributions: Routes crew page requests to S3 with cache headers; Lambda@Edge injects GA tags
- DynamoDB: Bookings table with GSI on confirmation status for test queries
- SNS Topic: Connected to ops phone via SMS delivery; used by watchdog and alert scripts
- EventBridge Rule: Triggers
jada-tests-watchdogoncron(0 */6 * * ? *)(every 6 hours)
What's Next
The test infrastructure is now deterministic and self-healing: if any component fails, automated alerts fire within 6 hours, and every fix is immediately verified by the next cron run. The next phase is expanding coverage to payment webhook handling, crew manifest generation, and calendar API sync — the remaining load-bearing surfaces in the booking pipeline.
```