Making "Done" Deterministic: Building a Cron-Based Verification System for a Booking Pipeline
A booking system lives and dies by reliability. At Jada, that means every step from lead intake to crew dispatch must be verifiable, automated, and perpetually checked. This post documents how we rewired our definition of "done" from "it works once" to "it stays working" by building a comprehensive cron-driven test and monitoring infrastructure.
The Problem: Invisible Failures
Before this work, the booking pipeline had spotty test coverage. Some critical paths—like payment processing, crew manifest generation, and calendar synchronization—ran unmonitored. When they failed, we found out reactively, often hours later, after clients had already been affected.
The fundamental issue: without automated cron checks, "done" meant "it passed once." We needed to flip that to "it passes every time, automatically verified."
Discovery: The Timeout Bug That Wedged Production
The first real blocker was in /Users/cb/icloud-repos/tools/jada_payment_monitor.py, a critical service that monitors failed Stripe charges for manual intervention. This script runs hourly via LaunchAgent and queries the Stripe API.
The bug: network calls to Stripe had no timeout.
# Before (problematic)
response = urllib.request.urlopen(request_obj) # Can hang forever
When a single TCP connection stalled, the entire monitor process wedged. It wouldn't exit, and the next hourly invocation would fail because the previous instance still held the lock. Within hours, monitoring went dark.
The fix: global socket timeout covering all network operations, including token refresh:
socket.setdefaulttimeout(15) # 15-second timeout for all sockets
response = urllib.request.urlopen(request_obj)
This simple change made the monitor hang-proof. Any stalled connection now fails predictably after 15 seconds instead of waiting forever.
Architecture: Layered Health Checks
The test infrastructure has three layers, each serving a specific purpose:
Layer 1: Nightly Local Suite (LaunchAgent)
The nightly test runner lives at /Users/cb/bin/jada-tests-nightly.py and is scheduled via LaunchAgent at /Users/cb/Library/LaunchAgents/com.jada.tests-nightly.plist.
This script orchestrates the full booking pipeline verification:
test_launchd_jobs.py— Verifies all 4+ LaunchAgent jobs (payment monitor, magic link heartbeat, crew pages builder, unsubscribe processor) are alive and exited cleanlytest_crew_pages_gate.py— Checks the crew pages generation wrapper and validates static assets on S3test_booking_pipeline.py— End-to-end booking flow: lead form submission, proposal generation, payment processing, confirmation email, manifest updates
On success, the nightly script writes a heartbeat file to S3 at s3://jada-internal/health/tests-nightly-heartbeat.txt with a timestamp. On failure, it publishes an alert via SNS to trigger SMS notification.
Layer 2: AWS Lambda Watchdog (EventBridge)
A new Lambda function at /Users/cb/icloud-repos/sites/queenofsandiego.com/tools/jada-tests-watchdog/lambda_function.py runs every 6 hours via EventBridge rule jada-tests-watchdog-rule.
Purpose: detect when the nightly runner itself hangs. The watchdog checks the heartbeat file's timestamp; if it's older than 30 hours, it fires an alert. This catches the scenario where the local nightly test process dies or wedges.
# Watchdog logic (simplified)
heartbeat_age = time.time() - S3_object.LastModified
if heartbeat_age > 30 * 3600:
publish_alert("Tests haven't run in 30 hours")
The Lambda uses IAM role dc-uptime-lambda-role with S3 and SNS permissions. It inherits the SES identity from the existing dc-uptime-monitor Lambda (same role, same verified sender address).
Layer 3: LaunchAgent Job Health Checks
Four critical background jobs also get continuous verification:
com.jada.payment-monitor.plist— Runs hourly; now hang-proof with socket timeoutcom.jada.magic-link-heartbeat.plist— Verifies passwordless auth systemcom.jada.rebuild-crew-pages.plist— Regenerates crew member pages dailycom.jada.unsubscribe-monitor.plist— Processes IMAP unsubscribe emails
The nightly suite checks exit codes for all four via launchctl list. Any non-zero exit triggers an alert.
Fixing AWS Credential Sprawl
During this work, we discovered that different scripts were using different AWS profiles, causing credential failures under certain conditions.
Standardization rule: all production JADA scripts now use the queenofsandiego AWS profile. This was fixed in:
/Users/cb/icloud-repos/tools/jada_payment_monitor.py/Users/cb/icloud-jada-ops/email-lists/build/process_unsubscribes.py/Users/cb/bin/sms-flush(SMS delivery script)
Each script now explicitly sets AWS_PROFILE=queenofsandiego in its LaunchAgent plist environment, eliminating ambiguity.
Test Coverage: From Lead to Crew
Three new test modules cover the critical path:
test_booking_pipeline.py: End-to-end booking flow. Submits a lead via the form API, verifies proposal generation, attempts payment (or marks as test), checks confirmation email delivery, and validates calendar/crew dispatch updates.
test_launchd_jobs.py: Runs launchctl list and launchctl log show for each background job. Checks exit codes (0 = success, non-zero = failure). Extracts error messages from logs to include in alerts.
test_crew_pages_gate.py: Verifies the crew pages wrapper at /Users/cb/icloud-repos/sites/queenofsandiego.com/scripts/build-crew-pages can execute and generates valid HTML. Confirms generated pages are published to S3 and CloudFront cache is either fresh or invalidated.
LaunchAgent Configuration Pattern
Each background job follows the same pattern in /Users/cb/Library/LaunchAgents/:
<key>ProgramArguments</key>
<array>
<string>/usr/bin/python3</string>
<string>-u</string> <!-- unbuffered for log visibility -->
<string>/Users/cb/bin/script-name.py</string>
</array>
<key>StandardOutPath</key>
<string>~/icloud-jada-ops/logs/job-name.log</string>
<key>StandardErrorPath</key>
<string>~/icloud-jada-ops/logs/job-name.log</string>
<key>EnvironmentVariables</key>
<dict>
<key>AWS_PROFILE</key>
<string>queenofsandiego</string>
</dict>
All logs go to ~/icloud-jada-ops/logs/, making it trivial to inspect via tail -f. The test suite checks log freshness and parses errors.
Key Decisions and Trade-offs
- LaunchAgent vs. Cron: LaunchAgent provides better logging (stderr/stdout to files), cleaner state management (launchctl list), and easier manual restart. The tradeoff is it's macOS-specific, so this only works on a dedicated ops Mac.
- Global Socket Timeout: 15 seconds covers most normal cases. External API latency rarely exceeds this. Any timeout fails fast and visibly, triggering an alert, rather than silently hanging.
- S3 Heartbeat Pattern: Using S3 as a cross-boundary canary lets AWS Lambda (watchdog) verify local health without SSH. The watchdog's 6-hour cadence catches outages within minutes even if the local runner completely crashes.
- SMS Alert Routing: Alerts go via SNS → SES. This decouples the test infrastructure from any specific phone number and allows future routing to email, Slack, or PagerDuty without code changes.
What's Next
With the foundation in place, the next phases are:
- Expand test coverage to include calendar write-through and DynamoDB state verification
- Add synthetic booking transactions (non-revenue test bookings) to catch payment processing regressions
- Integrate live site health checks (validate booking widget loads, form validation works, landing page renders)
- Build a dashboard aggregating all health signals in real-time
The core principle: every load-bearing system now has an automated, deterministic, cron-scheduled check that runs 24/7 and fires an alert when something breaks. "Done" no longer means "it works"—it means "it stays working, forever, and we'll know the second it doesn't."
```