```html

Building Deterministic Booking Pipeline Observability: From Ad-Hoc Checks to Continuous Verification

What Was Done

Redefined "done" for a booking pipeline system to mean: deterministic automated tests running on cron that continuously verify the entire critical path stays working. This session built out a comprehensive health and test suite covering lead intake → proposal → payment → confirmation → waivers/manifest → calendar/DDB → crew dispatch → live sites. Five distinct root causes were identified and fixed, three new test suites created, and a Lambda-based watchdog deployed with SMS alerting.

Technical Details: The System Architecture

Orchestration Layer: LaunchAgent-Based Cron

The system uses macOS LaunchAgents as the orchestration backbone, stored in /Users/cb/Library/LaunchAgents/. Each job is defined as a plist (e.g., com.jada.tests-nightly.plist) that runs shell wrappers in /Users/cb/bin/. This approach was chosen because:

  • Native to the deployment environment (macOS)
  • Integrated with system logging via log stream --predicate
  • Exit codes and timestamps recorded for post-failure analysis
  • No external cron service dependency

Test Suite Organization

Three new pytest modules were created in /Users/cb/icloud-repos/sites/queenofsandiego.com/tests/:

  • test_launchd_jobs.py — Verifies all critical LaunchAgent jobs ran recently and exited cleanly (checks plist definitions, log freshness, exit codes)
  • test_crew_pages_gate.py — Confirms crew-facing pages build correctly, deploy to S3, and exist behind CloudFront
  • test_booking_pipeline.py — End-to-end verification: database state (DynamoDB), payment processing, confirmation emails, and manifest generation

Each test is deterministic and hermetic — they read system state (logs, S3, DynamoDB, SES delivery logs) rather than triggering live actions. This lets them run frequently without side effects.

Payment Monitor: The Timeout Bug

The payment monitor (/Users/cb/icloud-repos/tools/jada_payment_monitor.py) was hanging indefinitely. Root cause: urllib.request.urlopen() called without a timeout. A single stalled TCP handshake to the payment processor wedged the entire monitor.

Fix applied:

socket.setdefaulttimeout(15)  # Global timeout for all socket operations
# Plus explicit timeout in urlopen(url, timeout=20) calls
# Plus retry logic with exponential backoff for transient failures

This pattern was applied to all external calls: Google API (analytics), payment processor, and DynamoDB. The timeout covers both initial connection and token refresh operations, preventing the entire job from wedging on a single slow upstream service.

Credential Configuration: Fixing AWS Profile Chaos

Multiple scripts were using different AWS profile names, causing silent failures. Audit found:

  • jada_payment_monitor.py — using default profile (unconfigured)
  • process_unsubscribes.py — using default profile
  • sms-flush — using default profile
  • Rebuilding crew pages — using default profile

All were corrected to use the queenofsandiego profile explicitly, which has the correct SES, S3, and DynamoDB permissions. The fix was applied consistently across shell wrappers and Python imports:

# Python
session = boto3.Session(profile_name='queenofsandiego')

# Shell wrapper
export AWS_PROFILE=queenofsandiego

Infrastructure Changes

Lambda Watchdog: Remote Health Verification

Created /Users/cb/icloud-repos/sites/queenofsandiego.com/tools/jada-tests-watchdog/lambda_function.py, a Lambda function that:

  • Runs every 6 hours via EventBridge rule jada-tests-watchdog-schedule
  • Checks S3 for heartbeat files published by local test jobs
  • Verifies heartbeat freshness (must be less than 8 hours old)
  • Sends SMS alerts via SNS if heartbeat is missing or stale
  • Provides remote verification that the local cron infrastructure is functioning

Why Lambda instead of checking the local machine directly? The booking system must remain available even if the monitoring machine restarts. Lambda provides an independent, always-on verification layer.

Heartbeat Publishing

The main test runner (/Users/cb/bin/jada-tests-nightly.py) now publishes a heartbeat object to S3 after each run:

s3.put_object(
    Bucket='queenofsandiego-bookings',
    Key='system/heartbeat/tests-nightly.json',
    Body=json.dumps({
        'timestamp': int(time.time()),
        'exit_code': 0,
        'failed_tests': []
    })
)

This provides a single source of truth for "did the critical tests run successfully" that both the local machine and remote Lambda can query.

SMS Alerting Chain

Alerts flow: LaunchAgent exits with code 1 → wrapper script calls /Users/cb/bin/sms-flush → SNS topic → SMS to ops phone. This pattern is consistent across all critical jobs, enabling real-time notification of pipeline failures.

Key Decisions

Why Pytest Over bash/shell scripts: Pytest provides deterministic output, structured failure reporting, and integration with CI systems. Each test is a clearly-named function with setup/teardown, making it obvious what's being verified.

Why Read State Rather Than Trigger It: The tests query existing system state (logs, databases, S3) instead of triggering test bookings. This eliminates side effects and allows frequent runs without polluting production data.

Why EventBridge + Lambda for Remote Verification: LaunchAgents are powerful but local-only. A remote watchdog ensures that infrastructure failures on the deployment machine don't go unnoticed. The 6-hour check interval balances detection latency against API costs.

Why Explicit Timeouts Everywhere: In a distributed system, "no response" is a valid response that means "this service is down." Waiting indefinitely is not a strategy. 15-20 second timeouts ensure that one slow upstream service doesn't cascade into multiple wedged processes.

What's Next

The system is now continuously verifying the entire booking pipeline. The next phase is hardening the test assertions themselves — ensuring they catch edge cases in payment reconciliation, manifest generation, and calendar sync. This session established the infrastructure for deterministic testing; the next work is deepening the coverage to catch the subtle failures that only emerge in production.

```