Debugging and Hardening the JADA Crew Dispatch Pipeline: PATH Issues, DDB Timeouts, and Operational Audit Findings
What Was Done
A comprehensive operational audit uncovered critical failures in the JADA charter business pipeline, with the most time-sensitive issue being a broken crew page rebuild process scheduled to serve a 30-guest charter departing tomorrow. The rebuild pipeline—responsible for fetching crew availability from DynamoDB, cross-referencing against Google Sheets crew assignments, and publishing a static HTML crew page—was failing silently due to a PATH environment variable issue combined with deeper DynamoDB scan timeout problems in the staging environment.
This post covers the debugging process, the architectural changes made to stabilize the pipeline, and the broader infrastructure findings that emerged during the audit.
The Problem: Multiple Failure Points
1. Immediate: Crew Page Rebuild Script Broken
The rebuild process invoked at `/Users/cb/icloud-jada-ops/crew-pages/rebuild.py` was failing with:
aws: command not found
Root cause: the script spawns AWS CLI calls via subprocess but was executing in an environment where the `aws` binary wasn't on the PATH. While LaunchAgent execution environments inherit some system state, they do not inherit shell-specific configurations (.zshrc, .bashrc). The AWS CLI installation in `/usr/local/bin/aws` was invisible to the subprocess call.
Immediate fix: Modified `/Users/cb/icloud-jada-ops/crew-pages/rebuild.py` to use an absolute path to the aws binary: `/usr/local/bin/aws` in all subprocess calls, eliminating the PATH dependency.
2. Secondary: DynamoDB Scan Timeouts Under Load
Even with the PATH issue resolved, dry-run tests revealed the rebuild hanging indefinitely during the DynamoDB scan phase. The script was scanning the `jada-crew-dispatch` table without pagination, attempting to load all crew records in a single request. With network latency and table growth, individual pages were taking 45–60 seconds to return, and the implicit 5-minute timeout was being exceeded silently.
Diagnostic steps:
- Timed a
COUNTscan of `jada-crew-dispatch` — 12 seconds for metadata, but full data retrieval hung - Probed different page sizes (50, 200, 800 items per page) with increasing timeout caps
- Walked the scan paginator to identify which page was hanging
- Performed a keys-only scan of the problem page followed by per-item fetches to isolate a poison record
The issue: a single crew record with oversized attributes (likely raw SMS dumps synced from Carole's phone) was causing serialization delays on that page, which then blocked the entire scan under the no-pagination model.
Technical Solutions Implemented
DynamoDB Scan Refactor: Chunked Pagination
Rewrote the DynamoDB scan logic in `rebuild.py` to fetch crew records in explicit chunks with per-chunk retry logic:
# Pseudo-code structure (actual implementation uses boto3 paginator)
items = []
page_size = 50 # Conservative page size
last_key = None
while True:
response = scan_crew_dispatch(page_size, last_key)
items.extend(response['Items'])
last_key = response.get('LastEvaluatedKey')
if not last_key:
break
# Per-chunk timeout/retry logic here
This approach:
- Processes 50 items at a time (empirically safe limit observed during testing)
- Returns control periodically to prevent hanging
- Allows per-chunk error recovery without losing progress
- Reduces memory pressure compared to loading all records upfront
AWS CLI Invocation Hardening
Created a retry-hardened AWS helper at `/Users/cb/icloud-jada-ops/bin/aws-cli-wrapper` that wraps all subprocess calls with:
- Exponential backoff (1s, 2s, 4s, 8s maximum)
- Explicit timeout enforcement (60-second hard cap per invocation)
- Structured error logging to `/var/log/jada-ops/aws-calls.log`
- Fallback to a cached crew page if all retries are exhausted (fail-safe for LaunchAgent execution)
The rebuild script now calls this wrapper instead of invoking aws directly, providing defense-in-depth against transient network issues.
Import Path Cleanup
The rebuild script was importing Google Sheets auth from a relative path that failed in LaunchAgent execution context. Refactored /Users/cb/icloud-jada-ops/crew-pages/rebuild.py to import from tools.jada_google, which lives in the shared `icloud-repos` package with proper module resolution.
Infrastructure and Operational Findings
The audit revealed several systemic issues requiring immediate attention:
- Plaintext AWS credentials: Long-term AWS access key hardcoded in `/Users/cb/Library/LaunchAgents/com.dangerouscentaur.dev-agent.plist`. Credential rotation and migration to AWS profiles required.
- Email campaign pipeline failure: July 2 campaign sent 0 of 3,640 emails due to control characters in the email list; no alerting caught this for 24 hours.
- Unsubscribe watcher offline: Broken for 30 days, creating CAN-SPAM compliance risk.
- Ledger data quality: Only 7 entries with 3 missing payment totals; reconciliation needed before next financial review.
- Privacy exposure: ~306KB of raw SMS logs syncing to iCloud Drive. Requires data minimization or encryption.
Key Decisions
Why chunked pagination over async workers: The audit timeline was tight (24 hours to fix before the Jul 4 charter). A chunked-scan approach is simpler to test and debug than introducing threading or async patterns, and it works within the existing LaunchAgent execution model without additional orchestration.
Why absolute paths instead of PATH manipulation: LaunchAgent environments are inherently constrained. Rather than fighting environment inheritance, we made the script self-contained by hardcoding the known location of the AWS CLI.
Why a wrapper over inline retry logic: Centralizing retry/timeout logic in a single wrapper allows the policy to evolve without touching the rebuild script repeatedly. It also creates a reusable component for other AWS operations in the JADA infrastructure.
What's Next
- Dry-run validation: The crew page rebuild is currently running a dry-run with chunked DDB scans to verify end-to-end success before the Jul 4 charter.
- Credential rotation: Rotate the plaintext AWS key and update LaunchAgent configuration to use AWS profiles (stored in `~/.aws/config` with temporary session tokens).
- Email alerting: Add CloudWatch alarms to the SES campaign pipeline; alert if zero emails are sent when the list size exceeds a threshold.
- Unsubscribe watcher: Rebuild the email list filter with robust error handling and monitoring.
- Data sanitization: Move raw SMS logs out of iCloud sync or apply encryption at rest.
Full audit findings and Board minutes have been saved to `/Users/cb/icloud-jada-ops/AUDIT-2026-07-03-full.md` and `/Users/cb/icloud-jada-ops/HANDOFF-2026-07-03.md` for operational continuity.