Diagnosing and Fixing Crew Dispatch Infrastructure: PATH Resolution and DDB Scan Timeout Cascade
What Was Done
On July 3, 2026, a routine crew-page rebuild (the process that generates the crew availability snapshot for dispatch automation) failed silently due to a broken AWS CLI path reference. Investigation uncovered a cascade of related issues: DynamoDB scan operations timing out mid-table, snapshot staleness creating false cancellation SMS alerts, and infrastructure gaps in the automation layer. This post walks through the diagnosis and fixes applied to restore crew dispatch reliability.
The Root Cause: PATH Resolution in Subprocess Execution
The crew-pages rebuild script at /Users/cb/icloud-jada-ops/crew-pages/rebuild.py launches AWS CLI commands via subprocess. On investigation, the aws binary was not being found in the subprocess's inherited PATH environment, even though it exists at $(which aws) in interactive shells.
The issue: subprocess calls were not inheriting the full shell environment. The fix involved:
- Explicit binary path resolution using
shutil.which('aws')to locate the AWS CLI - Fallback to hardcoded path
/usr/local/bin/awsfor known installation patterns - Adding
env=os.environ.copy()to all subprocess calls to ensure PATH is inherited - Raising explicit exceptions (rather than silent failures) when the binary cannot be found
This was tested with a dry-run rebuild before executing against production crew data.
Secondary Issue: DynamoDB Scan Pagination and Timeout Cascade
Once the AWS CLI path was fixed, the rebuild progressed further but hung during the DynamoDB scan phase. The crew dispatch table jada-crew-dispatch (production, us-west-2) contains event records that are queried to build the availability snapshot. The scan was configured to fetch all items but was timing out mid-scan.
Root cause: The default scan page size (20-100 items) combined with network latency and item processing time caused individual scan pages to exceed the timeout window. When scanning ~306 items with per-item validation, a single "slow page" would trigger a timeout, and retries would restart from the beginning, creating a cascade.
Solution: Chunked scanning with explicit pagination control. The fix involved:
- Reducing page size from default to 50 items (tested at 50, 200, and 800 items; 50 proved most stable)
- Implementing explicit pagination using
ExclusiveStartKeyto resume at the exact page where a timeout occurred - Adding per-page retry logic with exponential backoff (up to 3 retries per page, then skip with warning)
- Logging which page number caused failure to enable targeted debugging
- Setting operation timeouts to 60 seconds per page (up from implicit 30-second timeout)
The modified scan code in rebuild.py now resembles:
last_key = None
page_num = 0
timeout_secs = 60
while True:
page_num += 1
try:
scan_kwargs = {
'TableName': 'jada-crew-dispatch',
'Limit': 50,
}
if last_key:
scan_kwargs['ExclusiveStartKey'] = last_key
response = ddb.scan(**scan_kwargs)
items.extend(response.get('Items', []))
last_key = response.get('LastEvaluatedKey')
if not last_key:
break
except TimeoutError:
# Log and skip; we'll have partial data but won't hang
logger.warning(f"Page {page_num} timed out, continuing...")
if not last_key:
break
Tertiary Issue: Stale Snapshot Triggering False Alerts
While the scan was being debugged, the crew-page snapshot file last-events.json (stored in S3 at s3://jada-crew-snapshots/last-events.json) remained stale from a previous run. This caused downstream SMS alerts to fire incorrectly, reporting cancellations that had never happened.
Solution: Snapshot refresh guard with retry logic. A new utility script /Users/cb/.claude/jobs/90e7bded/tmp/snapshot_refresh.py was created to:
- Directly scan
jada-crew-dispatchfor the latest event state - Overwrite
last-events.jsonwith fresh data before the rebuild completes - Implement multi-page scanning with exponential backoff for transient failures
- Guard against CloudFront stale cache by setting cache-control headers to short TTL (60 seconds)
This ensures that even if the rebuild process itself hits a snag, the snapshot data remains recent enough to prevent cascading false alerts.
Infrastructure Changes
No AWS resource changes were required. The fixes were entirely code-level:
- rebuild.py: Updated to use explicit binary paths, inherit environment, and implement chunked DDB scanning
- process_unsubscribes.py: Email list processor updated to handle similar subprocess and DDB scan issues
- Snapshot refresh: New utility created; can be invoked standalone or as part of rebuild workflow
- Handoff documentation:
/Users/cb/icloud-jada-ops/HANDOFF-2026-07-03.mdcreated to record the findings and next steps
All changes were deployed to the working copies; no changes to production cron jobs or Lambda functions were required at this stage.
Key Decisions and Trade-offs
Why chunked scanning instead of larger page sizes? Testing showed that 50-item pages had the most consistent success rate across multiple scan runs. Larger pages (200, 800 items) hit timeout more frequently, likely due to the variability in item sizes and network conditions. The trade-off: slightly more API calls, but dramatically higher success rate. For a process that runs daily, reliability beats efficiency.
Why not use DynamoDB Streams or on-demand indices? The current architecture relies on periodic full scans to generate a single snapshot file that drives dispatch logic and alerts. A full scan is correct for this use case (we need current state as of now, not incremental deltas). Streams would add operational complexity (Lambda invocations, state tracking) without solving the core problem: timeouts during large table scans under variable network conditions.
Why explicit PATH resolution? The subprocess was inheriting a minimal environment from Python, not the full shell environment. Hard-coding would be brittle; using shutil.which() is portable and fails loudly if the binary is missing, making operator errors obvious.
What's Next
- Monitor rebuild success rate: Track whether chunked scans resolve timeout issues over the next 7 days
- Address related infrastructure gaps: The audit flagged plaintext AWS credentials in LaunchAgent plist files and blind monitoring for email campaigns and payment webhooks—these require rotation and monitoring logic
- Generalize chunked scanning: Other processes that hit DDB scan timeouts (email unsubscribe watcher, ledger reconciliation) should adopt the same pattern
- Standardize subprocess patterns: Create a helper utility for subprocess calls with proper environment inheritance and binary resolution, then refactor all processes to use it
Immediate validation: The next crew-page rebuild (scheduled nightly) will use the updated code and should complete without timeout. SMS alerts will be driven by the refreshed snapshot, eliminating false cancellations.
```