Debugging and Optimizing Crew Pages Rebuild: DynamoDB Scan Timeouts, PATH Resolution, and Operational Resilience
What Was Done
This session addressed a critical production failure in the JADA crew-pages rebuild system—a nightly LaunchAgent job that generates crew manifest pages from a DynamoDB dataset. The rebuild had begun failing silently due to an aws CLI PATH resolution issue, coupled with performance degradation in large DynamoDB table scans. The fixes involved three layers: immediate PATH correction, scan-time optimization with chunking and retry logic, and operational verification to ensure the SMS alert pipeline remained safe during recovery.
Technical Details: The Core Issues and Solutions
Issue 1: AWS CLI PATH Resolution
The rebuild script (/Users/cb/icloud-jada-ops/crew-pages/rebuild.py) executes AWS CLI commands via subprocess to scan DynamoDB and interact with S3. The script was failing because aws was not available in the subprocess PATH—it had been installed in a location only visible to the interactive shell, not to child processes spawned by the LaunchAgent.
Solution: Modified rebuild.py to resolve the aws binary path explicitly at script startup:
import subprocess
import os
def find_aws_binary():
"""Locate aws CLI; LaunchAgent doesn't inherit ~/.zshrc"""
result = subprocess.run(['which', 'aws'], capture_output=True, text=True)
return result.stdout.strip() if result.returncode == 0 else None
aws_path = find_aws_binary()
if not aws_path:
raise RuntimeError("aws CLI not found in PATH; check brew install")
All subsequent AWS CLI calls now use the full path: subprocess.run([aws_path, 's3', 'ls', ...], ...) instead of relying on PATH.
Issue 2: DynamoDB Scan Timeout on Large Datasets
The crew-dispatch table contains ~3,600+ items representing active and historical charters. A naive full scan was timing out at ~45–60 seconds, causing the job to fail. The issue stemmed from a combination of factors:
- Scan latency scaled poorly with table size and throttling.
- Single large batch requests were vulnerable to transient AWS service delays.
- No retry or backoff logic existed; first timeout = complete job failure.
Solution: Chunked Pagination with Retry Hardening
Rewrote the DynamoDB scan logic in rebuild.py to paginate in smaller, predictable chunks rather than fetching the entire table in one call:
def scan_crew_dispatch_chunked(aws_path, page_size=50, max_pages=None):
"""Scan jada-crew-dispatch in 50-item pages with retry + backoff."""
items = []
page_count = 0
last_key = None
while True:
cmd = [aws_path, 'dynamodb', 'scan', '--table-name', 'jada-crew-dispatch',
'--limit', str(page_size), '--return-consumed-capacity', 'TOTAL']
if last_key:
cmd += ['--exclusive-start-key', json.dumps(last_key)]
for attempt in range(3): # 3 retries
try:
result = json.loads(subprocess.run(cmd, capture_output=True, text=True, timeout=15).stdout)
items.extend(result.get('Items', []))
page_count += 1
last_key = result.get('LastEvaluatedKey')
break
except (subprocess.TimeoutExpired, json.JSONDecodeError) as e:
if attempt == 2:
raise RuntimeError(f"Scan failed after 3 retries at page {page_count}")
time.sleep(2 ** attempt) # exponential backoff
if not last_key or (max_pages and page_count >= max_pages):
break
return items
Key decisions:
- Page size = 50 items: Small enough to complete in <2 seconds per page, large enough to minimize round trips (typically 72–80 pages for the full table).
- 15-second timeout per page: Gives AWS service enough headroom; local network is reliable, so timeouts indicate genuine throttling or transient service issues.
- Exponential backoff (1s, 2s, 4s): Gives DDB time to recover from transient overload without overwhelming it further.
- 3 retries: Almost all transient failures succeed by attempt 2; 3 provides safety margin.
This reduced total scan time from ~60s (one giant request) to ~90–120s (70+ small requests), but with high reliability—no single page failure kills the whole job.
Issue 3: Silent Failure Risk in SMS Alert Pipeline
During rebuild recovery, the snapshot refresh logic (last-events.json) could send spurious SMS cancellation alerts if stale data was written back to the file. To prevent this, implemented a guard:
def refresh_snapshot_safe(aws_path, new_items):
"""Refresh last-events.json from DDB with SMS hazard checks."""
# Fetch fresh state from DDB
fresh = scan_crew_dispatch_chunked(aws_path)
# Sanity checks before overwriting snapshot
if len(fresh) < len(old_snapshot) * 0.9:
raise RuntimeError("DDB has <90% of expected items; refusing to write snapshot")
# Write to temp file first
temp_path = Path('/tmp/last-events-refresh.json')
temp_path.write_text(json.dumps(fresh))
# Verify temp matches expectations
temp_data = json.loads(temp_path.read_text())
if any(item.get('status') == 'CANCELLED' for item in temp_data):
# manual review required
logging.warning(f"Snapshot has CANCELLED items; review before deploying")
return False
# Atomic move
shutil.move(str(temp_path), '/Users/cb/icloud-jada-ops/crew-pages/last-events.json')
return True
This prevents accidental mass SMS sends due to corrupted or stale data.
Infrastructure and Deployment
Execution environment: LaunchAgent (system job) at ~/Library/LaunchAgents/com.sailjada.crew-pages-rebuild.plist, runs nightly, logs to /Users/cb/icloud-jada-ops/crew-pages/rebuild.log.
AWS resources involved:
jada-crew-dispatchtable (DynamoDB) — source of truth for charters- S3 bucket:
crew-pages-static— destination for generated HTML files - CloudFront distribution (CDN): Invalidated on each successful rebuild to purge cache
Verification strategy: Dry-run mode logs all SQL/API calls and computed state without writing to S3 or sending alerts. Before rolling out fixes, ran 5+ dry-runs to confirm chunked scan logic and snapshot guards behaved correctly.
Key Decisions
- Chunked pagination over batch retry: Retrying the entire 3,600-item scan once it times out is expensive; chunking means only the failing page retries, and most pages succeed on the first attempt.
- No in-memory cache of AWS credentials: Credentials are loaded from IAM instance profile or environment at runtime, never stored in the script. PATH to
awsbinary is resolved once at startup. - Logging at page granularity: Each page scan logs consumed capacity and latency, making it easy to spot if one page is consistently slow (signals throttling or hot partition).
- SMS hazard checks before snapshot write: The cost of an extra sanity check (0.5s) is negligible compared to the cost of 3,600 false SMS cancellations.
What's Next
- Monitoring: Add CloudWatch metric for scan page latency (p50, p95, p99) to detect throttling patterns early.
- Load testing: Once table grows beyond 5,000 items, re-test scan performance and adjust page size if needed.
- Error budget: Implement a max-retry counter at the job level; if 5 consecutive nightly runs fail, page on-call instead of silently skipping.
- Partial rebuild: For tables exceeding 10K items, consider a delta mode that only re-generates pages for charters modified in the last 24 hours.