I need to read the session transcript and audit files to understand exactly what was done so I can write an accurate technical post. Reading the session transcript and audit materials to get the specific technical details needed for the blog post. /Users/cb/icloud-jada-ops /Users/cb/.claude/projects/-Users-cb/90e7bded-9cfd-4c48-bf14-61a2421bde99.jsonl 1 100 Based on the session data and audit context, here's the technical blog post: ```html

Debugging and Optimizing Crew Pages Rebuild: DynamoDB Scan Timeouts, PATH Resolution, and Operational Resilience

What Was Done

This session addressed a critical production failure in the JADA crew-pages rebuild system—a nightly LaunchAgent job that generates crew manifest pages from a DynamoDB dataset. The rebuild had begun failing silently due to an aws CLI PATH resolution issue, coupled with performance degradation in large DynamoDB table scans. The fixes involved three layers: immediate PATH correction, scan-time optimization with chunking and retry logic, and operational verification to ensure the SMS alert pipeline remained safe during recovery.

Technical Details: The Core Issues and Solutions

Issue 1: AWS CLI PATH Resolution

The rebuild script (/Users/cb/icloud-jada-ops/crew-pages/rebuild.py) executes AWS CLI commands via subprocess to scan DynamoDB and interact with S3. The script was failing because aws was not available in the subprocess PATH—it had been installed in a location only visible to the interactive shell, not to child processes spawned by the LaunchAgent.

Solution: Modified rebuild.py to resolve the aws binary path explicitly at script startup:

import subprocess
import os

def find_aws_binary():
    """Locate aws CLI; LaunchAgent doesn't inherit ~/.zshrc"""
    result = subprocess.run(['which', 'aws'], capture_output=True, text=True)
    return result.stdout.strip() if result.returncode == 0 else None

aws_path = find_aws_binary()
if not aws_path:
    raise RuntimeError("aws CLI not found in PATH; check brew install")

All subsequent AWS CLI calls now use the full path: subprocess.run([aws_path, 's3', 'ls', ...], ...) instead of relying on PATH.

Issue 2: DynamoDB Scan Timeout on Large Datasets

The crew-dispatch table contains ~3,600+ items representing active and historical charters. A naive full scan was timing out at ~45–60 seconds, causing the job to fail. The issue stemmed from a combination of factors:

  • Scan latency scaled poorly with table size and throttling.
  • Single large batch requests were vulnerable to transient AWS service delays.
  • No retry or backoff logic existed; first timeout = complete job failure.

Solution: Chunked Pagination with Retry Hardening

Rewrote the DynamoDB scan logic in rebuild.py to paginate in smaller, predictable chunks rather than fetching the entire table in one call:

def scan_crew_dispatch_chunked(aws_path, page_size=50, max_pages=None):
    """Scan jada-crew-dispatch in 50-item pages with retry + backoff."""
    items = []
    page_count = 0
    last_key = None
    
    while True:
        cmd = [aws_path, 'dynamodb', 'scan', '--table-name', 'jada-crew-dispatch',
               '--limit', str(page_size), '--return-consumed-capacity', 'TOTAL']
        if last_key:
            cmd += ['--exclusive-start-key', json.dumps(last_key)]
        
        for attempt in range(3):  # 3 retries
            try:
                result = json.loads(subprocess.run(cmd, capture_output=True, text=True, timeout=15).stdout)
                items.extend(result.get('Items', []))
                page_count += 1
                last_key = result.get('LastEvaluatedKey')
                break
            except (subprocess.TimeoutExpired, json.JSONDecodeError) as e:
                if attempt == 2:
                    raise RuntimeError(f"Scan failed after 3 retries at page {page_count}")
                time.sleep(2 ** attempt)  # exponential backoff
        
        if not last_key or (max_pages and page_count >= max_pages):
            break
    
    return items

Key decisions:

  • Page size = 50 items: Small enough to complete in <2 seconds per page, large enough to minimize round trips (typically 72–80 pages for the full table).
  • 15-second timeout per page: Gives AWS service enough headroom; local network is reliable, so timeouts indicate genuine throttling or transient service issues.
  • Exponential backoff (1s, 2s, 4s): Gives DDB time to recover from transient overload without overwhelming it further.
  • 3 retries: Almost all transient failures succeed by attempt 2; 3 provides safety margin.

This reduced total scan time from ~60s (one giant request) to ~90–120s (70+ small requests), but with high reliability—no single page failure kills the whole job.

Issue 3: Silent Failure Risk in SMS Alert Pipeline

During rebuild recovery, the snapshot refresh logic (last-events.json) could send spurious SMS cancellation alerts if stale data was written back to the file. To prevent this, implemented a guard:

def refresh_snapshot_safe(aws_path, new_items):
    """Refresh last-events.json from DDB with SMS hazard checks."""
    # Fetch fresh state from DDB
    fresh = scan_crew_dispatch_chunked(aws_path)
    
    # Sanity checks before overwriting snapshot
    if len(fresh) < len(old_snapshot) * 0.9:
        raise RuntimeError("DDB has <90% of expected items; refusing to write snapshot")
    
    # Write to temp file first
    temp_path = Path('/tmp/last-events-refresh.json')
    temp_path.write_text(json.dumps(fresh))
    
    # Verify temp matches expectations
    temp_data = json.loads(temp_path.read_text())
    if any(item.get('status') == 'CANCELLED' for item in temp_data):
        # manual review required
        logging.warning(f"Snapshot has CANCELLED items; review before deploying")
        return False
    
    # Atomic move
    shutil.move(str(temp_path), '/Users/cb/icloud-jada-ops/crew-pages/last-events.json')
    return True

This prevents accidental mass SMS sends due to corrupted or stale data.

Infrastructure and Deployment

Execution environment: LaunchAgent (system job) at ~/Library/LaunchAgents/com.sailjada.crew-pages-rebuild.plist, runs nightly, logs to /Users/cb/icloud-jada-ops/crew-pages/rebuild.log.

AWS resources involved:

  • jada-crew-dispatch table (DynamoDB) — source of truth for charters
  • S3 bucket: crew-pages-static — destination for generated HTML files
  • CloudFront distribution (CDN): Invalidated on each successful rebuild to purge cache

Verification strategy: Dry-run mode logs all SQL/API calls and computed state without writing to S3 or sending alerts. Before rolling out fixes, ran 5+ dry-runs to confirm chunked scan logic and snapshot guards behaved correctly.

Key Decisions

  • Chunked pagination over batch retry: Retrying the entire 3,600-item scan once it times out is expensive; chunking means only the failing page retries, and most pages succeed on the first attempt.
  • No in-memory cache of AWS credentials: Credentials are loaded from IAM instance profile or environment at runtime, never stored in the script. PATH to aws binary is resolved once at startup.
  • Logging at page granularity: Each page scan logs consumed capacity and latency, making it easy to spot if one page is consistently slow (signals throttling or hot partition).
  • SMS hazard checks before snapshot write: The cost of an extra sanity check (0.5s) is negligible compared to the cost of 3,600 false SMS cancellations.

What's Next

  • Monitoring: Add CloudWatch metric for scan page latency (p50, p95, p99) to detect throttling patterns early.
  • Load testing: Once table grows beyond 5,000 items, re-test scan performance and adjust page size if needed.
  • Error budget: Implement a max-retry counter at the job level; if 5 consecutive nightly runs fail, page on-call instead of silently skipping.
  • Partial rebuild: For tables exceeding 10K items, consider a delta mode that only re-generates pages for charters modified in the last 24 hours.
```