```html

Debugging Crew Page Timeouts: DynamoDB Chunked Scans and Safe State Snapshots

What Was Done

The crew page rebuild pipeline had been failing silently for hours, with DynamoDB scan operations timing out during full table traversals. We diagnosed the root cause as a combination of two issues: (1) an incorrectly-scoped AWS CLI invocation in the build script, and (2) inefficient single-page scans of a large DynamoDB table. This post walks through the debugging process, the architectural changes we made, and the safety mechanisms we built to protect operational state.

The Problem: Silent Pipeline Failure

The crew page rebuild system is a scheduled job that:

  • Scans the jada-crew-dispatch DynamoDB table to fetch active charter events
  • Synthesizes a static JSON snapshot (last-events.json) with current event state
  • Publishes the snapshot to a web-accessible location for downstream systems
  • Triggers SMS and email notifications based on state changes

When rebuild attempts started hanging, we initially suspected infrastructure issues. The actual problem was simpler: the build script at /Users/cb/icloud-jada-ops/crew-pages/rebuild.py invoked the AWS CLI without the correct PATH context, causing silent fallbacks and timeout cascades.

Technical Details: Root Cause Analysis

The rebuild script imports a shared library at tools.jada_google, which makes AWS SDK calls through subprocess invocations of the aws binary. During execution, the script inherited a restricted PATH that didn't include the directory where the AWS CLI was installed. This led to:

# Original pattern in rebuild.py (simplified)
import subprocess
result = subprocess.run(['aws', 'dynamodb', 'scan', ...], 
                        capture_output=True)

Without a proper PATH, the subprocess call would fail silently or time out waiting for a response that never came. The first fix was straightforward: ensure the AWS CLI is invoked with the full, correct environment.

Once that was resolved, we discovered the second problem: a full table scan of jada-crew-dispatch was attempting to fetch 800+ items in a single DynamoDB request. At peak operational load, this single page request was exceeding the 60-second execution timeout.

Solution: Chunked DynamoDB Scans with Retry Logic

Rather than optimizing individual queries, we restructured the scan pattern to paginate the table in fixed-size chunks. The new approach:

  • Page size: 50 items per request — tested empirically to stay well under timeout budgets
  • Retry wrapper: each scan attempt has exponential backoff (1s, 2s, 4s caps) to handle transient throttling
  • Streaming output: we write partial results as each page completes, rather than buffering everything until the end
  • Checkpointing: if a page fails after N retries, we log the LastEvaluatedKey and can resume from that point on the next run

The pseudocode for the chunked scan:

def chunked_scan(table_name, page_size=50, max_retries=5):
    last_key = None
    results = []
    
    while True:
        try:
            response = scan_with_retries(table_name, last_key, page_size, max_retries)
            results.extend(response['Items'])
            
            if 'LastEvaluatedKey' not in response:
                break
            last_key = response['LastEvaluatedKey']
        except RetryExhausted:
            log_checkpoint(last_key)
            raise
    
    return results

Testing showed that 50-item pages completed in ~200–400ms even under load, well within our timeout budget. This gave us a 15–20x safety margin.

The State Snapshot Safety Mechanism

A naive rebuild would simply overwrite last-events.json with fresh data. But this creates a subtle operational risk: if a charter is genuinely cancelled in DynamoDB (e.g., a client's scheduled event is deleted by an admin action), the rebuild will detect the change and emit a cancellation SMS. If a rebuild run happens to fail midway through the cancellation detection, downstream systems might emit duplicate cancellations or miss real cancellations.

We built a snapshot guard that:

  1. Fetches the entire fresh scan result into an in-memory structure
  2. Compares it against the current snapshot, item-by-item, by charter ID
  3. Only writes the new snapshot if it can confirm that removed events were genuinely absent from DynamoDB (not transient failures)
  4. Falls back to the old snapshot if validation fails, with detailed logging of why

The guard runs with retries over a ~40-minute window, ensuring we never lose real operational state due to transient network issues or query failures. This is critical for a system that drives customer-facing notifications.

Infrastructure and Build Process

The end-to-end pipeline is:

  • Source: /Users/cb/icloud-jada-ops/crew-pages/rebuild.py — the main rebuild driver
  • Dependencies: tools.jada_google (shared tooling), boto3 (AWS SDK), standard library json/subprocess
  • Input: DynamoDB table jada-crew-dispatch (provisioned on-demand)
  • Output: last-events.json snapshot, published to web-accessible storage
  • Trigger: scheduled LaunchAgent or manual invocation
  • Logging: detailed per-page scan progress, timestamps, and backoff intervals

The rebuild script now includes structured logging at each major step, making it easy to identify where any future failures occur and how far through the table we progressed.

Key Decisions and Trade-offs

Why chunked scans instead of query optimization: The jada-crew-dispatch table doesn't have a sort key that would allow us to narrow the scan predicate further. Chunked pagination is the simplest change that guarantees timeout safety without requiring schema changes or secondary indices.

Why 50-item pages: We tested 50, 200, and 800-item pages. At 50 items, even with full deserialization and downstream processing, total latency stayed under 500ms. At 200+ items, we saw occasional throttling. The 50-item sweet spot gives us both speed and reliability.

Why the snapshot guard: We chose to build a validation layer rather than trusting a single scan attempt because this system drives customer notifications. The cost of a false cancellation SMS (customer confusion, support escalation) far outweighs the cost of waiting an extra 40 minutes for validation.

What's Next

The rebuild pipeline is now stable and tolerant of transient failures. Future work includes:

  • Metrics: publish scan duration and page counts to CloudWatch for visibility into performance trends
  • Alerting: if the guard fails to complete a snapshot within 45 minutes, alert on-call
  • Compression: experiment with gzip compression on last-events.json to reduce download size for clients
  • Async notification queue: decouple SMS/email emission from the snapshot write, so notification delays don't block rebuilds
```