Debugging and Fixing a Silent Pipeline Failure: DynamoDB Scan Optimization and PATH Resolution in the Crew Page Rebuild
What Was Done
The crew-page rebuild pipeline — a critical component that generates 15 USCG-compliant crew manifests and web pages every 30 minutes — failed silently for four weeks. The root cause was a PATH resolution issue in the LaunchAgent invoking /Users/cb/icloud-jada-ops/crew-pages/rebuild.py, combined with a separate timeout during DynamoDB table scans. This post documents the debugging methodology, the fixes applied, and the architectural decisions that make the pipeline resilient to recurrence.
The Failure Mode
The pipeline executes as a macOS LaunchAgent (registered in ~/Library/LaunchAgents/com.sailjada.crew-pages.plist) and queries a DynamoDB table named jada-crew-dispatch to fetch the canonical crew roster. The LaunchAgent ran on schedule, but the AWS CLI — invoked as a subprocess within rebuild.py — failed silently because the environment's PATH variable did not include the directory where aws was installed. Since the subprocess call was not wrapped in error handling that would propagate or log the failure, the pipeline appeared to succeed while producing no output.
Why this was dangerous: The crew-page build feeds USCG compliance manifests and email crew notifications. Four weeks of silent failure meant that an upcoming 30-guest charter had no manifest and no crew confirmation — a regulatory and operational crisis discovered only during the audit.
Technical Diagnosis
The debugging process involved three parallel tracks:
- Subprocess PATH isolation: Confirmed that the LaunchAgent's environment was executing in a restricted shell context without the user's standard
~/.zshrcinitialization. Theawsbinary (typically installed via Homebrew to/opt/homebrew/bin/awson Apple Silicon) was not accessible. Solution: explicitly passPATHto the subprocess call inrebuild.py. - DynamoDB scan timeout investigation: A full-table scan of
jada-crew-dispatch(approx. 300+ records) was timing out at ~45 seconds, with no progress logs. The default scan without pagination was attempting to fetch all items in a single batch, exceeding the LaunchAgent's execution timeout window. Diagnosis involved timing scans at various page sizes (50, 200, 800 items per page) to isolate the hanging page and identify a single "poison item" causing excessive processing. - Snapshot stale-data cascade: During troubleshooting, the crew-page rebuild compared live DDB state against a cached snapshot in
last-events.json. A stale snapshot entry (a canceled charter) caused the pipeline to generate false SMS cancellation alerts to crew. Fixing the pipeline alone would have filled crew inboxes with ghost cancellations.
Infrastructure Changes
1. PATH Resolution Fix
In rebuild.py, the subprocess call to the AWS CLI now explicitly constructs the environment:
import os
import subprocess
# Build a clean environment with the correct PATH
env = os.environ.copy()
env['PATH'] = '/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin'
result = subprocess.run(
['aws', 'dynamodb', 'scan', '--table-name', 'jada-crew-dispatch', ...],
env=env,
timeout=60,
check=True
)
This ensures that even if the LaunchAgent runs in a minimal environment, the AWS CLI is found reliably. The explicit timeout prevents indefinite hangs.
2. DynamoDB Scan Chunking and Retry Logic
The original scan fetched all items in a single request. The fix implements pagination with retry backoff:
- Page size: 50 items per DynamoDB request (balanced for API cost and latency).
- Retry strategy: On timeout or throttle, exponential backoff with a 5-second base and 3 max attempts.
- Logging: Each page fetch logs the LastEvaluatedKey and row count, enabling quick diagnosis if a page hangs again.
Example command used during testing:
aws dynamodb scan \
--table-name jada-crew-dispatch \
--projection-expression 'eventId, crewAssignment, eventDate' \
--limit 50 \
--page-size 50
3. Snapshot Refresh Guard
A new script, `snapshot_refresh.py`, runs before the main rebuild and fetches the latest 15 event records from DynamoDB directly into the snapshot file. This replaces the old behavior where the snapshot could be days out of sync. The guard includes:
- Atomic writes to
last-events.json(write-to-temp, then rename). - Verification that all 15 expected crew pages exist before committing the snapshot.
- A 90-second timeout to fail fast if DDB is unreachable.
Key Decisions
- Why pagination over a single scan: The full table can grow to 300+ items. Pagination limits memory per request and makes the pipeline restartable at any page boundary if interrupted. A single large scan is also more likely to be throttled by DynamoDB's burst capacity.
- Why explicit PATH instead of relying on LaunchAgent: LaunchAgents on macOS run in a restricted environment by design. Rather than fight the OS (e.g., by sourcing ~/.zshrc in a login shell — complex and fragile), we make the script self-contained. This also improves portability if the script ever needs to run in CI/CD or on another machine.
- Why a separate snapshot refresh: The crew-page rebuild is I/O-bound (querying DDB, rendering HTML, uploading to S3). Snapshot staleness is a different failure mode from rebuild failure. By isolating the snapshot refresh as a prerequisite with its own logging and timeout, we can detect and fix snapshot drift independently of the rest of the pipeline.
- Why retain retry logic even after fixing the timeout: DynamoDB and AWS API calls are inherently unreliable at scale. Retries with backoff are standard practice and cost almost nothing when the API succeeds on the first try. The risk of removing them later (if someone misinterprets the fix) exceeds the cost of keeping them.
Verification and Deployment
Post-fix validation included:
- Dry-run rebuild with AWS credentials and network connectivity confirmed.
- Live rebuild deployment and monitoring of the LaunchAgent's scheduled run in
rebuild.log. - Verification that all 15 crew pages (including the Jul 4 charter with 30 guests) were generated and uploaded to S3.
- Confirmation that zero false SMS alerts were triggered (snapshot was correctly refreshed).
What's Next
- Monitoring: Add CloudWatch metrics to the LaunchAgent run — track AWS CLI exit codes, DDB scan page counts, and S3 upload latency. Alert on rebuild failures (not just pipeline silence).
- Dependency auditing: Ensure all production scripts that rely on external binaries (aws, gsutil, python3) are wrapped with explicit PATH construction or use absolute paths (
/opt/homebrew/bin/aws, etc.). - Chaos testing: Simulate DynamoDB rate-limit scenarios and verify that retries succeed without manual intervention.
- Documentation: Update the crew-pages README with the PATH resolution pattern and DynamoDB pagination reasoning for future maintainers.