I'll write a technical blog post focused on the core technical work: diagnosing and fixing the silent failures in your crew-page automation pipeline, with emphasis on the PATH isolation and DynamoDB reliability challenges. ```html

Diagnosing Silent Failures: How a LaunchAgent Lost Its PATH and Crashed Crew Dispatch

On June 3, 2026, your crew-page rebuild stopped working. For 30 days, it ran silently every morning—no errors, no alerts, no trace. The system seemed alive. It was completely dead. This post covers the root cause (environment isolation in LaunchAgents), three cascading bugs that made diagnosis harder, and the architectural changes that caught the next failure hours before it sent false SMS to eight crew members.

The Problem: Four Weeks of Silent Failure

The crew-pages/rebuild.py script orchestrates three critical tasks:

  • Scan DynamoDB for recent event changes (new charters, cancellations, manifest updates)
  • Regenerate static HTML pages for each event
  • Trigger SMS alerts to crew via sms-flush if changes are detected

When run manually from the shell, it worked perfectly. When LaunchAgent triggered it at 06:00, it silently failed and logged nothing. The rebuild.py script would start, reach the AWS SDK call, and hang.

Root cause: LaunchAgent runs with a minimal environment—no PATH, no shell initialization. The script called external tools using bare commands (aws, ~/bin/...) that existed in the shell's PATH but were invisible to launchd.

Why This Went Undetected for 30 Days

Three compounding failures:

  • No heartbeat monitoring: LaunchAgent logged no output and exited 0 (a timeout, not a crash). With no alerting layer, a silent exit indistinguishable from success.
  • The snapshot cache masked the outage: rebuild.py keeps a JSON snapshot of the last known state. On failure, it re-ran the previous snapshot, regenerating yesterday's pages. The pages existed; crew got no alerts; no one checked if they were stale.
  • Timeout ambiguity: When the aws call hung, Python didn't crash—it blocked. After ~15 minutes, the system killed it. The script logged nothing. It looked like a completed run.

The Technical Fix

1. Resolve PATH isolation

LaunchAgent environment is sparse. The fix: make the script self-sufficient. Instead of calling bare aws, locate it explicitly:

import shutil
aws_path = shutil.which('aws')
if not aws_path:
    # Fallback: search known installation paths
    for candidate in ['/usr/local/bin/aws', '/opt/homebrew/bin/aws']:
        if os.path.isfile(candidate):
            aws_path = candidate
            break
if not aws_path:
    raise RuntimeError("aws CLI not found in PATH or standard locations")

The same logic applies to any tool the script imports—in this case, tools.jada_google, which itself calls ~/bin/google-token-refresh. When launchd can't find ~/bin, the entire auth chain breaks silently. Solution: use absolute paths and validate at startup.

2. Chunk DynamoDB scans for network resilience

A full DynamoDB table scan on unreliable networks can hang indefinitely. If a single page of results never arrives, the entire operation blocks. The fix: scan in small, bounded chunks with explicit timeouts.

def scan_ddb_paginated(table, page_size=50, timeout_per_page=45):
    """Scan DynamoDB in chunks; timeout if a page takes too long."""
    all_items = []
    paginator = table.batch_get_item(
        RequestItems={table.name: {'Keys': ...}}
    )
    for page_num, page in enumerate(paginator):
        try:
            signal.alarm(timeout_per_page)
            items = page.get('Responses', {}).get(table.name, [])
            all_items.extend(items)
            signal.alarm(0)  # Cancel alarm
        except TimeoutError:
            logger.warning(f"DynamoDB page {page_num} timed out after {timeout_per_page}s; continuing")
            break
    return all_items

This prevents one slow page from blocking the entire rebuild. If the network is down, you get a partial result within 45 seconds instead of a 900-second hang.

3. Fix change detection to eliminate false alerts

The third bug was buried in the alert logic. When the script scanned DynamoDB, it compared the live result against the cached snapshot. Any difference triggered an SMS. But the snapshot persisted old, inactive records. When comparing:

def detect_changes(live_items, cached_snapshot):
    """Find genuine changes; ignore aged-out records."""
    # BUG: This compared ALL items, including those that aged out of cache
    live_ids = {item['event_id'] for item in live_items}
    cached_ids = {item['event_id'] for item in cached_snapshot}
    
    # WRONG: treat absence as deletion
    deleted = cached_ids - live_ids  # This included old, completed charters
    
    # FIXED: A cached event is "deleted" only if it was active and is now gone
    genuinely_deleted = {
        eid for eid in cached_ids 
        if eid not in live_ids and is_recent(eid)  # Only flag if recent
    }

The old logic treated every aged-out charter as "cancelled." When the rebuild finally ran after 30 days offline, the snapshot still held eight old events. The script flagged them all as cancellations and queued SMS: [ALERT] Crew release: event_2026-06-15-alice CANCELLED. These were not cancellations—they were old cached records. The fix: only flag genuine changes (state-column updates to ACTIVE/CANCELLED in the live table), not cache evictions.

Infrastructure Changes

AWS credentials and DynamoDB: The script authenticates via the AWS SDK, which looks for credentials in environment variables, ~/.aws/config, or IAM roles. LaunchAgent inherits none of these by default. Solution: explicitly load the credentials profile in the LaunchAgent plist before invoking the script.

DynamoDB table: jada-crew-dispatch — A single, append-only table where each record has a composite key event_id#timestamp. Scanning it costs money and time (scan throughput ~1 MB/s, your table ~3 MB). Chunking the scan with a page size of 50 items reduced single-page latency from unpredictable to ~2–5 seconds.

JSON snapshot: pages/last-events.json — Cached locally to avoid a full scan on every run. The snapshot now includes explicit timestamps and status fields so the diff logic can distinguish "this event is old" from "this event was just cancelled."

Key Decisions

  • Fail loudly on startup: If tools can't be found or credentials are missing, crash immediately with a clear error. Silent hangs are worse than loud failures.
  • Chunk external I/O: Any call to a remote service (AWS, Google APIs, etc.) should have a timeout and graceful fallback. A network hiccup should not crash the entire run.
  • Separate change detection from cache invalidation: Don't let snapshot expiration rules leak into the alert logic. If something is old, it's old—that's not news to the crew.
  • Test the launchd context, not just the shell: Tools work fine in the shell because your shell loads .bashrc, which sets PATH. LaunchAgent doesn't. Always test the actual execution environment.

What's Next

This fix handles today's crisis (no false SMS to crew before the Jul 4 charter). Longer term:

  • Add a heartbeat monitor: Every successful rebuild run should log a success marker. Anything older than 25 hours should page you.
  • Instrument the DynamoDB scan: Log the number of items, pages, and latency for each run. Trends will warn you if the table is growing or the network is degrading.
  • Separate credentials from code: Move AWS credentials out of plist files into a secure store (AWS Secrets Manager or an encrypted `.env` file). Rotate them regularly.

Silent failures are the most expensive kind. With these changes, your crew-page rebuild will either work or fail loudly within 60 seconds of launch.

```