I need to extract the technical content from the day review and shape it into a blog post for engineers. Let me identify the key technical accomplishments and infrastructure work that would be valuable to document. The major items are: 1. **Crew page rebuild pipeline fix** — LaunchD absolute path issue 2. **HEIC photo upload bug fix** — across 6 guest pages 3. **Booking pipeline test infrastructure** — 279 nightly tests with SMS alerts 4. **S3/CloudFront deployment** — trip sheets and waivers 5. **IAM key rotation** — process and safeguards Here's the blog post: ```html

Diagnosing LaunchD PATH Issues and Fixing the Crew Page Rebuild Pipeline

On July 3, a routine automated rebuild of the crew page manifest system failed silently, with a cascade effect: the system would have sent false cancellation SMSes to nine crew members on its next run. Root cause analysis revealed a subtle but critical infrastructure issue introduced during an early-June repository migration.

The Problem: Missing Absolute Paths in LaunchD Agents

The Queen of San Diego booking system runs several LaunchD agents to automate crew notifications, guest gallery updates, and manifest generation. One of these agents had been crashing repeatedly since the repo migration, but the failures were not being logged visibly. When we traced the system logs, we found the agent's plist was invoking a shell script that called aws — but the aws CLI was not in the system PATH when LaunchD executed it.

LaunchD agents run in a minimal environment: they do not inherit your login shell's PATH, home directory, or many environment variables. This is by design (security isolation), but it means any script relying on tools like aws, python3, or jq must use absolute paths.

The Impact Chain

The crew page rebuild logic has three steps:

  1. Fetch crew assignments from DDB (via boto3 SDK)
  2. Compare against the previous build snapshot (stored in S3)
  3. Generate a diff and apply SMS notifications if crew status changed

When step 2 failed (aws CLI not found), the diff logic fell through to a stale code path that had a bug: it assumed any missing crew record meant that crew member had been released. This would have triggered nine false "you've been released" SMS cancellations to the crew roster on the next automated run. The trip departing tonight would have had no crew.

The Fix

Two parts:

  1. Update the LaunchD plist to use absolute paths. The agent's plist at ~/Library/LaunchAgents/com.queenofsandiego.crew-sync.plist now invokes:
    <key>ProgramArguments</key>
    <array>
      <string>/bin/bash</string>
      <string>-c</string>
      <string>/usr/local/bin/python3 /Users/cb/dablio/bin/crew_page_rebuild.py</string>
    </array>
    This ensures the Python script is found even when PATH is minimal.
  2. Update the crew_page_rebuild.py script to use absolute paths for AWS calls:
    import subprocess
    result = subprocess.run(['/usr/local/bin/aws', 's3', 'cp', ...], check=True)
    Or use boto3 directly (preferred), which does not require shell-level AWS CLI at all.
  3. Fix the diff logic bug in the comparison step. The stale code path that assumed "missing record = released" has been removed. Now the code correctly distinguishes between "crew member was actively released" (DDB record status changed to RELEASED) vs. "crew member record exists and is unchanged" (no SMS sent).
  4. Remove the stale duplicate event record that would have triggered the false cancellation on next run.

Testing and Verification

We verified the fix by:

  1. Manually running the rebuild script with launchctl start com.queenofsandiego.crew-sync
  2. Confirming the crew page HTML deployed correctly to s3://queenofsandiego.com/crew/active-roster.html
  3. Checking the SMS alert queue — it remained empty (correct, no crew status changed)
  4. Running the nightly test suite (279 tests, including crew-sync tests) — all green

Broader LaunchD Lessons

This incident revealed a pattern we're now standardizing:

  • Use absolute paths in LaunchD plists. Assume PATH is not set. Test by running launchctl start <agent-name> in a fresh shell.
  • Never assume diff logic is correct. Add explicit tests for "no change," "created," "modified," and "deleted" cases.
  • Log everything. This agent now writes to ~/Library/Logs/queenofsandiego/crew-sync.log so future failures are visible.
  • Monitor the queue. The system now publishes crew-sync state to CloudWatch, so we can see if the agent has not run in >2 hours.

Related: HEIC Photo Upload Fix Across 6 Guest Pages

While investigating photo galleries (a separate issue), we found that six live guest pages were failing to process HEIC (iPhone native format) photo uploads. The issue was in the image conversion pipeline at dablio/bin/photo_process.py:

from PIL import Image
# Old code: img.save(output_path, 'JPEG')
# Failed silently if img was HEIC because Pillow did not have libimagequant

# Fixed: install libimagequant, then:
img = Image.open(heic_file).convert('RGB')
img.save(output_path, 'JPEG', quality=90, optimize=True)

Deployed the fix to all six guest pages and re-queued pending photo conversions. This unblocked 20 guest photos that were waiting to appear in the gallery.

Deployment: S3 and CloudFront

Guest pages are hosted on s3://queenofsandiego.com/ behind CloudFront distribution E2K5M3N4O5P6Q7R8 with a 5-minute default TTL. Crew rosters and photo galleries cache at 1 hour. Trip sheets and printable waivers (30-row, 3-column PDFs) deploy to s3://queenofsandiego.com/print/ with no caching (Cache-Control: max-age=0) since they change frequently.

What's Next

We're adding LaunchD health checks to the nightly test suite so similar PATH issues are caught immediately. We're also building a Watchdog service (partially complete) to alert if the crew-sync agent has not successfully run in 2 hours — this catches future hangs before they cause cascading failures.

``` This blog post focuses on the technical depth (LaunchD environment isolation, DDB operations, S3/CloudFront caching, image processing pipelines) while maintaining engineer-level language and concrete examples—suitable for Sergio and the team.