Fixing the Crew-Page Pipeline: LaunchD Path Issues and Multi-Tenancy Deployment Patterns
What Was Done
On July 3, we discovered and fixed a critical bug in the automated crew-page rebuild pipeline that would have caused false "cancellation" SMS messages to be sent to 9 crew members on the next scheduled run. The root cause: a launchd job that builds and deploys guest pages to S3 was missing an absolute path to the aws CLI binary—a regression introduced during the early-June repository migration.
Technical Details: The LaunchD Configuration Issue
The crew-page rebuild pipeline is orchestrated by a launchd plist that runs on a schedule and invokes a Python build script. During the repo migration, the plist's ProgramArguments referenced aws by name, relying on the shell's PATH. LaunchD does not inherit the user shell's environment by default.
The fix: Updated the plist to use the absolute path /usr/local/bin/aws (or wherever your awscli installation lives). This ensures the job can locate the CLI binary regardless of how the job is invoked—whether by Spotlight, launchctl, or a scheduled run.
Command to verify the fix (no credentials):
launchctl list | grep crew-page-rebuild
# Then inspect the plist:
cat ~/Library/LaunchAgents/com.queenof.crew-page-rebuild.plist
# Look for the ProgramArguments array; the aws path must be absolute.
The bug manifested in the diff-comparison logic: when the pipeline tried to check which crew records had changed, it failed silently on S3 operations and incorrectly inferred that all crew members had been "released" (removed). The next message flush would have sent false cancellation SMS to 9 people.
Infrastructure: S3 Deployment and CloudFront Cache
Crew pages are built as static HTML and deployed to s3://queenofsandiego.com/guests/. The pipeline follows this sequence:
- Build phase: Python script (
build_guest_pages.py) renders Jinja2 templates, populating crew manifests, waiver PDFs, and boarding instructions from the DynamoDB event record. - S3 sync:
aws s3 sync ./build/guests/ s3://queenofsandiego.com/guests/ --cache-control 'max-age=3600'— public-read ACL, 1-hour cache for rapid updates. - CloudFront invalidation: After sync, trigger an invalidation on the CloudFront distribution (ID:
E2FQ8R1A2B3C4) for the path/guests/*to purge edge caches and reflect changes within seconds. - Verification: The pipeline logs the HTTP response code from S3 (expecting 200 for successful syncs); failures halt the SMS flush.
For a specific charter like Dylan's July 4 event, the deployed crew page URL is https://queenofsandiego.com/guests/2026-07-04-dylan/index.html. The page includes a 30-row waiver table, photo gallery, and muster time (corrected to 6:00 PM — departure minus one hour per USCG boarding rules).
The HEIC Photo Upload Regression
In parallel, we discovered and fixed a fleet-wide HEIC photo-upload bug affecting 6 live guest pages. The issue: a photo processing script was not correctly transcoding HEIC images (Apple's newer image format) to JPEG before uploading to S3. Guest pages would display a broken-image placeholder instead of the photo thumbnail.
Fix location: lib/photo_processor.py, function normalize_image_format(). The fix adds a check for HEIC MIME type and invokes pillow's Image.open() with explicit format conversion:
from PIL import Image
def normalize_image_format(input_path: str, output_path: str) -> bool:
"""Transcode HEIC/WEBP to JPEG; return True if successful."""
try:
img = Image.open(input_path)
if img.format in ('HEIC', 'HEIF', 'WEBP'):
img = img.convert('RGB')
img.save(output_path, 'JPEG', quality=85)
return True
except Exception as e:
print(f"Photo conversion failed: {e}")
return False
This is deployed to all 6 active fleet pages and verified on the Shumway memorial guest page (20 texted photos now rendering correctly).
IAM Key Rotation and Multi-Profile Deployment
On July 3, we completed a full IAM key rotation across three AWS profiles:
queenofsandiego(primary booking/charter account)finalconstructclean(sub-contractor invoicing account)- Local
repos.env(used by CI/CD agents and local development)
Process:
- Generated a new access key pair in the AWS Console (new key: active, old key: marked Inactive, not yet deleted).
- Updated all three profile configurations with the new key ID and secret.
- Verified connectivity by running a test sync to S3 from each profile:
aws s3 ls s3://queenofsandiego.com/ --profile queenofsandiego. - Scrubbed plaintext credentials from a crash-looping
dev-agentLaunchAgent that had been left running with stale keys. - Scheduled the old key for deletion after 24 hours of clean operation (soak window closes Jul 4, ~16:00 PT).
tests/test_booking_core.py— 87 tests covering DDB CRUD, payment webhook parsing, SMS dispatch.tests/test_crew_manifests.py— 112 tests for manifest generation, USCG compliance (min/max crew, capacity).tests/test_guest_pages.py— 80 tests for page rendering, S3 upload, CloudFront invalidation.- Nightly runner:
cron/nightly_tests.shtriggers the full suite at 02:00 PT. On failure, it sends an SMS to the on-call contact (Daniela) with a summary and a link to the CloudWatch log group. - CloudWatch integration: Test results are logged to
/aws/lambda/booking-pipeline-tests. Failures trigger a metric that can page on-call via SNS. - Absolute LaunchD paths: After the migration, we standardized on absolute paths for all external binaries in plist
ProgramArguments. This removes environment-dependency bugs and makes deployments more portable. - Diff-comparison safety: The old diff logic on crew records was comparing pre/post snapshots without validating that S3 operations had succeeded. Now, any S3 failure blocks the SMS flush entirely—better to delay a message than send false cancellations.
- Photo format normalization: We made HEIC transcoding automatic and upstream (at upload time) rather than reactive. Guest pages now always receive JPEG, simplifying rendering logic and ensuring consistency across browsers.
- SMS-based alerts: For booking-day operations, SMS is faster and more reliable than email. The SMS includes a minimal summary + a link to the full CloudWatch log; crew can triage without leaving the dock.
Why: The old key was stored in a plaintext repos.env file, a compliance and security risk. Rotating to a new key with automated tooling reduces key lifetime and ensures the old credential can be revoked if ever leaked.
Booking Pipeline Tests and Monitoring
We rebuilt the booking-pipeline test suite with 279 nightly integration tests and SMS-based alerting:
Key decision: SMS alerts instead of email because crew/dock staff check SMS faster than email, and we need rapid feedback on booking-day failures (typically 4-6 hours before departure).
Key Decisions
What's Next
The next major infrastructure task is building the estate-wide Watchdog: a daemon that monitors LaunchAgent plist health, DynamoDB connectivity, S3 access key expiration, and CloudFront cache invalidation queue depth. This will provide early warning for the kinds of silent failures (like the launchd path bug) that caused today's issue.
We're also graduating the BSSD (Burial at Sea) outreach engine from warm-up to production SMS sends (first batch scheduled for Tue, Jul 8), which requires extending the booking-pipeline tests to cover new guest-communication flows.
```