Preventing Ghost Cancellations: How a Rebuild Bug Nearly Texted Phantom Charter Updates
The Problem: A Silent Infrastructure Failure
When launchd agents fail silently for weeks, the damage often emerges only when you try to fix them. On July 3, 2026, we discovered that the crew-page rebuild system had been broken since the early June repository migration—but the real danger surfaced when we attempted to repair it.
The crew pages at queenofsandiego.com dynamically generate waiver links, trip sheets, and guest confirmations for each charter. When a charter is created or updated in DynamoDB, a launchd agent should detect those changes and rebuild the static HTML pages deployed to S3/CloudFront. For six weeks, this agent had been failing due to a PATH environment variable bug—so no pages were being regenerated, yet no alerts had fired. The system appeared to work because old pages were still live.
The Critical Discovery: False Positives in Change Detection
When we fixed the PATH issue and ran the agent for the first time, we almost sent 9 false "CANCELLED — you've been released" SMS messages to crew members. The root cause: detect_changes() in the rebuild daemon was comparing current DynamoDB records against stale local cache files. After six weeks of breakage, the cache was catastrophically out of sync with reality.
The function looked like this (simplified):
def detect_changes(ddb_records):
cached = load_cache('crew_charter_state.json')
changes = []
for charter_id, record in ddb_records.items():
if charter_id not in cached or cached[charter_id]['status'] != record['status']:
changes.append((charter_id, record))
return changes
The bug: a charter that existed in DDB but was missing from the 6-week-old cache would be interpreted as a "new event with status=CANCELLED" (the default zero-value). The rebuild agent would then queue SMS notifications to tell crew they'd been released from these "cancelled" charters. One duplicate Dylan evening charter phantom record from an earlier session made it through as a ninth false positive.
The Fix: Validation Before Notification
We added a guard condition in detect_changes():
def detect_changes(ddb_records):
cached = load_cache('crew_charter_state.json')
changes = []
for charter_id, record in ddb_records.items():
# Skip records with no cached entry unless they're new (created_at recent)
if charter_id not in cached:
if (datetime.now() - record['created_at']).seconds > 3600:
continue # Stale phantom record; don't trigger rebuild
if cached[charter_id]['status'] != record['status']:
changes.append((charter_id, record))
# Clear cache to prevent future misalignment
cache_from_ddb(ddb_records)
return changes
The logic: only treat a DDB record as "changed" if (1) it's in the cache and has a status mismatch, or (2) it's genuinely new (created in the last hour). Old phantom records are skipped. After validation, we rebuild the cache atomically from the canonical DDB source to prevent divergence.
Infrastructure and Deployment Details
File Structure:
/Users/cb/dablio/crew_rebuild_agent.py— launchd daemon entry point/Users/cb/dablio/core.py— containsdetect_changes()and DDB query logic~/.config/launchd/com.jada.crew-pages-rebuild.plist— launchd configuration (runs every 15 minutes)- S3 bucket:
queenofsandiego-static(region: us-west-2, versioning enabled) - CloudFront:
d1a2b3c4d5e6f7.cloudfront.netdistributesprint/and guest pages - DynamoDB: table
charterswith partition keycharter_idand GSI onstatus
Path Bug Root Cause: The June migration moved the repo from ~/dablio to /Users/cb/dablio, but the launchd plist still referenced Python via a relative PATH. When launchd spawned the agent in a minimal environment (no user shell initialization), Python 3.11 wasn't found. The fix was explicit in the plist:
<key>ProgramArguments</key>
<array>
<string>/usr/local/bin/python3</string>
<string>/Users/cb/dablio/crew_rebuild_agent.py</string>
</array>
Operational Context: The Dylan Charter July 4 Window
This discovery happened during preparation for a July 4 charter (organizer: Dylan Osborne, 1 captain + crew unresolved). The timing was critical because a false cancellation SMS during crew confirmation could have caused:
- Crew no-shows (confusion about release status)
- Last-minute staffing scrambles 24 hours before departure
- SMS spam penalties and reputation damage
- Trip liability (insufficient crew manifest)
We validated the entire crew messaging queue and manually verified that the Dylan record was intact in DDB before allowing any automated notifications to resume.
Key Decisions
- Why not rebuild the cache file immediately? Because we needed to audit which records were phantom vs. real. A full rebuild would have made the false positives "real" from launchd's perspective.
- Why add creation-time validation instead of just "skip unknowns"? Because a legitimate new charter should trigger a rebuild. We needed a heuristic to distinguish new-and-valid from old-and-stale.
- Why not just increase the rebuild interval? The 15-minute cycle is load-bearing for real-time guest confirmations. Longer intervals risk stale pages.
Testing and Verification
Before re-enabling automatic crew notifications:
- Queried DDB directly to verify Dylan record status and crew_manifest array
- Spot-checked 5 random past charters in both DDB and S3 to confirm sync
- Ran the rebuilt
detect_changes()function offline against prod data and confirmed 0 phantom changes - Restarted launchd agent and monitored syslog for errors over 2 cycles
- Manually triggered a crew message for Dylan and verified SMS delivery to Daniela (+1-442-258-0072)
What's Next
The Dylan charter crew situation remains unresolved (crew assignments TBD for July 4). Immediate priorities:
- Crew muster confirmations must be locked in by July 4 6:00 AM
- Waiver pages and trip sheets are live and verified
- SMS and email notification queue is safe to use for future charters