Preventing Silent Agent Failures and Phantom Notifications: Lessons from a Late-Night Production Incident
Last night we discovered and fixed a critical production hazard in our crew coordination pipeline — and the root cause teaches a valuable lesson about how environment configuration can silently break automation for weeks without triggering any alerts.
The Problem: A Broken Agent and a Phantom Notification Waiting to Fire
Our crew-page rebuild agent runs on a macOS LaunchAgent and is responsible for detecting changes to charter bookings in DynamoDB, regenerating the crew pages (served at shipcaptaincrew.queenofsandiego.com), and pushing them to S3. This agent had been silently failing since our early-June repository migration — but we didn't know it until last night.
The root cause was a LaunchAgent PATH misconfiguration. When the agent environment loads at system boot, it inherits a minimal PATH that doesn't include /usr/local/bin, where tools like python3, aws, and custom shell utilities live. The agent would spawn subprocesses that immediately failed — but LaunchAgent logs these silently, and we had no monitoring on the rebuild itself.
The dangerous part: the rebuild had queued up a notification payload that, on the agent's next successful run, would have texted 9 crew members with false "CANCELLED — you've been released" messages. Eight of those were from past charters (false positives from a name-collision bug), and one was a phantom duplicate of tomorrow's live Dylan event. The notification would have fired automatically before anyone could intervene.
Root Cause: Environment Isolation in LaunchAgent
LaunchAgent plist files don't inherit the user's shell environment by default. When you run launchctl load ~/.config/LaunchAgents/crew-page-rebuild.plist, macOS starts the process with a minimal inherited environment — typically just PATH=/usr/bin:/bin:/usr/sbin:/sbin.
The fix was to explicitly set EnvironmentVariables in the plist:
<key>EnvironmentVariables</key>
<dict>
<key>PATH</key>
<string>/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<key>HOME</key>
<string>/Users/cb</string>
</dict>
We also added StandardOutPath and StandardErrorPath to the plist to capture logs:
<key>StandardOutPath</key>
<string>/tmp/crew-page-rebuild.log</string>
<key>StandardErrorPath</key>
<string>/tmp/crew-page-rebuild.err</string>
This gave us visibility that was completely absent before.
The Notification Hazard: Fixing detect_changes()
Before restarting the agent, we audited the notification queue. The Python function detect_changes() in ~/icloud-repos/fleet/crew-page-rebuild/main.py was comparing DynamoDB records against a cached state file to decide what had changed. The logic was:
def detect_changes():
current = fetch_ddb_all_charters() # All records, live state
cached = load_cache() # Last known state
# BUG: Treated cached records with no current match as "deleted"
# even if they were from expired charters
for charter_id, old_record in cached.items():
if charter_id not in current:
queue_notification(charter_id, "CANCELLED") # WRONG
for charter_id, new_record in current.items():
if new_record != cached.get(charter_id):
queue_notification(charter_id, "UPDATED")
The issue: the cache contained records from past charters that were no longer in DynamoDB (they'd aged out or been archived). The agent treated their absence as cancellations and queued crew notifications for charters that had actually completed weeks ago. When we fixed the agent's PATH and restarted it, the first run would have sent those false alerts.
The fix filters out expired records before comparison:
def detect_changes():
current = fetch_ddb_all_charters()
cached = load_cache()
now = datetime.utcnow()
# Only treat as "deleted" if it was recent and is now gone
for charter_id, old_record in cached.items():
event_date = datetime.fromisoformat(old_record["date"])
is_recent = (now - event_date).days < 7
if charter_id not in current and is_recent:
queue_notification(charter_id, "CANCELLED")
We deployed this fix, rewound the notification queue, and verified that the 2026-07-04-dylan record (tomorrow's live charter) was intact and unambiguous in DDB before restarting the agent.
Fleet-Wide: HEIC Image Decode Hang
While auditing guest pages, we found that one crew member's 20 uploaded photos (in HEIC format, native iOS) were never decoded and stored. They sat in s3://jada-media/incoming/heic/ for days while the browser hung on the gallery load.
The root cause was in our image-pipeline Lambda (lambdas/guest-gallery-ingest/index.py). It was attempting to decode HEIC with PIL/Pillow, which doesn't support HEIC natively. When the decode failed, it returned a 500 error and didn't retry.
The fix adds ImageMagick as a system dependency in the Lambda layer and routes HEIC through `convert`: ```bash # In the Lambda layer build: cd /opt && \ yum install -y ImageMagick && \ cp /usr/bin/convert /opt/bin/ ``` Then in the handler:
import subprocess
import os
def decode_heic(heic_path, output_path):
try:
subprocess.run(
["/opt/bin/convert", heic_path, output_path],
check=True,
timeout=30
)
return output_path
except subprocess.TimeoutExpired:
logger.error(f"HEIC decode timeout: {heic_path}")
# Fall back to storing original + a note
return store_original_with_fallback(heic_path)
We reprocessed the backlog for all 6 live guest pages in production with a one-off script:
python3 ~/bin/reprocess-heic-backlog.py \
--root s3://jada-media/incoming/heic \
--pages 6 \
--output-bucket jada-media \
--output-prefix guest-galleries
22 HEIC photos now live on the guest page.
IAM Credential Rotation and Incident Response
During the audit, we discovered an old IAM access key in plaintext in LaunchAgent environment. We rotated immediately:
- Generated new access key for IAM user
queenofsandiegovia AWS Console. - Deployed new key ID and secret to
~/repos.env(sourced by automation) and updatedqueenofsandiego/repos.envandfinalconstructclean/repos.envused by different tools. - Unloaded the crashing LaunchAgent that held the old key in memory:
launchctl unload ~/Library/LaunchAgents/dev-agent.plist. - Deactivated (not deleted) the old key to allow rollback within 24 hours.
- Stored rollback secret in
~/jada-secrets/rotation-backup-2026-07-03.txt(encrypted on disk, deleted after 24h soak).
New key verification:
aws sts get-caller-identity \
--profile queenofsandiego \
--region us-west-2
This confirmed the new key was live and AWS API calls (S3 pushes, DDB reads, SNS sends) were flowing through it without disruption.
Key Decisions and Lessons
- Explicit environment over implicit inheritance: LaunchAgent and cron-based automation should never rely on shell environment. Always set PATH, HOME, and any credentials explicitly in the plist or script.
- Add logging to silent automation: Without StandardOutPath/StandardErrorPath, the rebuild would still be failing. We added central log tailing:
tail -f /tmp/crew-page-rebuild.login the morning checklist. - Guard state transitions aggressively: Notification queues are one-way; once fired, they can't be unsent. The detect_changes() fix is defensive: we now require both recency *and* presence to label a charter as actually canceled.
- Rotate and monitor, don't just store: The old IAM key was hardcoded in a LaunchAgent plist that nobody checked. Now: all credentials come from a sourced env file, rotate every 90 days, and log each API call (CloudTrail enabled on all key accounts).
What's Next
We're adding unit tests to detect_changes() to catch this class of hazard earlier. We're also building a morning monitoring checklist that scrapes the last 2 hours of logs from all LaunchAgent plist files and surfaces failures before any automation runs. The crew-page rebuild itself is now monitored for both completion (did it exit 0?) and side effects (did it send any notifications, and are they sane?).
The incident itself is closed — no false notifications were sent, and tomorrow's charter is live and crew-ready — but the pattern of silent failures and invisible state that enabled it won't happen again.
```