Silent Failures in Charter Ops: Architectural Gaps Discovered in the July 2026 JADA Audit
What Was Done
In early July 2026, we ran a comprehensive operational audit across JADA's charter systems, web properties, and infrastructure layer. The audit spanned four parallel survey tracks: charter operations (crew dispatch, manifests, proposals, ledger), web services (seven properties across CloudFront, S3, Route53), side ventures (QuickDumpNow, 2035ce, Lightsail agent box), and the automation/LaunchAgent layer. This post focuses on the charter ops findings, which surfaced critical silent failures that had gone undetected for weeks—and revealed gaps in our monitoring and deployment architecture.
Critical Operational Findings
The audit discovered three categories of problems:
Immediate Operational Risk
- Missing USCG Manifest (July 4, 2026 charter): The
~/icloud-jada-ops/passengers/manifests/directory was empty. The generator script (tools/generate_manifest.py) was ready, but the upstream passenger name capture (Step D2 in CHARTER-WORKFLOW.md) had no automation. Dylan Osborne's 30-guest charter had no way to produce the mandatory two-copy manifest for USCG compliance without manual intervention. - Crew Page Build Pipeline Broken: The crew rebuild script at
~/bin/jada-deployfailed silently because the AWS CLI was not in the subprocess PATH when invoked from launchd. The script uploads the dist to S3 bucketdc-sitesand invalidates CloudFront distributionE1P4PVXN8FJ07S; without AWS credentials in the environment, the entire rebuild chain stalled. - Payment Tracking Gap: $3,250 of $3,750 remained unpaid (40 hours before departure), with no automated reminder system wired into the ledger or GetMyBoat integration.
Silent Monitoring Failures (30+ Days)
- Unsubscribe Watcher Broken: The CAN-SPAM compliance watcher at
~/icloud-jada-ops/tools/unsubscribe_watcher.pyhad not run in 30 days. Email suppression lists in~/icloud-jada-ops/email-lists/suppression/(per-property CSV files) were not being updated from Gmail unsubscribe headers. Blast emails sent via local launchd (~/icloud-jada-ops/tools/jada_blast.py) use the suppression list in memory; stale data meant at-risk re-marketing to opted-out contacts. - Zelle Payment Watcher Blind: The payment watcher at
~/icloud-jada-ops/tools/zelle_watcher.pyrequires Full Disk Access to read the macOS chat.db (Messages app database). This permission had not been granted; the tool ran but produced no data, leaving manual deposit reconciliation as the only source of truth. - Test Suite Silent Failure: The nightly integration test suite (entry point:
tools/test_crew_pages.py) had been failing for three weeks. The test gate gates every crew page rebuild; failures should have blocked deployments, but no alerting was configured.
Data Integrity Issues
- Incomplete Ledger: The financial ledger at
~/icloud-jada-ops/ledger.jsoncontained only 7 entries, with 3 missing calculated totals. The ledger is the source of truth for board-confirmed revenue tracking; gaps here undermine both reconciliation and strategic visibility. - Privacy Exposure: Raw SMS dumps from Carole (approximately 306KB) were syncing to iCloud Drive daily via a launchd plist. These contained unstructured passenger data and payment info with no retention policy or encryption wrapper.
Root Causes: Architectural Patterns That Failed
Deployment Without PATH Hardening
The crew page rebuild uses subprocess.run(["aws", "s3", "sync", ...]) invoked from launchd. When launchd spawns a process, it inherits a minimal environment—PATH does not include /usr/local/bin (where Homebrew installs aws-cli). The script worked in manual testing (full shell environment) but failed in the scheduled context. Solution: explicitly set PATH in the launchd plist and use absolute paths for all CLI tools.
Monitoring Without Alerting
Test failures existed in logs (CloudWatch for Lambda tests, local launchd logs for the crew script). No alerting mechanism examined these logs; failures were silent. Health status is uploaded to S3 at s3://status.queenofsandiego.com/crew-health/ (CloudFront dist E1P4PVXN8FJ07S), but no dashboard or digest aggregated failures into a single view.
Permission Gates Without Ownership
Tools requiring system-level permissions (Full Disk Access for Zelle watcher, keychain access for Google auth, LaunchAgent placement) were documented but not wired into the onboarding. The Zelle watcher was built but non-functional for weeks because the permission gate was a one-time manual click only CB could grant—and the blocking dependency was never surfaced.
Infrastructure Details
Crew Deployment Pipeline
Tool: ~/bin/jada-deploy
- Calls: python3 ~/icloud-jada-ops/tools/crew_pages_rebuild.py
- Outputs crew page HTML to: ~/icloud-repos/sites/queenofsandiego/crew/
- AWS Sync: aws s3 sync ./crew/ s3://dc-sites/crew/ --no-progress
- CloudFront Invalidation: aws cloudfront create-invalidation \
--distribution-id E1P4PVXN8FJ07S --paths "/crew/*"
- No-cache paths: /print/*, /g/* (CachingDisabled in dist config)
- Monitoring: test_crew_pages.py validates HTML output before sync
Email Compliance Stack
Blast Path:
- Template: ~/icloud-jada-ops/email-lists/blast_template.txt
- Suppression: ~/icloud-jada-ops/email-lists/suppression/{slug}.csv
- Tool: ~/icloud-jada-ops/tools/jaja_blast.py
- Watcher: ~/icloud-jada-ops/tools/unsubscribe_watcher.py (reads Gmail unsubscribe headers)
- Sync: LaunchAgent com.jada.unsubscribe-watcher (plist at ~/Library/LaunchAgents/)
Payment Watchers:
- Zelle: ~/icloud-jada-ops/tools/zelle_watcher.py (reads ~/Library/Messages/chat.db)
- Stripe: integrated into Lambda backend (shipcaptaincrew /api/webhook)
Key Decisions
Why These Gaps Existed
Deterministic scripts without observability: Our infrastructure philosophy (see ICM default working method in memory) is to build once, trigger on events—not re-improvise flows with LLM tokens. We succeeded at building (BSSD blog publisher, CHARTER-WORKFLOW.md, proposals pipeline all solid), but observability was treated as a post-launch concern. Failed watchers, broken test gates, and silent launchd failures all share the same root: no metrics, no alerting, no health dashboard.
Permission gates without escalation: We correctly identified system permission requirements (Full Disk Access, keychain, LaunchAgent placement) but documented them without building a dependency tracker. When a permission gate blocked a tool, the blocker was invisible to automated systems—it required CB to notice the tool was broken.
Ledger as manual process: The ledger is deterministic (append-only JSON) but entry creation was manual: `ledger.py add
What's Next
- Fix July 4 charter: Implement Step D2 automation to capture passenger names at confirmation; generate manifest from DDB (shipcaptaincrew, Lambda-backed).
- PATH hardening in launchd: Update all plist files to export
PATH=/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbinand use absolute paths (/usr/local/bin/aws) in script invocations. - Alerting layer: Aggregate health status (crew page tests, watcher logs, ledger completeness) into a single dashboard; wire CloudWatch alarms for test failures.
- Permission dependency tracking: Encode required permissions in tool metadata; report unmet permissions at startup.
- Ledger automation: Wire DDB charter records → automatic ledger entry creation on payment confirmation.
The audit revealed that our deployment and ops infrastructure is correct in design (one-time scripts, event-triggered, minimal re-execution) but missing the observability layer that should catch breakage before it reaches guests. The fixes are architectural, not tactical: add metrics, add alerting, add dependency tracking. The next pass will convene the full Board to map these fixes against the web properties and side-ventures audits.
```