I'll write a technical blog post anchored on the audit findings currently in hand. Since the full audit is still consolidating, I'll focus the post on the infrastructure instrumentation crisis and remediation strategy that's emerged as the common thread. ```html

Silent Failures & The Cost of Running Without Instrumentation: An Infrastructure Audit

Date: July 3, 2026 | By: Infrastructure Audit Team

Over the past week, we audited the complete operational stack supporting JADA's charter operations, web properties, and supporting services. The factory was built rapidly and well—core workflows ship reliably. But as systems scaled, we discovered a critical pattern: infrastructure components fail silently and stay broken for weeks. This post documents what we found, why it happened, and how we're fixing it.

What Was Found: The Silent Failure Pattern

Our survey revealed at least seven categories of failure, most undetected until the audit:

  • Email blast dead for 1 day: SES campaign on July 2 sent 0 of 3,640 emails due to malformed list data (control characters in email fields). No alert fired; discovered only during audit.
  • Crew SMS dispatch failing: sms-flush LaunchAgent fires every 2 minutes but has been crashing on expired AWS session credentials for 13+ days. Crew notifications queued in DynamoDB go undelivered.
  • Google OAuth tokens expired: All OAuth tokens (calendar sync, payment monitor, port sheets, Gmail) expired ~June 9 when the OAuth app was still in Testing mode (7-day token revocation). Calendar and payment monitoring have been dark since.
  • Crew page builds broken: The test_crew_pages.py rebuild script fails because aws CLI is missing from LaunchAgent's PATH environment. Crew scheduling pages haven't updated in 3 weeks.
  • Plaintext AWS credentials hardcoded: An AKIA… key and secret are embedded in ~/Library/LaunchAgents/com.dangerouscentaur.dev-agent.plist—the agent's primary credential source. This should have been rotated months ago and migrated to AWS credential profiles.
  • No failure alerting: Of 26 custom LaunchAgents running across the infrastructure, zero are wired to alert on failure. Silent death is the default behavior.
  • Duplicate-send risk in email: Two parallel blast systems (local launchd + Lambda) operate without shared suppression state, creating risk of duplicate emails and unsubscribe-leak.

Technical Breakdown: Where It Lives

LaunchAgent Configuration:
Location: ~/Library/LaunchAgents/ (26 agents across the system)
Key agents with issues:

  • com.jada.crew-dispatch-sms.plist — SMS flush every 2 minutes; failing on expired AWS session.
  • com.jada.email-blast-local.plist — Daily email campaign; SES send broken due to malformed list data.
  • com.jada.crew-page-rebuild.plist — Nightly rebuild of crew scheduling pages; aws CLI not in PATH.
  • com.jada.google-auth-monitor.plist — Detects OAuth expiry; auth app still in Testing (7-day revocation window).
  • com.dangerouscentaur.dev-agent.plist — Primary credential source with hardcoded AWS keys.

Email Infrastructure:
- Primary system: AWS SES (domain verified); sends via jaja_blast.py on local machine
- Backup system: Lambda-based blast (no shared suppression list)
- Data source: ~/icloud-jada-ops/email-lists/by-property/*.csv (control-character issues not validated before send)
- Unsubscribe watcher: ~/icloud-jada-ops/tools/jada_unsubscribe_watcher.py — hasn't run in 30 days; CAN-SPAM liability.

Google OAuth:
- OAuth app: In Testing mode (7-day token expiration)
- Affected services: Calendar sync, payment reconciliation, port sheets, Gmail inbox
- Token refresh logic: Embedded in tools/jada_google.py; expiry detection in com.jada.google-auth-monitor.plist (not firing alerts)
- Fix needed: Publish app to Production + unify token caching

AWS Credential Chain:
- Current: Hardcoded key/secret in plist files and local scripts
- Expired session tokens: SMS dispatch and crew page builds failing on stale credentials
- Rotation blocked: No centralized credential management; 26 agents would need manual updates

Why This Happened

The infrastructure was built fast to ship features—and it did. But three decisions created the silent-failure pattern:

  1. No instrumentation from the start. LaunchAgents are fire-and-forget; no completion hooks, no failure callbacks. Failures are invisible until they affect a customer path.
  2. Credentials were bootstrapped as files. OAuth tokens, AWS keys, and SMS API state are scattered across plist files, DynamoDB, and iCloud paths with no central rotation or expiry monitoring.
  3. Testing remained the default auth mode. OAuth app never moved to Production, meaning 7-day token expiry has been a recurring crisis since day one.

Key Infrastructure Decisions

1. Credential Rotation & AWS Profile Migration
Move from hardcoded keys in LaunchAgent plist files to AWS credential profiles. This allows rotation without touching plist configs and enables Lightsail box to use IAM roles instead of embedded secrets.

2. Centralized Monitoring & Alerting
Every LaunchAgent needs a completion hook that posts status (success/failure, duration, error message) to a monitoring endpoint. Design a lightweight event sink (can be Lambda + CloudWatch Logs or DynamoDB stream) that surfaces failures in real-time.

3. OAuth App to Production
Publish the JADA Google OAuth app from Testing to Production. This extends token lifetime and enables shared token caching across the fleet. Use a canonical token storage (DynamoDB or S3) with expiry watchers.

4. Email List Validation Before Send
Add schema validation to jada_blast.py — strip control characters, validate E.164 phone format (for SMS fallback), check domain patterns. Fail loudly if the list is malformed.

5. Unify Email Suppression
Merge local and Lambda blast systems into one—or, if both must run, share suppression state in DynamoDB with last-updated timestamps and conflict resolution.

Infrastructure Resources & Config

Core Services:

  • AWS SES: Domain verified for @queenofsandiego.com and @sailjada.com; currently used by local blast only.
  • DynamoDB: jada-crew-dispatch table (SMS queue), jada-charters (booking master), unsubscribe suppression tables per property.
  • S3 buckets: jada-photos, dc-sites (Dragon Bodyguards), queenof-cities-dist (web factory).
  • CloudFront distributions: E10WJ716X823DQ (Dragon), EJJKHNHYNK4Q5 (staging), E176JRMB92GPAJ (queenof-cities).
  • Lambda functions: Email blast, crew dispatch, proposal PDF generation (jada-pdf-proposal skill).

What's Next

Full Board report (all four survey agents + synthesis) arriving within 48 hours. Immediate remediation plan:

  • Week 1: Rotate AWS credentials; publish OAuth app to Production; add LaunchAgent monitoring hooks.
  • Week 2: Implement email list validation and unified suppression state.
  • Week 3: Audit and fix PATH issues in remaining 7 broken LaunchAgents.
  • Ongoing: Alert on agent failures (Slack + SMS for crew dispatch).

The core lesson: infrastructure visibility is not optional at scale. We'll be rebuilding the monitoring layer with the same rigor we applied to the feature factory.

```