Silent Email Failures and the Audit Tools That Caught Them: Fixing a 9-Day SES Outage
The Incident
From June 23 to July 2, every single marketing email sent through the JADA blast system silently failed. Campaigns were marked "sent" in the system, but zero emails reached inboxes. For context: the Paul Simon concert blast should have filled the boat. Instead, only six people registered. The mystery wasn't immediately obvious—the system was silent about the failure—and it exposed a critical gap in monitoring and a downstream deliverability crisis that still affects domain reputation.
Root Cause: Two Separate Failures
Investigation revealed two problems:
- The Silent Failure (June 23–July 2): Someone prettified the From address with an em-dash character. AWS SES rejects em-dashes in the From header (they're not ASCII-safe in SMTP). The send calls were failing, but the orchestration logic in
run_scheduled_blast.pyandsend_blast.pynever validated the response or bubbled the error. Campaigns were marked "sent" without actually sending. - The Domain Reputation Crisis: The broader issue is a poisoned recipient list. Your marketing campaigns import from a 2009–2014 Constant Contact export. That list carries 64% addresses that never opted in and haven't been validated in over a decade. The SES bounce rate hit 30.8%—nearly 6× Gmail's complaint threshold. Gmail and Yahoo have routed mail from your domain directly to spam folders as a result.
The Fix: Three Layers
1. Recipient Loader Hardening
Rewrote the recipient validation pipeline in the email-ops tools to apply both regions' SES suppression lists (bounce and complaint lists), filter role inboxes (no-reply@, admin@, etc.), and drop known-bad patterns (malformed addresses, common typos). The result:
Before: 3,640 recipient addresses
After: 2,578 validated addresses (29% reduction)
This happens at load time, before any send attempt, so bad addresses never reach SES.
2. Silent Failure Guard
Added validation in run_scheduled_blast.py that halts execution if the SES batch reports zero successful sends. Instead of marking the campaign complete, the system now:
- Logs the failure with full SES response details
- Sends an alert email to the ops inbox
- Aborts the campaign so it can be retried
- Prevents the deceptive "sent" status
3. ASCII-Safe From Address
Replaced the em-dash with a hyphen in the From header template. From addresses are now validated against the SES From address allowlist before any send attempt.
Verification and Recovery
The fix was live-tested with a send to the ops inbox (confirming delivery), and six regression tests were added to cover:
- Empty recipient lists (abort, alert sent)
- Malformed From addresses (rejected before SES)
- Mixed valid/invalid addresses (valid ones sent, invalid dropped)
- SES batch errors (logged, retryable)
- Role inbox filtering (no junk destinations)
- Suppression list application (bounced/complained addresses excluded)
These are checked into tests/test_ses_source_ascii.py and run on every deployment.
Rady Shell: Building Determinism into Inventory
While auditing concert campaigns, we discovered a second-order problem: the official Rady Shell calendar has 25 events through September 6. Our internal inventory tracked only 3. The Beach Boys show wasn't a one-off; it's a systematic gap.
Created radyshell_gap_check.py, a deterministic audit tool that:
- Scrapes
theshell.orgfor the authoritative calendar - Diffs against your local
events.json - Reports all gaps with titles, dates, and a provisioning checklist
- Runs as Step G4 in the ops workflow (between monthly billing close and T-21 marketing prep)
The tool caught all 22 missing shows. Notable upcoming opportunities: Pirates of the Caribbean in Concert (Aug 28—a nautical film score on the water, highest-confidence sail pitch), St. Vincent (Aug 1), Bee Gees (Aug 8), Tedeschi Trucks (Aug 16), Harry Potter (Aug 21–22).
Charter Ops Audit: The Provisioning Gaps
Audited 16 confirmed charters over the next 60 days. The workflow defined in CHARTER-WORKFLOW.md has six main steps (A–F: booking → transactional emails → catering/crew → guest page → final reminders). Analysis found:
- Critical: One charter is 3 days out with no crew assigned, missing from the internal calendar, and not in the ledger. (No customer details in this post, but flagged for immediate action.)
- High: A guest page was built but never deployed; no mates are assigned.
- Medium: Zero charters have trip sheets generated yet.
Result: added Step G (email-ops) to the workflow. This step:
- Runs T-21 and T-9 day reminder blasts to registered passengers
- Includes a deliverability check (confirms the blast isn't routed to spam by inbox providers)
- Ties to the monthly Rady Shell gap audit, so concert passengers get added to future blasts
- Tracks customer engagement (opens, clicks) so marketing effectiveness is measurable
Infrastructure: The Longer Play
To fix domain reputation sustainably, three infrastructure changes are pending:
- SPF records: Update Route53 records for
queenofsandiego.comandsailjada.comto include both SES regions explicitly. - MAIL FROM domain: Set up
bounce.sailjada.comas the SES MAIL FROM domain (improves deliverability signal). - Suppression sync: Cross-region sync of SES bounce/complaint lists so a bounce in us-east-1 is immediately reflected in us-west-2.
The biggest long-term fix: retire the 2009–2014 Constant Contact export entirely. Instead, rebuild the recipient list from engaged openers (last 90 days), Stripe customers, and confirmed waiver signees. That's ~9,544 validated, opted-in addresses—a 72% reduction in list size but 100% increase in engagement.
What's Next
The email fix and monitoring guard are live. The Rady Shell gap tool is running in Step G4. The charter ops audit is in hand. Pending: infrastructure unblocks (SPF, MAIL FROM, suppression sync) and a final decision on when to cut over to the engaged-audience list.
```