Building an Automated Email Outreach Engine for BSSD: From Harvesting to Scheduled Delivery
What We Built
We constructed an end-to-end email outreach automation system for burialsatseasandiego.com that harvests contact information from 76+ funeral homes, churches, hospices, and cremation societies across San Diego County, validates emails, deduplicates across sources, and delivers personalized outreach via AWS SES on a scheduled cadence. The system runs daily via macOS LaunchAgent and includes comprehensive test coverage, suppression list management, and deployment gates.
Architecture Overview
The system follows a three-stage pipeline:
- Harvest Stage:
harvest_emails.pyscrapes contact information from websites, handles Cloudflare email obfuscation, follows contact-link redirects, and validates MX records - Merge Stage:
merge_found_emails.pydeduplicates across multiple email-finder agents and consolidates results into canonical CSVs - Send Stage:
send_bssd_outreach.pyconstructs personalized messages, respects suppression lists, and delivers via AWS SES
Each stage is independently testable and can run on demand or scheduled. LaunchAgent triggers the full pipeline daily via /Users/cb/Library/LaunchAgents/com.jada.bssd-outreach.plist.
Email Harvester Implementation
The harvester at /Users/cb/icloud-jada-ops/bssd-crm/harvest_emails.py solves a non-trivial problem: funeral home websites protect emails aggressively. Our solution:
- HTTP fallbacks: Try HEAD request first (fast), fall back to GET with accept-encoding bypass if redirected
- Cloudflare obfuscation decoding: Parse
data-cfemailattributes directly instead of relying on JavaScript rendering - Contact-link following: Detect contact/mailto links in HTML, follow them recursively with depth limit
- MX validation: Query DNS for MX records on harvested domains to filter garbage addresses
- Status tracking: Return HTTP status for each URL (200, 403 blocked, 401 auth required, etc.) to diagnose site blocking patterns
The harvester runs against a CSV of target organizations (columns: name, type, website, city) and outputs email addresses with confidence scores and source attribution. This decouples finding logic from data — we can re-run against old datasets to catch emails missed on first pass.
Deduplication & Merge Strategy
We run parallel email-finder agents on chunked datasets (chunk-aa, chunk-ab, etc.). Each agent discovers emails independently, then merge_found_emails.py consolidates results into canonical CSVs like /Users/cb/icloud-jada-ops/bssd-crm/funeral-homes.csv and churches.csv. The merge function:
- Preserves existing rows and contact info
- Appends new emails to the
emailscolumn as semicolon-separated values - Logs which finder discovered which address for audit trails
- Respects suppression lists (see below)
This approach lets us safely re-run finders without losing prior work or duplicating effort.
Outreach Sender & Test Suite
The sender at /Users/cb/icloud-jada-ops/bssd-crm/send_bssd_outreach.py constructs and delivers emails. The test suite at /Users/cb/icloud-repos/sites/burialsatseasandiego.com/tests/test_bssd_outreach.py validates:
- Template rendering: Verify personalization tokens (business name, contact name) substitute correctly
- Suppression enforcement: Confirm emails in
/Users/cb/icloud-jada-ops/email-lists/suppression/burialsatseasandiego.csvare never sent to (prior bounces, opt-outs, competitors) - CSV parsing: Handle missing columns gracefully; skip rows lacking email or name
- SES behavior: Dry-run mode constructs all messages without calling AWS, allowing safe testing against live CSVs
All 11 tests pass. The sender supports a --dry-run flag for preview mode and --csv-path to target specific datasets. Emails are constructed with subject line, body copy personalized to business type (funeral home vs. church vs. hospice), and BSSD footer with GA-tagged link for attribution.
Infrastructure & Scheduling
LaunchAgent Configuration: /Users/cb/Library/LaunchAgents/com.jada.bssd-outreach.plist runs the full pipeline daily. LaunchAgent provides reliability (restart on failure), logging to ~/Library/Logs, and tight macOS integration. We followed the hotel outreach agent pattern, which is proven in production.
AWS SES & Email Verification: Verified burialsatseasandiego.com identity in SES. SendGrid was considered but SES has lower cost at scale and tighter AWS integration. All sending happens via boto3 with IAM role.
S3 & Deployment: Homepage updates (adding outreach footer link) deploy to S3 bucket burialsatseasandiego.com via jada-deploy tool. Changes flow through staging for verification before prod push. CloudFront distribution caches homepage; invalidations handled via deploy script.
Email List Management: Contact CSVs live in /Users/cb/icloud-jada-ops/email-lists/ with README documenting format and source. Suppression list at burialsatseasandiego.csv tracks bounces, opted-out addresses, and known non-targets (competitors, internal). This is read by the sender on every run.
Key Decisions & Rationale
- Parallel discovery, centralized merge: Running independent finders on chunked data maximizes coverage and handles site-blocking gracefully. Merge consolidation ensures deduplication and audit trails without re-scanning. Cost: added merge complexity; benefit: resilient discovery.
- Cloudflare decoding in the harvester: We could wait for JavaScript rendering, but that adds latency and requires a headless browser. Direct
data-cfemailparsing handles ~80% of protected addresses instantly. Fallback to Selenium for remaining cases. - MX validation: Some harvested addresses are honeypots or typos. DNS MX queries filter those out cheaply before sending. Reduces bounce rate, improves SES sender reputation.
- LaunchAgent over cron: Cron works, but LaunchAgent provides restart-on-failure and better logging for a macOS-based deployment. If scaling to servers, this moves to Lambda + EventBridge or a container orchestrator.
- Test-first sender: We wrote tests before implementation. This clarified CSV format requirements, suppression logic, and dry-run behavior. Tests now prevent regressions as we add new contact categories (celebrants, grief counselors).
Metrics & Results
First harvest run across 76 funeral homes yielded 5 direct email hits; HTTP status sweep revealed 11 sites block automated access, 23 use contact forms only, 37 have emails but obscured. Parallel finders across church/hospice dataset discovered 31 organizations with 24 email addresses (77% coverage). Suppression list filters ~8 addresses per run (prior bounces, opt-outs). Daily sends currently 15–20 emails to verified contacts.
What's Next
- Expand contact categories: Integrate celebrant and grief-support organization data from recent research; merge into outreach flow
- A/B testing: Segment outreach by organization type; test subject lines and body copy variants; track clicks via GA-tagged links
- Bounce handling: Auto-add bounces to suppression list; track soft vs. hard bounces separately
- Response automation: Build auto-responder for common questions (pricing, service area, booking flow)
- Scale to multi-region: Harvest contacts from other coastal cities; adapt templates for regional variation