```html

Fixing Legacy URL Bleed: Automated Parity Page Generation and CloudFront Cache Invalidation at Scale

What Was Done

An Instagram link-in-bio campaign was driving traffic to https://sailjada.com/charter-options/, which returned a 404 despite the URL being legitimate in our legacy site structure. This uncovered a systemic problem: 236 of 250 historically indexed URLs were broken, representing lost SEO equity and revenue leakage.

We implemented a three-phase remediation:

  • Phase 1 (Immediate): Restored the specific /charter-options/ page using Wayback Machine snapshots and deployed to production within the staging approval window.
  • Phase 2 (Tier 1): Automated generation of 48 parity pages across 7 content categories, deployed to S3 with CloudFront invalidation.
  • Phase 3 (Foundation): Rebuilt the XML sitemap, submitted to Google Search Console API, and established versioning for future audits.

Technical Details: Wayback Machine Integration and Content Recovery

Rather than manually recreating legacy pages, we automated content recovery from Internet Archive:

# Pull legacy URL list from Wayback CDX API
curl -s "https://cdx.archive.org/api/v1/search?url=sailjada.com/*&output=json&collapse=urlkey&matchType=prefix&filter=statuscode:200" \
  | jq '.[] | select(.[1] | contains("sailjada.com"))' > legacy_urls.json

# Fetch specific Wayback snapshot
curl -s "https://archive.org/wayback/available?url=sailjada.com/charter-options/×tamp=202605" \
  | jq -r '.archived_snapshots.closest.url'

This approach preserved original title tags, H1 text, and meta descriptions—critical for maintaining search ranking signals. The audit script (/tmp/sailjada-legacy-audit.py) categorized 250 URLs into seven buckets:

  • Legacy charter permalinks (~20): Old product pages; stubbed with 301 redirects to canonical modern equivalents.
  • SEO landing pages (~10): Whale watching, dinner cruises, seasonal content; rebuilt with modern design + Wayback content.
  • WordPress blog posts (~42): 2010–2024 archive; consolidated to /blog/ hub with category canonicals.
  • Event/calendar URLs (~32): Dead calendar plugin URLs; 301 to /sd-sailing-calendar/ or events hub.
  • Dated archives (~30): /2010/06/post-slug/ taxonomy; low-value, stubbed to homepage or skipped.
  • Whale watching (~17): Consolidated to single /whale-watching/ hub page with legacy canonical preservation.
  • Other (~25): Broken asset links, malformed URLs; handled case-by-case or 404'd intentionally.

Parity Page Generation: Automating at Scale

Generating 48 pages manually was untenable. We built three generator scripts that use Wayback content + modern templates:

# Generate tier-1 stub pages (7 SEO categories)
python3 /tmp/build_parity_stubs.py \
  --input-audit /tmp/sailjada-legacy-audit.json \
  --template-dir ./templates/parity \
  --output-dir /tmp/parity-out/ \
  --category tier1_seo_landing

# Generate whale-watching hub + 22 legacy stubs
python3 /tmp/build_parity_whale.py \
  --wayback-snapshots /tmp/whale-snapshots.json \
  --canonical-map ./whale-canonical-map.csv \
  --output-dir /tmp/parity-out/whale-watching/

Each generator:

  • Fetches Wayback snapshot metadata (title, meta description, content excerpt).
  • Injects into Handlebars templates with modern design (dark navy + gold, responsive grid, modal booking widget).
  • Sets canonical links to the modern equivalent page (e.g., <link rel="canonical" href="/charter-types/" />).
  • Adds structured data (Schema.org Organization, LocalBusiness, Event where applicable).
  • Outputs valid HTML5 with preconnect/prefetch hints for analytics and third-party scripts.

Infrastructure: S3 Deployment and CloudFront Cache Invalidation

Files were deployed to S3 bucket s3://sailjada.com/ with the following workflow:

# Sync 48 parity pages to staging first (read-only validation)
aws s3 sync /tmp/parity-out/ s3://sailjada-staging.com/ \
  --exclude "*" \
  --include "*/index.html" \
  --metadata "cache-control=no-cache,public" \
  --profile sailjada-deploy

# Verify staging renders (basic auth required)
curl -u staging:$STAGING_PASS \
  https://staging.sailjada.com/whale-watching/index.html \
  -I | grep "200 OK"

# Promote to production with CloudFront invalidation
aws s3 cp /tmp/parity-out/ s3://sailjada.com/ \
  --recursive \
  --exclude "*" \
  --include "*/index.html" \
  --cache-control "public, max-age=3600, must-revalidate" \
  --profile sailjada-deploy

# Invalidate CloudFront cache (distribution ID: E2K7ABC123XYZ)
aws cloudfront create-invalidation \
  --distribution-id E2K7ABC123XYZ \
  --paths "/*" \
  --profile sailjada-deploy

This two-stage approach (staging → prod) prevented cache poisoning and allowed manual verification before public exposure. The invalidation ensured browsers and edge nodes flushed old 404 responses within seconds.

Sitemap Regeneration and Search Console Submission

After deploying 48 parity pages, we rebuilt the sitemap to include new URLs:

# Generate sitemap from S3 directory structure
python3 ./scripts/generate_sitemap.py \
  --bucket sailjada.com \
  --output ./sitemap.xml \
  --exclude-patterns "**/draft/*, **/*.tmp"

# Deploy to S3
aws s3 cp ./sitemap.xml s3://sailjada.com/sitemap.xml \
  --content-type "application/xml" \
  --cache-control "public, max-age=86400"

# Submit to Google Search Console via API
curl -X POST \
  "https://www.google.com/ping?sitemap=https://sailjada.com/sitemap.xml"

The new sitemap expanded from ~150 URLs to ~200+, with explicit lastmod and priority tags to signal freshness to Google's crawler.

Key Decisions and Rationale

Why Wayback, not manual rewrite? Recovering content from Wayback preserves historical meta tags and H1 text, maintaining SEO signals that would be lost