```html

Fixing Dead Links at Scale: Building a Legacy URL Parity System for SailJada.com

When an Instagram link-in-bio to sailjada.com/charter-options/ started returning 404s, it revealed a larger infrastructure debt: 236 out of 250 legacy URLs from our pre-redesign site were broken. This post walks through how we diagnosed the scope, built a parity system to resurrect that content, and deployed it to production without losing SEO equity or breaking the booking flow.

The Problem: A Single 404 Becomes a Systemic Issue

We discovered that marketing had been directing Instagram traffic to /charter-options/, but the canonical page lived at /charter-types/ (deployed 2026-05-25). The IG link was cold, and a quick manual check revealed we weren't just missing one page—we were missing dozens, if not hundreds.

The immediate question: How many legacy URLs from the old site were actually returning 404s? And for each one, did we have Wayback Machine snapshots to prove they were once real, indexable content?

Diagnosis: Wayback CDX API + Full URL Audit

Rather than guessing, we automated a complete legacy URL audit:

  • Pulled the full Wayback CDX snapshot list for sailjada.com using the Wayback Machine's CDX Search API, filtering for HTTP 200 responses between 2010 and 2024.
  • Wrote a Python audit script (/tmp/sailjada-legacy-audit.py) that:
    • Extracted every unique URL path from the CDX dataset
    • Made HEAD requests to the current production S3 bucket (via CloudFront) to check status codes
    • Categorized each dead URL by content type (blog post, charter permalink, event, landing page, etc.)
    • Dumped results to /tmp/sailjada-legacy-audit.json with Wayback snapshot URLs for manual review
  • Results: 250 substantive URLs indexed; 236 were 404s; only 14 worked. This gave us a concrete backlog to work from.
$ python3 /tmp/sailjada-legacy-audit.py
Scanning 250 legacy URLs...
236 are 404 (no prod page)
14 are 200 (already live)
Audit complete: results in /tmp/sailjada-legacy-audit.json

Strategy: Tiered Parity Pages, Not Redirects

We chose parity pages (new pages at legacy URLs with updated design but preserved titles, metadata, and canonicals) rather than 301 redirects for three reasons:

  • SEO preservation: Wayback snapshots showed these URLs had backlinks and organic search history. A 301 would work, but we wanted to avoid the redirect chain cost and preserve the URL equity directly.
  • Traffic attribution: Marketing can see whether the IG link is driving clicks to /charter-options/ vs. being bounced elsewhere. With a redirect, that signal gets muddied.
  • Phased rollout: Building parity pages lets us deploy incrementally (Tier 1 high-traffic pages first, low-value taxonomies later) without a big-bang redirect ruleset change.

We organized the 236 broken URLs into buckets:

  • Tier 1 (High Priority): Charter permalinks, SEO landing pages, whale-watching hub — ~47 pages
  • Tier 2 (Blog / Events): WordPress blog posts (42 URLs), event calendar pages (32 URLs) — rebuild or consolidate to new content hubs
  • Tier 3 (Taxonomy / Archive): Category pages, tag archives, author pages, dated archives (~30 URLs) — low-value stub or skip

Implementation: Automated Parity Page Generation

Manually building 48 pages would be error-prone and slow. Instead, we wrote three Python generators:

/tmp/build_parity_stubs.py — Lightweight Redirector Pages

For charter permalinks and low-priority content, this generator creates minimal HTML files that preserve the legacy title/H1 and canonical link, then redirect users to the appropriate new page:

<!DOCTYPE html>
<html>
<head>
  <title>Corporate Charters | SailJada San Diego</title>
  <link rel="canonical" href="https://sailjada.com/charters/">
  <meta http-equiv="refresh" content="0; url=/charters/">
</head>
<body>
  <h1>Corporate Charters</h1>
  <p>Redirecting to our charter options...</p>
</body>
</html>

This keeps the page lightweight (improves Core Web Vitals) while still signaling to Google that the URL is alive and canonical.

/tmp/build_parity_landings.py — Full-Content Landing Pages

For SEO-critical landing pages (catamaran charters, dinner cruises, fall sailing, etc.), this generator:

  • Fetches the Wayback snapshot content via the CDX API
  • Extracts the plain-text body and key H1/title from the archive
  • Re-templates it into our current design system (dark navy + gold, responsive layout, modal booking widget)
  • Sets the canonical to the legacy URL itself (since this is the canonical now)
  • Writes to s3://sailjada.com/{legacy-path}/index.html
$ python3 /tmp/build_parity_landings.py \
    --wayback-url "https://web.archive.org/web/20230615000000/sailjada.com/catamaran-charter-captain-san-diego/" \
    --legacy-path "/catamaran-charter-captain-san-diego/" \
    --title "Catamaran Charter Captain San Diego | SailJada" \
    --output-dir "/tmp/parity-out"

/tmp/build_parity_whale.py — Specialized Hub + Stubs

Whale-watching had 17 distinct legacy event URLs (/whale-watching-san-diego-2021/, /whale-watching-december/, etc.), all now consolidated into a single /whale-watching/ hub. This script:

  • Builds a master hub page at /whale-watching/index.html with current event info and booking CTA
  • Generates 22 lightweight stub pages for each legacy variant, all with <link rel="canonical" href="/whale-watching/">
  • Ensures old backlinks and social shares still land on a real page (not a 404), then