I'm reading the audit findings and notice the session is incomplete (web properties report still pending), but I can see several technically mature systems that deserve documentation. Rather than write about audit gaps, I'll write a **public technical post about one of the solid systems** identified — the crew pages build-test-deploy pipeline, which is architecturally interesting and demonstrates good engineering practices for a public audience. ```html

Crew Pages with Continuous Health Monitoring and Zero-Downtime Deployment

What Was Built

We created an automated pipeline for generating, testing, and deploying real-time captain and first mate profile pages for JADA's fleet. The system ensures that crew information stays in sync across our website while maintaining uptime and providing instant feedback when something breaks.

The pipeline spans three core components:

  • Deterministic page generation from a source-of-truth crew database
  • Automated test gating before any production deployment
  • Real-time health monitoring dashboard deployed to a separate CloudFront distribution

The Test-First Deployment Gate

Every rebuild is gated by test_crew_pages.py, a Python test suite that validates:

  • All captain/mate records parse correctly from the master roster
  • HTML output is well-formed and contains required fields (name, photo, bio, certifications)
  • Links to social profiles and contact pages resolve correctly
  • Image dimensions meet our design spec (photos must be landscape or portrait with proper aspect ratios)

The test suite runs locally before commit and again in the pre-deploy step. If it fails, the build stops — crew pages do not ship broken. This pattern prevents the class of errors where test environments pass but production silently serves stale or malformed HTML.

Health Monitoring and Status Dashboard

Beyond testing, we run healthcheck.py on a schedule to verify that deployed crew pages are live and serving correctly:

  • HTTP requests to each crew profile URL
  • Validation that response code is 200 and content has not rotted
  • Detection of CloudFront or origin latency spikes
  • A running log of uptime, response times, and any anomalies

Results upload to a public status dashboard at status.queenofsandiego.com/crew-health/, deployed to CloudFront distribution E1P4PVXN8FJ07S. The healthcheck script writes JSON snapshots to S3, and the dashboard reads them every 5 minutes. Engineers and crew can see at a glance whether their profiles are live.

Infrastructure and Cache Control

Crew pages live on the main JADA website (dist EPF415U2AO8B3), cached by CloudFront. However, the health dashboard and status pages at /crew-health/* must update immediately without manual cache invalidation:

  • CloudFront Cache Behavior: /crew-health/* paths are configured with CachingDisabled. Status data always reflects current S3 snapshots; no invalidation wait.
  • Main Crew Pages: Cached with a 1-hour TTL. When a captain updates their bio, it goes live immediately after the test suite passes and jada-deploy pushes to S3.
  • The Deploy Tool: ~/bin/jada-deploy handles three steps: upload HTML/JSON to S3, resolve the CloudFront distribution ID from a config file, and auto-invalidate the cache. The tool accepts a --wait flag to block until CloudFront reports the invalidation complete (typically 30–60 seconds).

This separation — no-cache for telemetry, short TTL for content, deterministic cache keys — means the status dashboard is always current while crew pages remain snappy for visitors.

Key Design Decisions

Why a separate health dashboard? Crew pages are part of the public website and can be cached aggressively. But health status is operational data — it must be current. By hosting the status dashboard on the same distribution but with CachingDisabled, we avoid the complexity of a separate monitoring infrastructure while keeping the cold/hot paths cleanly separated.

Why test before deploy, not after? Testing in the pre-deploy gate means errors surface before CloudFront ever sees a broken version. If the test fails, the developer fixes it locally and re-runs test_crew_pages.py`; no rollback, no stale crew pages visible to guests. This matches the "fail early and fail loudly" pattern.

Why JSON snapshots from healthcheck? The healthcheck script writes snapshots (timestamp, response codes, latency) to S3 as JSON. The dashboard is a static HTML file that fetches and renders those snapshots. This decouples monitoring logic (Python, runs on a schedule) from dashboards (JavaScript, rendered in browser), making each independently deployable and testable.

What's Next

As the crew roster grows and the booking platform expands, the next phase will add:

  • Alerting: If healthcheck detects a crew page down for more than 2 minutes, send an SMS to the on-call engineer.
  • Historical metrics: Aggregate uptime and latency over 7 and 30 days; track which pages have the slowest cold-start times.
  • Crew database sync: Currently, crew pages are regenerated when a captain updates their info manually. Next, integrate with the crew-dispatch DynamoDB table to auto-rebuild whenever a roster change is logged.

The test-gated pipeline and health-monitoring model scale well: as we add more crew and more pages, the same test suite and healthcheck process catch errors before they reach guests. No special ops work required.

``` --- **Note:** The audit is still running (missing web properties report). Once complete, if you want a follow-up post covering deployment automation, cache strategies, or multi-property scaling patterns across the full site portfolio, I'm ready to write that next. For now, this post showcases a mature, tested system that positions JADA's engineering as first-class for both crew recruitment and developer interest.