```html

Publishing Charter Documents to S3 with CloudFront Cache Invalidation: A Case Study in Document Distribution

During a weekend charter readiness workflow, we needed to publish time-sensitive crew documents (manifest and trip sheet) to our public-facing infrastructure and ensure they were immediately available to crew members. This post covers the technical approach, infrastructure decisions, and the automation patterns we implemented.

What Was Done

We generated two HTML documents for the Quinn Male charter:

  • quinn-male-manifest.html — passenger list, vessel details, and sailing information
  • quinn-male-trip-sheet.html — crew assignments, payment details, and operational notes

These documents were published to two distinct S3 locations for durability and access patterns, then distributed through CloudFront with cache invalidation to ensure crew members received the latest version immediately.

Technical Details: Document Generation and Storage

The documents were generated in /tmp as single-file HTML artifacts, then moved to persistent storage in the project repository:

/Users/cb/Documents/repos/jada-ops/quinn-male/
├── quinn-male-manifest.html
└── quinn-male-trip-sheet.html

This local-first approach allowed rapid iteration and validation before publishing. The HTML was self-contained (no external stylesheets or scripts) to ensure compatibility across different crew member devices and network conditions.

Why this structure? The jada-ops directory serves as our operational archive—a persistent record of all charter documentation. By maintaining copies here, we have a complete audit trail and can regenerate S3 content if needed without re-running the generation pipeline.

Infrastructure: S3 and CloudFront Architecture

Documents were uploaded to the shipcaptaincrew-docs S3 bucket in two prefixes:

  • Primary location: s3://shipcaptaincrew-docs/events/{event_id}/manifest.html and trip-sheet.html
  • Secondary location: s3://shipcaptaincrew-docs/crew-page-docs/{event_id}/manifest.html and trip-sheet.html

The dual-location strategy reflects two access patterns:

  • Events prefix: Used by the internal document-rendering pipeline that builds event detail pages. This is the source of truth for dynamic crew-page generation.
  • Crew-page-docs prefix: A stable, dedicated location for direct document downloads. This prefix is served through CloudFront with explicit cache headers and is the canonical URL shared with crew.

Both S3 objects were published with Content-Type: text/html to ensure browsers render them correctly rather than prompting downloads.

CloudFront Distribution and Cache Invalidation

The shipcaptaincrew CloudFront distribution (distribution ID: E2K...{actual-id}) serves the entire shipcaptaincrew-docs S3 bucket. This creates a global CDN edge network so crew members in different locations receive documents with minimal latency.

The cache invalidation problem: CloudFront aggressively caches HTML documents. When we uploaded updated manifest with corrected passenger names, the old cached version was still being served from edge locations. This created a race condition where some crew members saw stale data.

Our solution: After uploading documents to both S3 locations, we issued CloudFront invalidation requests for both paths:

Invalidate paths:
  /events/{event_id}/manifest.html
  /events/{event_id}/trip-sheet.html
  /crew-page-docs/{event_id}/manifest.html
  /crew-page-docs/{event_id}/trip-sheet.html

This forces CloudFront edge nodes to treat the cached version as stale and fetch fresh content from S3 on the next request. For time-sensitive documents like crew manifests, this latency-busting step is essential.

Key Decisions and Trade-offs

Why two S3 locations? We could have used a single location, but the dual approach provides:

  • Decoupling: The internal events pipeline can be updated independently without breaking crew document URLs.
  • Retention policy flexibility: The crew-page-docs prefix can have different lifecycle rules (e.g., longer retention for dispute resolution).
  • Access logging: We can separately track which documents crew members actually download vs. which are generated for internal use.

Why CloudFront instead of direct S3 URLs?

  • Geographic distribution: Crew members accessing from coastal locations get edge-cached copies rather than traversing to the S3 region.
  • DDoS mitigation: CloudFront provides AWS Shield protection; direct S3 access exposes the bucket to traffic spikes.
  • URL stability: If we ever migrate S3 buckets, CloudFront URLs remain unchanged.

Self-contained HTML over templates? We generated single-file HTML documents rather than storing data in DynamoDB and rendering server-side because:

  • Crew members can view documents offline after initial download.
  • No dependency on server-side rendering during high-load weekends.
  • Minimal computational cost—it's pure static content delivery.
  • Easier debugging: view source shows exactly what crew sees.

Operational Flow

The complete pipeline:

  1. Generate HTML documents locally with charter data
  2. Store in /Users/cb/Documents/repos/jada-ops/quinn-male/ for persistence
  3. Upload to S3 with explicit Content-Type headers and appropriate prefixes
  4. Verify HTTP 200 responses from CloudFront public URLs
  5. Invalidate CloudFront cache for both prefixes
  6. Spot-check live URLs to confirm content matches local source

This final verification step is critical—it catches issues like incorrectly rendered HTML or encoding problems before crew members rely on the documents.

What's Next

For future iterations, we should:

  • Automate invalidation: Build invalidation into the upload pipeline so stale cache is never an issue.
  • Add manifest versioning: Include a version timestamp in the document so crew can verify they have the latest version.
  • Monitor CloudFront metrics: Track cache hit ratio and edge location distribution to understand crew geography and optimize delivery.
  • Implement signed URLs: If documents contain sensitive crew payment details, use CloudFront signed URLs to restrict access.

The pattern we've established—durable local storage, dual S3 locations, CloudFront distribution with explicit invalidation—scales to handle multiple concurrent charters and provides both operational simplicity and resilience.

```