Publishing Charter Documents to S3 with CloudFront Cache Invalidation: A Multi-Stage Deployment Pattern
During this development session, we implemented a complete document publishing pipeline for the JADA operations charter system, moving from temporary staging in /tmp to permanent S3 storage with CloudFront CDN caching. This post covers the technical decisions, architecture patterns, and exact infrastructure components involved.
The Problem: Ephemeral Documents and Crew Page Synchronization
The Quinn Male charter required two critical documents—a manifest and trip sheet—to be generated dynamically and made available to crew members through the shipcaptaincrew web application. Initial generation placed these files in /tmp/quinn-male-manifest.html and /tmp/quinn-male-trip-sheet.html, which are temporary and would be lost on system reboot. The crew page needed to reference these documents with consistent URLs that would remain stable across deployments and survive infrastructure restarts.
Solution Architecture: Multi-Destination Publishing
We implemented a three-stage publication strategy:
- Stage 1: Local durability — Copy generated documents to
/Users/cb/Documents/repos/jada-ops/quinn-male/for version control and local backup - Stage 2: S3 primary storage — Upload documents to the
shipcaptaincrewS3 bucket under thedocs/prefix withtext/htmlcontent-type headers - Stage 3: CDN invalidation — Trigger CloudFront cache invalidation to ensure live URLs reflect the latest content immediately
Technical Implementation Details
Document Storage Structure
Documents were published to two S3 prefixes within the same shipcaptaincrew bucket:
s3://shipcaptaincrew/docs/quinn-male-manifest.html— Standard crew documentation locations3://shipcaptaincrew/crew-page/docs/quinn-male-manifest.html— Parallel storage for redundancy and legacy compatibilitys3://shipcaptaincrew/crew-page/docs/quinn-male-trip-sheet.html— Trip sheet in the same structure
The dual-location approach was intentional: it maintains backward compatibility with existing document-loading code while establishing a new primary location structure. The crew page Lambda function in shipcaptaincrew/tools/ queries both locations when rendering event pages, providing a graceful migration path.
S3 Upload with Proper Content-Type
Critical detail: manifest and trip sheet files must be served as HTML, not generic binary objects. The upload process explicitly sets:
Content-Type: text/html; charset=utf-8
This prevents S3 from defaulting to application/octet-stream, which would trigger downloads instead of in-browser rendering. The decision to use text/html over application/xhtml+xml was pragmatic—browsers handle HTML more universally, and the documents are generated as valid HTML5.
CloudFront Distribution and Cache Invalidation
The shipcaptaincrew project uses a CloudFront distribution (discovered via AWS console and existing tool configuration) that caches all S3 objects. When documents are updated, the CloudFront edge caches must be invalidated to serve fresh content immediately.
Invalidation was performed for both document paths:
/docs/quinn-male-manifest.html/docs/quinn-male-trip-sheet.html/crew-page/docs/quinn-male-manifest.html/crew-page/docs/quinn-male-trip-sheet.html
The specific distribution ID is stored in the shipcaptaincrew project configuration. Each invalidation request creates a separate cache-busting operation, typically completing within 30-60 seconds globally. This approach avoids the performance penalty of invalidating wildcard patterns (/docs/*), which would clear unrelated cached content.
Integration with Crew Page Backend
The shipcaptaincrew application maintains a Lambda function (located in the project's tools directory, invoked during build_event_pages operations) that renders charter event pages. This function:
- Receives the event ID from Route53 DNS records mapping charter names to database lookups
- Queries for associated documents in both the legacy
docs/and newcrew-page/docs/prefixes - Generates HTML links in the event detail page pointing to CloudFront URLs
- Relies on CloudFront to serve cached versions, with origin fallback to S3
The decision to maintain both document prefixes in the backend code (handle_get_doc function) meant we could publish to the new structure without requiring immediate code changes. The crew page discovers documents dynamically by listing the S3 prefix, so new documents appear automatically once uploaded.
Key Operational Decisions
Why Dual-Location Publishing?
Publishing to both docs/ and crew-page/docs/ protects against breaking changes if any code path expects documents in a specific location. It's redundant but provides safe migration semantics—if one location becomes the canonical source, the other continues to work as fallback.
Why Manual CloudFront Invalidation Instead of Automatic?
Automatic invalidation on every upload would generate unnecessary cache-busting operations for documents that rarely change. Manual invalidation ties to the publishing workflow, ensuring we only clear cache when documents actually require immediate updates. This is cost-effective (CloudFront invalidations have per-request charges) and keeps the operation explicit and auditable.
Why Store Both in Git and S3?
The jada-ops repository at /Users/cb/Documents/repos/jada-ops/quinn-male/ serves as the durable record of what was published. S3 is the runtime source of truth. Separating these concerns means git history captures the document state, while S3 provides the actual served content. If S3 objects are accidentally deleted, the git history allows reconstruction.
Verification and Live Testing
Post-publication verification confirmed:
- CloudFront URLs return HTTP 200 with correct content-type headers
- Passenger names and trip details render correctly in live manifests
- Both Quinn Male documents are discoverable via the crew page event detail endpoint
- Content matches the local jada-ops repository versions
The spot-check validated that the manifest HTML included correct passenger names (initially missing in the first generation), proving the full pipeline works end-to-end from document generation through cache invalidation to live serving.
What's Next
Future improvements could include:
- Automated document generation triggered by calendar event creation in the JADA internal calendar
- S3 lifecycle policies to archive older charter documents to Glacier
- Versioning strategy to maintain historical manifests for audit trails
- Lambda@Edge integration to add security headers or access control to document delivery
The foundation established here provides a reliable, scalable pattern for publishing dynamic charter documents to a