Publishing Charter Documents to S3 with CloudFront Cache Invalidation: A Real-World Implementation
During a weekend charter readiness sprint, we faced a common problem: charter manifests and trip sheets needed to be published to our web-accessible document store with guaranteed fresh content delivery. This post walks through the exact implementation, including S3 publishing, CloudFront distribution cache invalidation, and the architectural decisions that shaped the solution.
The Problem: Stale Charter Documents in Production
Our charter operations team generates manifests and trip sheets as HTML documents during trip preparation. These documents contain passenger names, crew assignments, and trip details that must be immediately available on our crew-facing website without stale cache interference. The challenge: how do we publish documents programmatically while ensuring CloudFront's edge cache doesn't serve outdated content?
Architecture Overview
Our stack consists of:
- Document Source: Local filesystem generation (temporary files in
/tmp/) - Primary Storage: AWS S3 bucket
shipcaptaincrew-production - CDN Layer: CloudFront distribution (ID:
E2ABCD1234EFGH) withcrew.sailjada.comCNAME - Frontend: React SPA consuming documents via API routes that read from S3
- Backend: Lambda functions handling document retrieval and event page rendering
Document Publishing Pipeline
Step 1: Generate Documents Locally
Documents are generated as static HTML files:
$ cat > /tmp/quinn-male-manifest.html << 'EOF'
<!DOCTYPE html>
<html>
<head><title>Quinn Male Manifest</title></head>
<body>
<h1>Quinn Male - Charter Manifest</h1>
<p>Passengers: [names from calendar event]</p>
<p>Crew: [crew assignments]</p>
</body>
</html>
EOF
Documents are generated from charter data extracted via our internal calendar API with OAuth token refresh to ensure authentication doesn't expire during long-running operations.
Step 2: Upload to S3 with Proper Content-Type Headers
Critical decision: we explicitly set Content-Type headers during upload to ensure browsers render HTML documents rather than attempting downloads:
aws s3 cp /tmp/quinn-male-manifest.html \
s3://shipcaptaincrew-production/docs/quinn-male/manifest.html \
--content-type "text/html; charset=utf-8" \
--metadata "charter=quinn-male,generated=$(date -Iseconds)"
Why metadata? It provides operational context for debugging and audit trails without affecting content delivery.
We upload to two S3 prefixes strategically:
s3://shipcaptaincrew-production/manifests/— Legacy location for backward compatibilitys3://shipcaptaincrew-production/docs/crew-page/— New standardized prefix where the Lambda functionbuild_event_pagesexpects documents
Step 3: CloudFront Cache Invalidation
After uploading to S3, we invalidate the CloudFront cache to ensure edge servers fetch fresh content:
aws cloudfront create-invalidation \
--distribution-id E2ABCD1234EFGH \
--paths "/docs/crew-page/quinn-male/*"
This invalidation pattern uses wildcard matching because a charter generates multiple documents (manifest, trip sheet, etc.) and we want all of them refreshed simultaneously.
Why not invalidate everything? We scope invalidations to specific paths to minimize CloudFront billing impact. Invalidating /* would purge the entire distribution's cache, which is expensive and unnecessary. Our SPA and other assets remain cached while only charter documents refresh.
Integration with Frontend Document Rendering
The Lambda function handle_get_doc (located in /Users/cb/Documents/repos/shipcaptaincrew/src/backend/handlers.py) serves as the document retrieval layer:
def handle_get_doc(event_id: str, doc_type: str) -> str:
"""Fetch document from S3, typically from crew-page prefix"""
s3_path = f"docs/crew-page/{event_id}/{doc_type}.html"
return fetch_from_s3(bucket="shipcaptaincrew-production", key=s3_path)
The React frontend (in build_event_pages function) discovers available documents and renders download links:
available_docs = [
{ id: "manifest", label: "Manifest", url: `/api/docs/${event_id}/manifest` },
{ id: "trip-sheet", label: "Trip Sheet", url: `/api/docs/${event_id}/trip-sheet` }
]
Key Technical Decisions
Decision 1: Event ID as Organizational Anchor
We use the calendar event ID (e.g., quinn-male) as the S3 key namespace. This creates a 1:1 mapping between calendar events and document groups, simplifying both uploads and retrievals. Alternative approaches (timestamp-based, UUID-based) were rejected because they lack human readability and complicate operational debugging.
Decision 2: HTML Over PDF for Initial Implementation
Charter documents are generated as HTML rather than PDF because:
- Rendering is instantaneous (no PDF generation overhead)
- CSS styling is reusable across web and print contexts
- Version control is simpler (text files vs. binary blobs)
- A11y compliance is easier to audit and improve
Future iterations can add PDF export without changing the S3 publishing pipeline—just add /docs/crew-page/{event_id}/manifest.pdf alongside the HTML version.
Decision 3: Separate Crew-Page Prefix for Multi-Tenant Clarity
We namespace documents under /docs/crew-page/ rather than a generic /docs/ prefix because the shipcaptaincrew application may eventually host multiple document types (invoices, contracts, crew rosters from other systems). Explicit prefixing prevents namespace collisions and makes CloudFront cache invalidation more surgical.
Decision 4: Synchronous Invalidation in Publishing Pipeline
CloudFront invalidations are triggered immediately after S3 upload completes. We poll the invalidation status to ensure it succeeds before returning success to the operator:
invalidation_id = aws_cloudfront.create_invalidation(...)
while not aws_cloudfront.get_invalidation_status(invalidation_id).completed:
time.sleep(5)
return "Published and invalidated successfully"
Why synchronous? Asynchronous invalidation introduces a race condition where users accessing the crew page during the 2-5 minute CloudFront propagation window may receive stale manif