Automating Charter Document Publishing: Building a Document Pipeline for JADA Operations
This session focused on building a reliable document publishing pipeline for charter operations at JADA. The core challenge: charter manifests and trip sheets needed to be generated, validated, published to multiple locations, and cached appropriately—all while maintaining data consistency across a distributed system.
What Was Done
We created an end-to-end document pipeline that:
- Generated charter manifest HTML from trip data (Quinn Male and Jonathan charters)
- Published documents to two S3 locations with proper content-type headers
- Invalidated CloudFront cache to ensure immediate availability
- Validated live URLs returned correct HTTP 200 responses
- Created a durable local archive in
/Users/cb/Documents/repos/jada-ops/
Technical Details
Document Generation
Charter manifests were generated as self-contained HTML files with embedded passenger information. The files were created in two temporary locations initially:
/tmp/quinn-male-manifest.html/tmp/quinn-male-trip-sheet.html/tmp/jonathan-manifest.html/tmp/jonathan-trip-sheet.html
The manifest format includes passenger names, trip details, and crew assignments—critical information that must remain synchronized across all distribution channels. A key insight: passenger names needed to appear consistently in published documents, so we validated content at the source before publishing.
S3 Publishing Strategy
Documents were published to two distinct S3 prefixes within the shipcaptaincrew bucket:
- Manifests prefix:
s3://shipcaptaincrew/manifests/— canonical location for all charter documents - Crew-page docs prefix:
s3://shipcaptaincrew/crew-page-docs/— integration point for the crew-facing SPA
Both locations received the same content with identical content-type headers (text/html; charset=utf-8). This redundancy serves different purposes: the manifests prefix is the archival location, while crew-page-docs is the active serving location used by the frontend application.
Example publishing command structure (credentials omitted):
aws s3 cp quinn-male-manifest.html \
s3://shipcaptaincrew/manifests/quinn-male-manifest.html \
--content-type "text/html; charset=utf-8"
aws s3 cp quinn-male-manifest.html \
s3://shipcaptaincrew/crew-page-docs/quinn-male-manifest.html \
--content-type "text/html; charset=utf-8"
CloudFront Cache Invalidation
After publishing, we invalidated the CloudFront distribution to ensure edge caches didn't serve stale content. The distribution ID for shipcaptaincrew was used to invalidate both document paths:
aws cloudfront create-invalidation \
--distribution-id \
--paths "/manifests/quinn-male-manifest.html" \
"/crew-page-docs/quinn-male-manifest.html"
CloudFront invalidation is critical in this workflow because:
- Edge caches would otherwise serve outdated manifests for up to 24 hours
- Crew members access documents immediately before charter departure
- Stale data creates operational risk (incorrect passenger counts, crew assignments)
Infrastructure Architecture
The publishing pipeline integrates several AWS services:
- S3 (shipcaptaincrew bucket): Primary storage with two conceptual zones (manifests/ and crew-page-docs/)
- CloudFront: Edge caching layer serving documents globally with sub-second latency
- Lambda (shipcaptaincrew-tools): Existing function in
/Users/cb/Documents/repos/sites/queenofsandiego.com/tools/shipcaptaincrew/lambda_function.pythat handles document serving and crew page rendering
The Lambda function's build_event_pages and handle_get_doc routes consume documents from the crew-page-docs prefix. This separation of concerns—publishing to a dedicated serving location rather than manifests—prevents accidental overwrites of the archival record.
Key Decisions and Rationale
Dual S3 Locations
Publishing to both manifests/ and crew-page-docs/ seemed redundant initially, but serves distinct purposes:
- Manifests: Immutable audit trail for compliance and troubleshooting
- Crew-page-docs: Active serving location with potential for future mutations (corrections, amendments)
This prevents a stale document from being "re-published" and corrupting the audit trail.
Content-Type Headers
Explicit content-type metadata was critical. S3 defaults to binary/octet-stream for unknown extensions, which would cause browsers to download rather than display manifests. Setting text/html; charset=utf-8 ensures documents render inline.
Local Archive Durability
Beyond S3 publishing, documents were also saved locally to:
/Users/cb/Documents/repos/jada-ops/quinn-male/quinn-male-manifest.html/Users/cb/Documents/repos/jada-ops/jonathan/jonathan-manifest.html
This provides a second copy independent of AWS infrastructure and simplifies document generation auditing. The local repo structure mirrors the operational structure: one directory per charter with all related documents co-located.
CloudFront Invalidation Timing
Invalidation was performed after confirming content was live in S3 and verified to contain correct data. This prevents invalidating caches before documents are ready, which would cause cache misses at scale.
Operational Tools Created
Two Python utilities were created to automate this workflow:
send_charter_emails.py— prepares and sends charter confirmation emails with embedded document linkscharter_provisioner.py— orchestrates the full pipeline: manifest generation, S3 publishing, and cache invalidation
These tools encode the workflow as code, ensuring consistency and reducing manual steps that could introduce errors.
Validation and Testing
Live URL verification was performed by:
- Fetching manifests from CloudFront URLs and confirming HTTP 200 responses
- Spot-checking manifest content for correct passenger names and trip details
- Verifying both S3 prefixes served identical, current content
This validation loop ensures the pipeline output is correct before operational use.