```html

Automating Charter Document Publishing: Building a Document Pipeline for JADA Operations

This session focused on building a reliable document publishing pipeline for charter operations at JADA. The core challenge: charter manifests and trip sheets needed to be generated, validated, published to multiple locations, and cached appropriately—all while maintaining data consistency across a distributed system.

What Was Done

We created an end-to-end document pipeline that:

  • Generated charter manifest HTML from trip data (Quinn Male and Jonathan charters)
  • Published documents to two S3 locations with proper content-type headers
  • Invalidated CloudFront cache to ensure immediate availability
  • Validated live URLs returned correct HTTP 200 responses
  • Created a durable local archive in /Users/cb/Documents/repos/jada-ops/

Technical Details

Document Generation

Charter manifests were generated as self-contained HTML files with embedded passenger information. The files were created in two temporary locations initially:

  • /tmp/quinn-male-manifest.html
  • /tmp/quinn-male-trip-sheet.html
  • /tmp/jonathan-manifest.html
  • /tmp/jonathan-trip-sheet.html

The manifest format includes passenger names, trip details, and crew assignments—critical information that must remain synchronized across all distribution channels. A key insight: passenger names needed to appear consistently in published documents, so we validated content at the source before publishing.

S3 Publishing Strategy

Documents were published to two distinct S3 prefixes within the shipcaptaincrew bucket:

  • Manifests prefix: s3://shipcaptaincrew/manifests/ — canonical location for all charter documents
  • Crew-page docs prefix: s3://shipcaptaincrew/crew-page-docs/ — integration point for the crew-facing SPA

Both locations received the same content with identical content-type headers (text/html; charset=utf-8). This redundancy serves different purposes: the manifests prefix is the archival location, while crew-page-docs is the active serving location used by the frontend application.

Example publishing command structure (credentials omitted):

aws s3 cp quinn-male-manifest.html \
  s3://shipcaptaincrew/manifests/quinn-male-manifest.html \
  --content-type "text/html; charset=utf-8"

aws s3 cp quinn-male-manifest.html \
  s3://shipcaptaincrew/crew-page-docs/quinn-male-manifest.html \
  --content-type "text/html; charset=utf-8"

CloudFront Cache Invalidation

After publishing, we invalidated the CloudFront distribution to ensure edge caches didn't serve stale content. The distribution ID for shipcaptaincrew was used to invalidate both document paths:

aws cloudfront create-invalidation \
  --distribution-id  \
  --paths "/manifests/quinn-male-manifest.html" \
                "/crew-page-docs/quinn-male-manifest.html"

CloudFront invalidation is critical in this workflow because:

  • Edge caches would otherwise serve outdated manifests for up to 24 hours
  • Crew members access documents immediately before charter departure
  • Stale data creates operational risk (incorrect passenger counts, crew assignments)

Infrastructure Architecture

The publishing pipeline integrates several AWS services:

  • S3 (shipcaptaincrew bucket): Primary storage with two conceptual zones (manifests/ and crew-page-docs/)
  • CloudFront: Edge caching layer serving documents globally with sub-second latency
  • Lambda (shipcaptaincrew-tools): Existing function in /Users/cb/Documents/repos/sites/queenofsandiego.com/tools/shipcaptaincrew/lambda_function.py that handles document serving and crew page rendering

The Lambda function's build_event_pages and handle_get_doc routes consume documents from the crew-page-docs prefix. This separation of concerns—publishing to a dedicated serving location rather than manifests—prevents accidental overwrites of the archival record.

Key Decisions and Rationale

Dual S3 Locations

Publishing to both manifests/ and crew-page-docs/ seemed redundant initially, but serves distinct purposes:

  • Manifests: Immutable audit trail for compliance and troubleshooting
  • Crew-page-docs: Active serving location with potential for future mutations (corrections, amendments)

This prevents a stale document from being "re-published" and corrupting the audit trail.

Content-Type Headers

Explicit content-type metadata was critical. S3 defaults to binary/octet-stream for unknown extensions, which would cause browsers to download rather than display manifests. Setting text/html; charset=utf-8 ensures documents render inline.

Local Archive Durability

Beyond S3 publishing, documents were also saved locally to:

  • /Users/cb/Documents/repos/jada-ops/quinn-male/quinn-male-manifest.html
  • /Users/cb/Documents/repos/jada-ops/jonathan/jonathan-manifest.html

This provides a second copy independent of AWS infrastructure and simplifies document generation auditing. The local repo structure mirrors the operational structure: one directory per charter with all related documents co-located.

CloudFront Invalidation Timing

Invalidation was performed after confirming content was live in S3 and verified to contain correct data. This prevents invalidating caches before documents are ready, which would cause cache misses at scale.

Operational Tools Created

Two Python utilities were created to automate this workflow:

  • send_charter_emails.py — prepares and sends charter confirmation emails with embedded document links
  • charter_provisioner.py — orchestrates the full pipeline: manifest generation, S3 publishing, and cache invalidation

These tools encode the workflow as code, ensuring consistency and reducing manual steps that could introduce errors.

Validation and Testing

Live URL verification was performed by:

  • Fetching manifests from CloudFront URLs and confirming HTTP 200 responses
  • Spot-checking manifest content for correct passenger names and trip details
  • Verifying both S3 prefixes served identical, current content

This validation loop ensures the pipeline output is correct before operational use.

What's Next