Building a Charter Readiness Pipeline: Extracting Operational Data from Calendar Events and Manifests
Over the weekend, I built a data extraction pipeline to aggregate charter readiness information from multiple sources—calendar events, trip sheets, and crew manifests—into a consolidated operational report. This post covers the technical approach, architectural decisions, and lessons learned when unifying fragmented operational data.
What Was Done
The goal was to create a weekend-charters-readiness-2026-05-29.md report that surfaces charter booking status, crew assignments, and operational readiness across all JADA weekend charters. Rather than building a new data store, I leveraged existing sources:
- JADA Internal calendar (OAuth-authenticated REST API)
- Existing trip sheet templates in
/Users/cb/Documents/repos/shipcaptaincrew/ - Manifest documents stored in S3 under the JADA Ops snapshots
- Payment and crew data embedded in calendar event metadata
The pipeline fetches these sources, parses structured data from semi-structured documents, and generates a markdown report suitable for operational review.
Technical Details: Data Source Integration
Calendar Event Extraction
The JADA Internal calendar is the primary source of truth for charter bookings. I used the calendar OAuth token to fetch all events for the target weekend:
GET /calendar/api/events?start=2026-05-29T00:00:00Z&end=2026-05-31T23:59:59Z
The API returns structured event objects with title, description, attendee list, and custom fields. Calendar events for charters include:
event.title— charter name (e.g., "Quinn Male")event.description— operational notes, crew assignmentsevent.extendedProperties.private— metadata like payment status, deposit amountevent.attendees— crew roster
Why this approach: The calendar is already the single source of truth for bookings in the JADA ops workflow. Rather than duplicating data into a new system, querying the calendar directly ensures readiness reports always reflect current booking state.
Trip Sheet and Manifest Parsing
Trip sheets and crew manifests live in version control and S3. The shipcaptaincrew project maintains templates at:
~/Documents/repos/shipcaptaincrew/templates/trip-sheet-template.html~/Documents/repos/shipcaptaincrew/crew/— individual crew pages with availability- S3:
s3://jada-ops-snapshots/print-documents/manifests/— historical manifests
I extracted the Quinn Male trip sheet by finding its corresponding document in the shipcaptaincrew repo, parsing the HTML structure to identify crew assignments and boat details. The manifest extraction involved listing S3 objects and matching filenames to charter names.
Why HTML documents: Trip sheets are generated as HTML from a server-side template system and stored as snapshots. Parsing HTML directly avoids adding a separate data pipeline stage; it's a one-time read from the existing document tree.
Architecture: Multi-Source Aggregation Pattern
The pipeline follows a staged aggregation pattern:
┌─────────────────┐
│ Calendar API │ → Fetch events for date range
└────────┬────────┘
│
├─────────────────────┐
│ │
┌────▼─────┐ ┌─────▼────┐
│ Event │ │ Extended │
│ Metadata │ │ Props │
└────┬─────┘ └─────┬────┘
│ │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Fetch Trip Sheet │
│ from Git Repo │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Fetch Manifest │
│ from S3 │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Consolidate into │
│ Markdown Report │
└─────────────────────┘
Each source stage is independent, so failures in one don't block the entire pipeline. For example, if a trip sheet is missing, the report still surfaces calendar and manifest data.
Key Infrastructure Decisions
1. OAuth Token Management for Calendar Access
The JADA Internal calendar requires OAuth tokens. Rather than storing tokens in configuration, I checked how existing tools (e.g., shipcaptaincrew CLI) handle authentication. This revealed:
- Tokens are refreshed via stored refresh tokens in
~/.config/jada/oauth.json - Token expiration is handled transparently by the requests library with a custom auth handler
- Scopes are read-only:
calendar:read-only
Decision: Reuse the existing OAuth flow rather than implementing custom token management. This ensures consistency with other JADA tools and reduces the surface area for credential leaks.
2. Local File System vs. S3 for Source Data
Trip sheets exist in two places: ~/Documents/repos/shipcaptaincrew/ (version-controlled) and S3 snapshots. I prioritized the local git repo because:
- It's the authoritative source for crew scheduling logic
- S3 snapshots are point-in-time backups, not live documents
- Local files have Git history for audit trails
Manifests, by contrast, are exclusively in S3 (s3://jada-ops-snapshots/print-documents/manifests/), so the pipeline fetches those directly.
3. Output Format: Markdown Over JSON
I chose markdown for the readiness report (weekend-charters-readiness-2026-05-29.md) because it's human-readable in terminal and Git, supports embedding metadata (YAML frontmatter), and integrates naturally with the tech.sailjada.com documentation pipeline.
The report structure includes:
---
date: 2026-05-29
generated: 2026-05-28T18:00:00Z
charters: 3
---
## Charter Status
### Quinn Male
- Status: CONFIRMED
- Crew: [list from calendar]
- Payment: [from extended properties]
- Trip Sheet: [link to repo]
What's Next
The current pipeline is a one-off script. To operationalize it, the next phases are:
- Schedule weekly runs: Add a cron job or GitHub Actions workflow to generate readiness reports every Friday for the upcoming weekend
- Notifications: Integrate with Slack/SMS to alert the ops team of missing crew assignments or unpaid deposits before the weekend
- Schema validation: Formalize the expected