Production Safety Patterns: Lessons from Charter Ops Code Review
A nightly review of sailing charter production scripts surfaced several critical patterns that deserve deeper examination. These aren't isolated bugs—they're systematic failure modes that recur across payment systems, communications, and safety gates. This post covers real findings from a solo-operator ops infrastructure and the architectural decisions that prevent them from becoming incidents.
The Crew Manifest Auth Vulnerability: Tokens in Query Parameters
The new crew manifest route was implemented to serve guest-PII pages (crew names, contact info, dietary restrictions) with authentication via a query parameter token: ?t=session-token-here.
This pattern creates three failure modes:
- CloudFront/Lambda access logs: The full token persists in AWS access logs and CloudFront request logs, visible to anyone with read access to the logging S3 bucket. Token theft from logs is common in real incidents.
- Browser history: Any browser with the link in its history exposes the token. Shared devices become a vector.
- Link forwarding: Crews screenshot or forward the manifest link—the token travels with it.
Why this happened: The manifest page was built as a quick read-only view with minimal auth surface. Query parameters felt simpler than introducing cookies or signed URLs, especially for a guest-facing, time-limited feature.
The fix: Replace the direct token with a POST-only or cookie-based session, or use short-lived AWS CloudFront signed URLs (valid for 1–4 hours). If query parameters must be used, scrub sensitive patterns from access logs using CloudFront Lambda@Edge or ALB listener rules to strip ?t=.* before logging.
Command reference: Generate a signed URL for a CloudFront distribution with:
aws cloudfront get-distribution-config \
--id E1A2B3C4D5E6F7 | jq .DistributionConfig.Origins[0].DomainName
Then use the AWS CloudFront Signer (boto3 or NodeJS SDK) to generate 1-hour-validity tokens bound to the manifest path.
Silent Gate Failures: Stale Snapshot Syndrome
Three critical gates protect charter execution: MT-01 (manifest readiness), MT-02 (compliance flags), and MT-03 (account balance). All three implement their logic against a cached snapshot file, last-events.json, rather than querying live DynamoDB.
The risk is stark: if last-events.json becomes stale (last updated 6+ hours ago, or missed a scheduled refresh), the gates will silently pass on charters they never saw. A balance sheet update, a compliance flag flip, or a crew roster change won't reach the gates—and charters proceed into payment or liability exposure.
This is a repeat of incidents I-11 and I-37, both caused by stale snapshots masking critical state changes.
Why the snapshot exists: DynamoDB queries during peak booking periods add latency and cost. The snapshot was introduced to cache the query results and serve gates sub-millisecond. It's a valid optimization.
The problem: No staleness assertion. The gates trust whatever last-events.json contains, without verifying when it was generated.
The fix: Add a test assertion in test_charter_readiness.py that validates last-events.json mtime is within N hours (e.g., 2 hours) before gate evaluation:
def test_gates_reject_stale_snapshot(self):
snapshot_age_hours = (time.time() - os.path.getmtime('last-events.json')) / 3600
self.assertLess(snapshot_age_hours, 2,
f"last-events.json is {snapshot_age_hours:.1f}h old; gate safety unavailable")
self.assertTrue(charter_gates_pass(self.test_charter_id))
If the snapshot refresh job fails, the test fails red before any charter logic runs.
IAM Key Rotation Hygiene
A decision document, jada-ops/decisions/2026-07-03-iam-key-rotation.md, outlined a rotation of AWS access keys. During the rotation, an access key ID appeared in the diff before being redacted. Key IDs themselves aren't credentials (the secret access key is), but the ops docs have a rule: "no secret values in docs."
The finding: Redacting the key ID after commit doesn't remove it from git history. It's still recoverable with git log -p. While a key ID alone doesn't grant access, it violates the stated policy and increases the surface for downstream leaks if paired with a secret elsewhere.
The resolution: Don't attempt history rewriting. Instead, rely on the planned key deletion immediately after the soak period. Confirm that both the old key and the rotated-in key are deleted (not merely marked Inactive) within 24 hours of rotation completion:
aws iam list-access-keys --user-name deploy-user \
--query 'AccessKeyMetadata[?Status==`Inactive`]'
This verification should be part of the rotation playbook and tracked in an incident log.
Incident Tracking: Duplicate Numbering
The incident ledger, jada-ops/docs/FIRES.md, contains 35 documented production failures. During the review, two separate incidents (an S3 copy silent no-op deploy and a phone OAuth redirect dead-end) were both assigned incident number I-37. This numbering collision will corrupt future test-coverage references and incident correlation.
The fix: Renumber the phone-OAuth incident to I-38 and update the summary count in FIRES.md to 36 total incidents. Ensure future incidents are assigned sequentially and validated in git pre-commit hooks.
Key Takeaways for Production Ops
- Tokens belong in headers or cookies, not URLs. Query parameters persist in logs and history.
- Cached state needs a freshness guarantee. An age check before use prevents silent failures that repeat past incidents.
- Redaction doesn't erase history. Plan key deletion and verification instead.
- Incident tracking must be deterministic. Enforce unique numbering with tooling, not manual discipline.
Next steps: A full diff review including launchd plist and shell script changes under ~/bin/ is needed to complete the review surface. The truncated diff points to executable changes that weren't yet analyzed for interpreter paths and PATH dependencies.