I need to check the memory system and read some of the key decision docs to understand the full technical context before writing this post. Reading the Board verdict and decision documents now to extract the specific technical details.

Infrastructure Audit and Toolchain Decisions: Rejecting Complexity, Keeping Observability

What Was Done

Over the past week, we audited the entire infrastructure stack across 15 production static sites, the DragonBodyGuards API surface, and Queen's Fleet content infrastructure. Instead of adopting a retired DevOps engineer's recommended Kubernetes/Jenkins/Docker pipeline, we conducted a systematic review of all 39 production incidents in FIRES.md and mapped them to root causes. The verdict: reject nearly all complex tooling, optimize for the actual failure modes we observe.

Technical Audit Methodology

The analysis started with a complete incident catalog. Every entry in /Users/cb/icloud-jada-ops/FIRES.md was categorized by failure type:

  • Silent S3 no-ops — deployments that appeared successful but didn't actually sync (root: expired AWS credentials or region mismatches)
  • Token expiration — Lambda functions failing mid-request due to stale credentials
  • DNS propagation delays — CloudFront distribution URLs not resolving during rollout windows
  • Configuration drift — CloudFront behaviors, cache policies, or Route53 records changed outside of version control
  • Health check timeouts — monitoring scripts failing on degraded network conditions

None of the 39 incidents were container-related. None were scaling failures. None required orchestration. The pattern was clear: our actual risk surface is configuration accuracy and credential lifecycle, not deployment throughput.

Toolchain Evaluation

We evaluated each proposed tool against the observed failure modes:

  • Kubernetes + microservices framework: Rejected. We run 15 static sites served from S3 + CloudFront. Zero containers in production. Zero scaling incidents. Microservices patterns don't map to our deployment model.
  • Jenkins: Rejected. Our deploy cycle is already measured in seconds, with automatic ETag==md5 verification baked in. Jenkins adds a server to babysit for zero observability gain.
  • Docker: Rejected. Revisit only if a Lambda function requires a container image (none currently do). Containerizing static sites adds build complexity without solving any observed problem.
  • Robot Framework + Selenium: Rejected. Our existing nightly test suite uses curl and Python to exercise actual failure modes—expired credentials, region mismatches, silent API failures. Selenium tests browser rendering; it doesn't catch the incidents we've seen.
  • Terraform: Adapted, not adopted. The principle is sound—capturing infrastructure as code—but Terraform's state management and drift detection don't fit a single-operator estate. Instead: a read-only config-dump script.

The Config-Dump Architecture

In place of Terraform, we're building a nightly audit script (ticket queued in ~/dablio/TICKETS.md) that will:

  1. Read CloudFront distributions from all regions and verify cache behaviors, TLS certificates, origin configurations match expected state
  2. Audit Route53 hosted zones — DNS records, health checks, traffic policies
  3. Inventory IAM policies applied to Lambda execution roles and S3 bucket policies
  4. Compare against version-controlled baselines in /Users/cb/icloud-jada-ops/infrastructure/
  5. Alert on drift — if the production CloudFront distribution ID E2ABCDEF12345 has a cache behavior that differs from the baseline, the script flags it with path and old/new values

This solves the actual gap: if the Mac with launchd jobs dies, we can rebuild the entire AWS footprint from git. No state files to manage, no drift detection tool to debug—just diffs against version control and human review.

Credential and Deployment Safety

Our existing deploy pipeline already has the safety properties we need:

  • S3 sync with ETag verification: aws s3 sync with --metadata-directive COPY and md5 checksum validation ensures idempotent uploads
  • Lambda credential rotation: Nightly audit of jada-crew-dispatch and related DynamoDB tables for token expiration using boto3 key-only scans
  • CloudFront invalidation safety: Deployment scripts lock on specific distribution IDs (e.g., E2XXXXX for thequeensfleet.com CDN) before invalidating paths, preventing cross-domain accidents

What We're Not Adding

No CI/CD server. No container runtime. No Kubernetes. No additional abstraction layers. The incident data doesn't justify them.

What's Next

The config-dump script goes into /Users/cb/icloud-jada-ops/scripts/nightly-aws-audit.py and runs via launchd. It will:

  • Fetch live state from CloudFront, Route53, and IAM
  • Compare against /Users/cb/icloud-jada-ops/infrastructure/baselines/
  • Log diffs to stdout for manual review (nothing auto-fixes)
  • Serve as the source of truth for disaster recovery

This is infrastructure as observability, not infrastructure as code. It gives us recoverability if the Mac dies, without introducing the operational burden of state files, drift detection tools, or continuous reconciliation loops. The data earned this choice.