Infrastructure Audit and Toolchain Decisions: Rejecting Complexity, Keeping Observability
What Was Done
Over the past week, we audited the entire infrastructure stack across 15 production static sites, the DragonBodyGuards API surface, and Queen's Fleet content infrastructure. Instead of adopting a retired DevOps engineer's recommended Kubernetes/Jenkins/Docker pipeline, we conducted a systematic review of all 39 production incidents in FIRES.md and mapped them to root causes. The verdict: reject nearly all complex tooling, optimize for the actual failure modes we observe.
Technical Audit Methodology
The analysis started with a complete incident catalog. Every entry in /Users/cb/icloud-jada-ops/FIRES.md was categorized by failure type:
- Silent S3 no-ops — deployments that appeared successful but didn't actually sync (root: expired AWS credentials or region mismatches)
- Token expiration — Lambda functions failing mid-request due to stale credentials
- DNS propagation delays — CloudFront distribution URLs not resolving during rollout windows
- Configuration drift — CloudFront behaviors, cache policies, or Route53 records changed outside of version control
- Health check timeouts — monitoring scripts failing on degraded network conditions
None of the 39 incidents were container-related. None were scaling failures. None required orchestration. The pattern was clear: our actual risk surface is configuration accuracy and credential lifecycle, not deployment throughput.
Toolchain Evaluation
We evaluated each proposed tool against the observed failure modes:
- Kubernetes + microservices framework: Rejected. We run 15 static sites served from S3 + CloudFront. Zero containers in production. Zero scaling incidents. Microservices patterns don't map to our deployment model.
- Jenkins: Rejected. Our deploy cycle is already measured in seconds, with automatic ETag==md5 verification baked in. Jenkins adds a server to babysit for zero observability gain.
- Docker: Rejected. Revisit only if a Lambda function requires a container image (none currently do). Containerizing static sites adds build complexity without solving any observed problem.
- Robot Framework + Selenium: Rejected. Our existing nightly test suite uses curl and Python to exercise actual failure modes—expired credentials, region mismatches, silent API failures. Selenium tests browser rendering; it doesn't catch the incidents we've seen.
- Terraform: Adapted, not adopted. The principle is sound—capturing infrastructure as code—but Terraform's state management and drift detection don't fit a single-operator estate. Instead: a read-only config-dump script.
The Config-Dump Architecture
In place of Terraform, we're building a nightly audit script (ticket queued in ~/dablio/TICKETS.md) that will:
- Read CloudFront distributions from all regions and verify cache behaviors, TLS certificates, origin configurations match expected state
- Audit Route53 hosted zones — DNS records, health checks, traffic policies
- Inventory IAM policies applied to Lambda execution roles and S3 bucket policies
- Compare against version-controlled baselines in
/Users/cb/icloud-jada-ops/infrastructure/ - Alert on drift — if the production CloudFront distribution ID
E2ABCDEF12345has a cache behavior that differs from the baseline, the script flags it with path and old/new values
This solves the actual gap: if the Mac with launchd jobs dies, we can rebuild the entire AWS footprint from git. No state files to manage, no drift detection tool to debug—just diffs against version control and human review.
Credential and Deployment Safety
Our existing deploy pipeline already has the safety properties we need:
- S3 sync with ETag verification:
aws s3 syncwith--metadata-directive COPYand md5 checksum validation ensures idempotent uploads - Lambda credential rotation: Nightly audit of
jada-crew-dispatchand related DynamoDB tables for token expiration using boto3 key-only scans - CloudFront invalidation safety: Deployment scripts lock on specific distribution IDs (e.g.,
E2XXXXXfor thequeensfleet.com CDN) before invalidating paths, preventing cross-domain accidents
What We're Not Adding
No CI/CD server. No container runtime. No Kubernetes. No additional abstraction layers. The incident data doesn't justify them.
What's Next
The config-dump script goes into /Users/cb/icloud-jada-ops/scripts/nightly-aws-audit.py and runs via launchd. It will:
- Fetch live state from CloudFront, Route53, and IAM
- Compare against
/Users/cb/icloud-jada-ops/infrastructure/baselines/ - Log diffs to stdout for manual review (nothing auto-fixes)
- Serve as the source of truth for disaster recovery
This is infrastructure as observability, not infrastructure as code. It gives us recoverability if the Mac dies, without introducing the operational burden of state files, drift detection tools, or continuous reconciliation loops. The data earned this choice.