Rejecting Terraform: Why We Kept Shell Scripts and Deterministic Testing for Multi-Cloud Infrastructure
\n\nTL;DR: After evaluating a complete CI/CD toolchain migration to Terraform, we deliberately chose to keep our shell-script-based deployment model with git mirroring and nightly deterministic tests. This post explains the infrastructure constraints that shaped this decision and how we're strengthening the current approach instead of replacing it.
\n\nThe Problem We Investigated
\n\nThe jada-estate infrastructure landscape spans approximately 15 static sites hosted on S3 + CloudFront, single-file Lambda functions, launchd/cron-based nightly jobs, and several Lightsail instances for peer network operations. Our deployment model relied on:
\n\n- \n
- Verified shell scripts for infrastructure provisioning \n
- Git repository mirroring for state tracking \n
- Deterministic nightly test suites validating the entire estate \n
- Manual Route53 DNS management \n
- CloudFront distribution updates via AWS SDK calls \n
The proposal on the table: migrate everything to Terraform-managed infrastructure, using a centralized state backend and automated CI/CD pipelines. On paper, this sounded like operational maturity. In practice, it introduced risks we couldn't justify.
\n\nWhy We Rejected the Full Terraform Migration
\n\nReason 1: State Management Complexity in a Multi-Account Environment
\n\nOur infrastructure spans multiple AWS accounts and Lightsail fleet management. Terraform's remote state backend introduces a single point of failure. If the state becomes corrupted—a known issue in concurrent environments—recovery is non-trivial. Our shell-script approach stores state in git (versioned, auditable, recoverable), making rollback straightforward.
\n\nReason 2: Nightly Deterministic Testing Catches Real Failures
\n\nOur current system runs comprehensive integration tests every night against live infrastructure. These tests verify:
\n\n- \n
- S3 bucket policies and CloudFront cache invalidation \n
- Lambda function availability and permission boundaries \n
- Lightsail peer sync mechanisms \n
- DNS resolution via Route53 \n
- Certificate validity (checked via
queenof_certs.pystate tracking) \n
Terraform doesn't catch these problems—it only declares what should exist. Our tests verify what actually works. During the recent stall-watchdog deployment, nightly tests detected timing issues that Terraform would have missed entirely.
\n\nReason 3: Lambda Functions Don't Scale Well in Terraform for Our Workload
\n\nOur Lambda functions are single-file deployments, often just a few hundred lines of Python using boto3 for AWS API interactions. Terraform's file-based resource declarations add noise without benefit. Our shell scripts directly zip the function code, validate imports, and deploy via AWS CLI—all in readable, auditable steps.
\n\nWhat We're Doing Instead: Strengthening the Current Model
\n\nGit Mirror as Single Source of Truth
\n\nRather than adopting Terraform state, we're treating git as our infrastructure state store. Every infrastructure change is committed with clear messaging. The repository structure mirrors AWS account topology:
\n\n- \n
/Users/cb/dablio/— main estate coordination and board decisions \n/Users/cb/dablio/peer/lightsail/— Lightsail fleet operations and peer-sync logic \n/Users/cb/icloud-jada-ops/— operations tooling and state files (certificate tracking, failure domain planning) \n
Infrastructure changes are deployed via shell scripts that read from git, validate state, and apply changes atomically.
\n\nDeterministic Nightly Test Suite
\n\nOur FAILURE-DOMAINS-PLAN.md documents known failure modes and their detection strategies. Nightly tests (triggered via launchd on our ops instance) run:
# Validate all static sites are reachable and cacheable\nfor bucket in $ESTATE_S3_BUCKETS; do\n aws s3api head-bucket --bucket \"$bucket\" || echo \"FAILURE: $bucket unreachable\"\ndone\n\n# Check CloudFront distribution cache behavior\nfor dist_id in $CLOUDFRONT_DISTS; do\n aws cloudfront get-distribution --id \"$dist_id\" | jq '.Distribution.Status'\ndone\n\n# Verify Lambda permissions and execution role validity\nfor function_name in $LAMBDA_FUNCTIONS; do\n aws lambda get-function --function-name \"$function_name\" | jq '.Configuration.Role'\ndone\n\n\nEach test is idempotent and reports in JSON for automated alerting.
\n\nboto3-Based Infrastructure Updates
\n\nFor dynamic infrastructure changes (scaling Lightsail instances, rotating certificates, updating Lambda code), we use boto3 Python scripts. The queenof_certs.py state tracker in /Users/cb/icloud-jada-ops/state/ manages certificate lifecycle:
- \n
- Tracks certificate issue/expiry dates \n
- Validates Route53 TLS records before issuing new certificates \n
- Generates audit logs of all state transitions \n
Infrastructure Specifics: What We Actually Run
\n\nS3 + CloudFront Static Sites
\n\nApproximately 15 S3 buckets, each paired with a CloudFront distribution. Cache behavior is controlled via bucket metadata stored in git-tracked JSON configuration files. Example flow:
\n\n# Read bucket config from git\nBUCKET_CONFIG=$(cat infrastructure/s3-buckets/tech-sailjada.json)\n\n# Apply bucket policies via AWS CLI\naws s3api put-bucket-policy \\\n --bucket tech-sailjada-prod \\\n --policy file://infrastructure/policies/tech-sailjada-policy.json\n\n# Invalidate CloudFront cache on deployment\naws cloudfront create-invalidation \\\n --distribution-id E1ABCD1234EFGH \\\n --paths \"/*\"\n\n\nSingle-File Lambda Deployments
\n\nEach Lambda function lives in its own git directory with vendored dependencies. Deployment script:
\n\n#!/bin/bash\nFUNCTION_DIR=$1\ncd \"$FUNCTION_DIR\"\n\n# Validate Python imports\npython3 -m py_compile *.py\n\n# Zip with dependencies\nzip -r function.zip . -x \"*.git*\" \"*.pyc\"\n\n# Deploy to AWS\naws lambda update-function-code \\\n --function-name $(basename $FUNCTION_DIR) \\\n --zip-file fileb://function.zip\n\n\nLightsail Peer Sync and Stall Watchdog
\n\nThe recent Lightsail twin deployment (see git commit 8bc52f7) added autonomous peer-sync logic. A nightly cron job on the Lightsail instance compares state with the git mirror and self-corrects if drift is detected. The stall-watchdog monitors this sync mechanism and alerts if peer replication falls behind by more than 5 minutes.
Key Decision: Prioritize Observability Over Infrastructure-as-Code
\n\nTerraform promises declarative infrastructure, but only if your actual infrastructure matches the declarations. We instead invest in:
\n\n- \n
- Verification: Nightly tests that run against live infrastructure \n
- Traceability: Every change is git-committed with commit messages explaining the why \n
- Auditability: Shell scripts are human-readable; boto3 state tracking logs all mutations \n
- Recovery: Git history makes rollback trivial; no state file corruption to recover from \n
This isn't avoiding modern tooling—it's matching tools to constraints. We run approximately 20 nightly deterministic tests that catch real failures Terraform never would.
\n\nWhat's Next: Strengthening the Gaps
\n\nUpcoming work to solidify this approach:
\n\n- \n
- AWS Config Dump Tooling: Automated export of actual AWS state (all S3 buckets, CloudFront distributions, Lambda functions) for comparison against git-tracked intended state. This catches manual changes in the AWS Console. \n
- DBG Licensure Pre-Check: Before deploying new infrastructure or code, validate that all external dependencies (libraries, APIs, cloud services) comply with our licensing requirements. \n
- Queen's Fleet Review: Comprehensive audit of all Lightsail instances, their peer-sync status, and migration readiness for the upcoming July 18 deployment window. \n
Conclusion
\n\nWe rejected Terraform not from stubbornness, but because our estate's constraints (multiple accounts, nightly verification, single-file Lambdas, git-based audit trails) are better served by shell + boto3 + deterministic testing. This isn't a permanent stance—if we scale to hundreds of resources or multi-region active-active setups, we'll revisit. Until then, we're investing in the observability and auditability that actually catches production failures.
\n\nFor questions about infrastructure decisions or the estate architecture, reach out to the board or ops team via the Dablio coordination system.
"}