Rejecting Terraform, Keeping Git-Driven Deploys: Why We Scaled Our CI/CD Without a Framework
\n\nDate: July 5, 2026 | Context: Board decision on estate infrastructure tooling; 15+ static sites, Lambda functions, and deterministic nightly test suite.
\n\nWhat We Did
\n\nOur infrastructure board (convened via dablio's convene_board tool) evaluated a formal CI/CD toolchain proposal—specifically, adopting Terraform for infrastructure-as-code. After reviewing our estate shape, failure modes, and deployment patterns, we rejected the full toolchain and instead doubled down on our existing git-driven, boto3-based deployment model. This decision wasn't risk-averse; it was pragmatic.
In parallel, we clarified our path forward for the DBG mobile application, settling on a Lambda+gateway architecture rather than traditional containerized deployment.
\n\nWhy Terraform Didn't Fit (Yet)
\n\nOur estate currently consists of:
\n\n- \n
- ~15 static S3 + CloudFront sites (verified shell deploys) \n
- Single-file Lambda functions deployed via boto3 scripts \n
- Nightly deterministic test suite (launchd/cron-driven) \n
- Git mirror infrastructure for artifact versioning \n
- Route53 DNS managed via Python boto3 clients \n
Terraform's strength is managing complex, multi-resource dependencies. Our current workload—mostly stateless compute, immutable S3-backed sites, and explicit boto3 orchestration in /Users/cb/icloud-jada-ops/state/queenof_cf_create.py—doesn't require that abstraction layer yet. Adding Terraform would introduce:
- \n
- State management overhead: Remote state files, locking, drift detection—non-trivial for our team size. \n
- Operational coupling: Changes to CloudFront distributions, S3 buckets, or Lambda IAM roles would require Terraform plan/apply cycles. \n
- Learning curve: HCL syntax and Terraform semantics are orthogonal to what our engineers already know (Python, boto3, shell scripting). \n
- Fragmentation risk: While Terraform manages CloudFront, our Lambda deploys stay in Python scripts. That inconsistency is worse than homogeneous boto3. \n
The board's verdict: Scale the existing model horizontally before introducing a new framework.
\n\nWhat We're Keeping: Git-Driven, Deterministic Deploys
\n\nOur deployment pipeline already embeds the principles Terraform promises—just implemented directly:
\n\n- \n
- Infrastructure as code: All CloudFront distribution creation, S3 bucket policies, and Route53 records are defined in Python scripts (e.g.,
queenof_cf_create.py) checked into git. \n - Versioning: Git commit hashes tie infrastructure changes to code changes; our git mirror system (via
jada-estate-git-system) ensures artifact traceability. \n - Repeatability: Nightly deterministic tests (scheduled via launchd and cron) validate that deploys are idempotent; re-running a deploy script produces the same result. \n
- Rollback capability: Checking out a prior git commit and re-running the deployment script reverts infrastructure to that state. \n
This approach is less magical than Terraform, but for our scale, it's more transparent. Engineers see exactly which AWS API calls are made, in what order, and why—no abstraction layer hiding behavior.
\n\nInfrastructure Patterns in Use
\n\nCloudFront Distribution Management: The queenof_cf_create.py script uses boto3's CloudFront client to:
client = boto3.client('cloudfront')\ndistributions = client.list_distributions()\n# Inspect DistributionList, extract distrib IDs\n# Update origins, behaviors, cache policies as needed\n# Call create_distribution_with_tags or update_distribution\n\nEach distribution is idempotent—the script checks current state and only applies deltas. Distribution IDs are tracked in version control alongside their configurations.
\n\nLambda Deployment via Git Hooks: Our nightly test suite includes deterministic Lambda verification. Before promotion, the test harness:
\n\n- \n
- Reads the Lambda function code from git (commit hash). \n
- Packages and uploads to S3 staging bucket. \n
- Invokes boto3
update_function_code()with the S3 key. \n - Runs synthetic tests against the deployed version. \n
- On success, tags the git commit; on failure, logs the divergence and alerts on-call. \n
Route53 as Deployment Target: DNS records are updated via Python scripts that call boto3's Route53 client. Changes to A records, CNAME aliases, or health check configurations are applied in atomic batches, with git tracking which team member authorized each change.
\n\nDBG Mobile App: Lambda + API Gateway Path
\n\nFor the DBG (Dashboards Business Group) mobile application, we decided against container orchestration (ECS/EKS). Instead:
\n\n- \n
- Request handler: API Gateway endpoint routes requests to Lambda. \n
- Compute: Python/Node.js Lambda functions handle business logic (no Docker image overhead). \n
- State: Data persisted to DynamoDB or S3, depending on access patterns. \n
- Deployment: Follows the same git-driven pattern—zip the function code, upload to S3, update via boto3. \n
This sidesteps container networking complexity, simplifies IAM (Lambda execution roles are simpler than pod service accounts), and aligns with our existing operational model.
\n\nKey Architectural Decisions
\n\nWhy boto3 over CloudFormation templates? Python scripts are easier to version-control with meaningful diffs. CloudFormation JSON/YAML templates obscure conditional logic; boto3 lets us write explicit control flow.
\n\nWhy nightly deterministic tests? Our FAILURE-DOMAINS-PLAN identified infrastructure drift as a top risk. Nightly test runs catch configuration mismatches (e.g., a CloudFront cache policy manually changed in the console) before they cause user-facing issues. Tests are deterministic: same seed, same infrastructure state, reproducible outcomes.
\n\nWhy git mirror infrastructure? Artifact traceability. Every Lambda deployment, CloudFront distribution, and DNS change is tagged with the git commit that triggered it. If production behaves unexpectedly, we can ask: \"What changed in the past 24 hours?\" and get a full audit trail from git history.
\n\nWhat's Next
\n\n- \n
- Expand Python boto3 library: Standardize on a set of reusable functions for common operations (CloudFront invalidation, Lambda version tagging, Route53 batch updates). \n
- AWS Config monitoring: Detect manual console changes and alert when they diverge from git-defined state. (AWS-config-dump ticket drafted.) \n
- DBG licensure pre-check: Before mobile app launch, validate that our Lambda execution roles meet data residency and encryption requirements. (Ticket in progress.) \n
- Evaluate Terraform again at 50+ resources: Once our estate grows significantly, the overhead of managing state and HCL syntax may be justified. For now, boto3 scales with our team. \n
Our infrastructure isn't managed by a declarative framework—it's managed by intentional, auditable code. That's a trade-off we're willing to make at this scale.
", "save_location": "blog-posts"}