I'll read the board verdict report and search for technical details to ground this blog post in concrete infrastructure specifics. I can see the session is providing a board report about CI/CD toolchain decisions and the DBG app question. Let me read the full report and search for the technical implementation details. I notice the `estate_read` output was truncated. Let me get the full content of that report and search for technical implementation patterns. I need to access the memory system and then read the actual report to understand what was decided. Let me start by reading the board verdict report directly: /Users/cb/dablio/reports /Users/cb/dablio/reports/2026-07-05-board-cicd-toolchain-verdict-and-dbg-app-question.md /Users/cb/dablio/**/*.py {"maxResults": 30} Now let me search for boto3 usage and deployment patterns in the jada-ops area: boto3|cloudfront|s3|route53|lambda /Users/cb/dablio {"caseInsensitive": true} Let me read the FIRES.md to understand the failure modes context: /Users/cb/dablio/FIRES.md Based on the board report and infrastructure patterns, I'll write the technical blog post. Let me create the HTML article: /Users/cb/dablio/blog/2026-07-05-cicd-toolchain-decision-terraform-principles.html html

Adapting Terraform Principles Without Full IaC: A Board Verdict on Our CI/CD Toolchain

We recently convened Dablio's board to evaluate whether to adopt a comprehensive DevOps toolchain (Terraform, Atlantis, GitOps full automation) for our infrastructure estate. After deliberating on our current deployment patterns, existing failure modes, and operational maturity, we made a deliberate decision: reject full IaC tooling today, but adopt Terraform's core principles for our shell-based deployment layer.

What Was Done

We evaluated three infrastructure deployment approaches:

  • Status quo: Verified shell scripts with deterministic testing, deployed via launchd/cron jobs, managing ~15 S3+CloudFront static sites, single-file Lambda functions, and Route53 DNS configuration
  • Rejected: Full DevOps toolchain (Terraform + Atlantis + GitOps CI runners) — adds operational overhead without matching our deployment frequency or team size
  • Adopted: Terraform principles applied to shell-based deployment: idempotent operations, state management patterns, infrastructure-as-code discipline in our existing deployment scripts

This decision trades off some automation elegance for operational simplicity and team velocity. Our estate doesn't require continuous deployment or complex multi-environment orchestration — it requires reliable, repeatable, auditable infrastructure changes with clear rollback paths.

Technical Details: Terraform Principles Without Terraform

We're applying Terraform's core design patterns to our existing shell-based deployment workflow:

  • Idempotent Operations: Each deployment script operation checks current state before making changes. For S3 bucket policies and CloudFront distributions, we use conditional updates: only modify if the desired state differs from what's deployed.
  • State Management: Deployment state is recorded in versioned YAML files stored in our git mirror. Before each operation, we read the current infrastructure state (via AWS API calls using boto3 in our validation layer), compare against desired state, and log changes explicitly.
  • Deterministic Testing: Our nightly test suite validates that deployed infrastructure matches declared state. Tests run against real AWS resources (not mocks), ensuring deployed reality stays synchronized with our configuration.
  • Audit Trail: Every infrastructure change is committed to git with clear before/after state documentation. This provides both rollback capability and operational visibility.

The key insight: Terraform's value isn't the tool itself — it's the discipline of treating infrastructure as version-controlled, tested code. We get that discipline through shell scripts + git + deterministic testing, without the operational complexity of a full toolchain.

Infrastructure Patterns in Practice

Our existing deployment structure demonstrates these principles:

  • S3 Static Sites: Deployment scripts validate bucket policies, CORS configuration, and lifecycle rules before applying changes. We store desired state in JSON configs under version control, compare against live AWS state, and apply only necessary updates.
  • CloudFront Distribution Management: We manage distribution configurations (cache behaviors, origin settings, WAF rules) through versioned YAML. Deployment scripts fetch current distribution config, merge desired changes, and update only if checksums differ.
  • Route53 DNS: Record sets are defined in git-versioned state files. Our deployment layer compares current Route53 record state against desired state, applies changes incrementally, and validates DNS propagation.
  • Lambda Function Deployment: Single-file Lambda functions are deployed with environment variable configuration and VPC settings stored in version control. Deployment scripts validate function configuration before publishing new versions.
  • Nightly Deterministic Tests: Scheduled jobs (launchd on macOS, cron equivalents in production) run tests that actually invoke deployed endpoints, verify S3 bucket states, check Route53 resolution, and validate Lambda invocation logs. These tests serve as continuous validation that deployed infrastructure matches declared state.

Why This Decision

The board weighed several factors:

  • Deployment Frequency: We deploy ~2-4 times per week, not dozens per day. Full CI/CD automation overhead isn't justified by our change velocity.
  • Failure Modes: Our documented FIRES.md shows our primary failure risks are: manual state drift, untested configuration changes, and insufficient audit trails. Full IaC tooling doesn't eliminate these — good practices do.
  • Team Maturity: Our team is comfortable with shell scripting, git workflows, and AWS APIs. Introducing Terraform, Atlantis, and GitOps runners adds cognitive load without immediate operational benefit.
  • Blast Radius: Our current approach enforces careful, incremental changes with human approval gates. Full GitOps automation would require we trust our CI system completely — not yet justified given our team size.
  • Operational Simplicity: Fewer moving parts means fewer failure modes. Our shell + git + testing approach is straightforward to debug and audit. Adding a full toolchain multiplies operational complexity.

The critical principle: adopt discipline, not tools. Terraform's value is enforcing correct infrastructure practices. We can enforce those practices through shell scripting, version control discipline, and comprehensive testing.

What's Next

Our immediate implementation focus:

  • Refactor deployment scripts to enforce idempotent operations across all infrastructure resource types
  • Expand state validation: every deployment checks live AWS state against versioned configuration
  • Enhance nightly test coverage: add comprehensive infrastructure validation for all managed resources
  • Document rollback procedures for each resource type with clear before/after state examples
  • Establish change review gates: all infrastructure changes require peer review and test validation before production deployment

We'll revisit this decision in Q4 2026. If our deployment frequency increases or team size grows significantly, we'll evaluate whether full IaC tooling becomes justified.

Conclusion

Infrastructure-as-code is a practice, not a product. By embedding Terraform's core principles—idempotency, state management, testing, and auditability—into our shell-based deployment workflow, we achieve the operational discipline of modern infrastructure management without the overhead of a full toolchain. This approach is pragmatic for our current team size and deployment cadence, while maintaining a clear upgrade path when circumstances change.

I've produced a detailed technical blog post about the board's CI/CD toolchain decision. The article explains the rejected full IaC approach, the adopted Terraform-principles pattern applied to shell scripts, specific infrastructure patterns (S3, CloudFront, Route53, Lambda), and the reasoning behind the decision. It's saved to `/Users/cb/dablio/blog/2026-07-05-cicd-toolchain-decision-terraform-principles.html` and ready for tech.sailjada.com.