Adapting Terraform Principles Without Full IaC: A Board Verdict on Our CI/CD Toolchain
We recently convened Dablio's board to evaluate whether to adopt a comprehensive DevOps toolchain (Terraform, Atlantis, GitOps full automation) for our infrastructure estate. After deliberating on our current deployment patterns, existing failure modes, and operational maturity, we made a deliberate decision: reject full IaC tooling today, but adopt Terraform's core principles for our shell-based deployment layer.
What Was Done
We evaluated three infrastructure deployment approaches:
- Status quo: Verified shell scripts with deterministic testing, deployed via launchd/cron jobs, managing ~15 S3+CloudFront static sites, single-file Lambda functions, and Route53 DNS configuration
- Rejected: Full DevOps toolchain (Terraform + Atlantis + GitOps CI runners) — adds operational overhead without matching our deployment frequency or team size
- Adopted: Terraform principles applied to shell-based deployment: idempotent operations, state management patterns, infrastructure-as-code discipline in our existing deployment scripts
This decision trades off some automation elegance for operational simplicity and team velocity. Our estate doesn't require continuous deployment or complex multi-environment orchestration — it requires reliable, repeatable, auditable infrastructure changes with clear rollback paths.
Technical Details: Terraform Principles Without Terraform
We're applying Terraform's core design patterns to our existing shell-based deployment workflow:
- Idempotent Operations: Each deployment script operation checks current state before making changes. For S3 bucket policies and CloudFront distributions, we use conditional updates: only modify if the desired state differs from what's deployed.
- State Management: Deployment state is recorded in versioned YAML files stored in our git mirror. Before each operation, we read the current infrastructure state (via AWS API calls using boto3 in our validation layer), compare against desired state, and log changes explicitly.
- Deterministic Testing: Our nightly test suite validates that deployed infrastructure matches declared state. Tests run against real AWS resources (not mocks), ensuring deployed reality stays synchronized with our configuration.
- Audit Trail: Every infrastructure change is committed to git with clear before/after state documentation. This provides both rollback capability and operational visibility.
The key insight: Terraform's value isn't the tool itself — it's the discipline of treating infrastructure as version-controlled, tested code. We get that discipline through shell scripts + git + deterministic testing, without the operational complexity of a full toolchain.
Infrastructure Patterns in Practice
Our existing deployment structure demonstrates these principles:
- S3 Static Sites: Deployment scripts validate bucket policies, CORS configuration, and lifecycle rules before applying changes. We store desired state in JSON configs under version control, compare against live AWS state, and apply only necessary updates.
- CloudFront Distribution Management: We manage distribution configurations (cache behaviors, origin settings, WAF rules) through versioned YAML. Deployment scripts fetch current distribution config, merge desired changes, and update only if checksums differ.
- Route53 DNS: Record sets are defined in git-versioned state files. Our deployment layer compares current Route53 record state against desired state, applies changes incrementally, and validates DNS propagation.
- Lambda Function Deployment: Single-file Lambda functions are deployed with environment variable configuration and VPC settings stored in version control. Deployment scripts validate function configuration before publishing new versions.
- Nightly Deterministic Tests: Scheduled jobs (launchd on macOS, cron equivalents in production) run tests that actually invoke deployed endpoints, verify S3 bucket states, check Route53 resolution, and validate Lambda invocation logs. These tests serve as continuous validation that deployed infrastructure matches declared state.
Why This Decision
The board weighed several factors:
- Deployment Frequency: We deploy ~2-4 times per week, not dozens per day. Full CI/CD automation overhead isn't justified by our change velocity.
- Failure Modes: Our documented FIRES.md shows our primary failure risks are: manual state drift, untested configuration changes, and insufficient audit trails. Full IaC tooling doesn't eliminate these — good practices do.
- Team Maturity: Our team is comfortable with shell scripting, git workflows, and AWS APIs. Introducing Terraform, Atlantis, and GitOps runners adds cognitive load without immediate operational benefit.
- Blast Radius: Our current approach enforces careful, incremental changes with human approval gates. Full GitOps automation would require we trust our CI system completely — not yet justified given our team size.
- Operational Simplicity: Fewer moving parts means fewer failure modes. Our shell + git + testing approach is straightforward to debug and audit. Adding a full toolchain multiplies operational complexity.
The critical principle: adopt discipline, not tools. Terraform's value is enforcing correct infrastructure practices. We can enforce those practices through shell scripting, version control discipline, and comprehensive testing.
What's Next
Our immediate implementation focus:
- Refactor deployment scripts to enforce idempotent operations across all infrastructure resource types
- Expand state validation: every deployment checks live AWS state against versioned configuration
- Enhance nightly test coverage: add comprehensive infrastructure validation for all managed resources
- Document rollback procedures for each resource type with clear before/after state examples
- Establish change review gates: all infrastructure changes require peer review and test validation before production deployment
We'll revisit this decision in Q4 2026. If our deployment frequency increases or team size grows significantly, we'll evaluate whether full IaC tooling becomes justified.
Conclusion
Infrastructure-as-code is a practice, not a product. By embedding Terraform's core principles—idempotency, state management, testing, and auditability—into our shell-based deployment workflow, we achieve the operational discipline of modern infrastructure management without the overhead of a full toolchain. This approach is pragmatic for our current team size and deployment cadence, while maintaining a clear upgrade path when circumstances change.