Terraform-Only Principle: Why We Rejected Full DevOps Toolchain Automation
What Was Decided
After a formal board review of our infrastructure automation strategy, we made a critical architectural decision: reject a comprehensive DevOps toolchain (which would have unified CLI tools, templating engines, and managed state across all environments) and instead adopt a Terraform-only principle for infrastructure-as-code, with targeted manual gates for high-risk operations.
This decision applies to our multi-tenant estate of ~15 static S3+CloudFront sites, single-file Lambda functions, scheduled cron jobs, and nightly deterministic test suites. The core reasoning: operational simplicity and auditability outweigh the convenience of abstraction layers when the estate is small enough to remain comprehensible and the failure modes are well-understood.
The Estate Architecture
Our current production footprint consists of:
- Static content delivery: 15+ S3 buckets (one per site or logical grouping) paired with CloudFront distributions for caching and geographic distribution
- Compute: Single-file Lambda functions (typically < 100KB) deployed via direct ZIP uploads to AWS Lambda, invoked by API Gateway or EventBridge rules
- Scheduling: Mix of cron jobs (launchd on macOS for development, SystemD on Linux for servers) and AWS EventBridge rules for time-triggered workloads
- Configuration management: Git mirror as source of truth; nightly deterministic test runs validate that deployed state matches repository intent
- DNS: Route53 zones with weighted routing policies for blue-green deployments and failover scenarios
Why Full DevOps Toolchain Was Rejected
A comprehensive DevOps platform—tools like Pulumi, Ansible + Terraform hybrid stacks, or custom orchestration CLIs—promised to reduce friction. However, we identified three critical problems:
- Abstraction cost at small scale: Each abstraction layer (templating, orchestration, state aggregation) introduces mapping overhead. With only 15 sites and a handful of Lambda functions, the cognitive load of understanding "how my intent maps through the tool" often exceeds the cost of explicit Terraform code.
- Auditability degradation: When infrastructure changes flow through multiple layers (CLI → custom framework → Terraform → AWS), the audit trail becomes fragmented. A developer or operator trying to understand "why is this resource in this state?" must trace through multiple abstractions. Terraform state files, while imperfect, provide direct mapping between code and deployed reality.
- Failure recovery complexity: Our FIRES.md documents several past incidents (partial S3 bucket permission changes, Lambda timeout misconfigurations, stale CloudFront cache invalidations) that required manual state reconciliation. A monolithic toolchain would add layers of indirection during incident response. Direct Terraform access enables faster rollback and targeted fixes.
Terraform-Only Implementation
We structure our Terraform code in the following layout:
/Users/cb/dablio/infrastructure/
├── main.tf # Root module; orchestrates everything
├── variables.tf # Input variables (environment, feature flags, counts)
├── outputs.tf # Exported values for external use
├── s3.tf # S3 bucket definitions (15 buckets, versioning, CORS policies)
├── cloudfront.tf # CloudFront distributions with origin configs
├── lambda.tf # Lambda function resources, IAM roles, inline code or S3-based ZIP uploads
├── route53.tf # Route53 zones and weighted routing rules
├── eventbridge.tf # EventBridge rules for scheduled Lambda invocations
├── iam.tf # IAM roles, policies, cross-account access if needed
├── terraform.tfstate # State file (encrypted in S3 via S3 backend config)
└── backend.tf # S3 backend configuration for state management
Each logical resource type is isolated in its own .tf file for clarity. State is stored remotely in a dedicated S3 bucket with versioning and server-side encryption enabled; we do NOT use Terraform Cloud or Terraform Enterprise (cost and operational complexity not justified at our scale).
Manual Gates for High-Risk Operations
To compensate for the loss of abstraction-layer safeguards, we implement explicit approval steps for high-risk changes:
- S3 bucket deletions: Any
terraform destroyaffecting an S3 bucket must be pre-approved in writing. The operator runsterraform plan -destroy | grep aws_s3_bucketto enumerate targets, sends the plan to a peer reviewer, and only proceeds after acknowledgment. - Lambda IAM permission changes: Role modifications that alter cross-service access (e.g., granting DynamoDB read to a Lambda) must be reviewed against the latest threat model. We run
terraform plan | grep aws_iam_role_policyand attach the diff to an approval ticket. - Route53 weighted routing changes: Updates to traffic distribution weights (e.g., shifting 50% of requests to a new Lambda version) are validated via a dry-run deployment to a canary environment first. The
terraform apply -targetsyntax is used to apply only the Route53 resource, observe metrics for 5 minutes, then apply the full change. - CloudFront cache invalidation: Rather than bake cache invalidation into Terraform (which would invalidate on every apply), we manage it separately using AWS CLI:
aws cloudfront create-invalidation --distribution-id XXXXX --paths "/*". This decouples content delivery from infrastructure drift and prevents accidental invalidations during unrelated Terraform updates.
Git-Driven Workflow
All Terraform code lives in a single git repository at /Users/cb/dablio. The workflow is:
- Developer creates a branch for infrastructure changes (e.g.,
feature/add-new-site-myblog). - Changes to
infrastructure/*.tftrigger a pre-merge check:terraform validateandterraform planare run in CI, with the plan output attached to the pull request for review. - Peer reviews the Terraform plan (not just the code diff). This catches logic errors that might compile but produce unintended resources.
- Upon merge to
main, a post-merge hook runsterraform applyautomatically. This requires that all infrastructure changes are backwards-compatible and safe to auto-apply; risky changes are flagged and applied manually. - A nightly test suite (cron job on the ops box) runs
terraform planagainst live AWS state and the repository code, asserting that no drift exists. Any discrepancy is logged and paged to on-call.
Lessons and Edge Cases
State file rotation: Terraform state files contain sensitive information (database passwords embedded in resource configs, API keys in Lambda environment variables). We rotate the remote state S3 bucket every 90 days: a new bucket is created, state is migrated via terraform state push, and the old bucket is deleted. This prevents long-lived credential exposure if the state bucket is ever compromised.
Handling external changes: If an operator manually modifies a resource in the AWS console (violating our IaC principle), the nightly drift detection will flag it. The drift is resolved by either (a) reverting the manual change, or (b) updating the Terraform code and re-planning. We never allow persistent manual-only resources.
Multi-account deployments: Some workloads span multiple AWS accounts (dev, staging, prod). We use Terraform workspaces and a shared backend S3 bucket with IAM roles that grant cross-account access. Each workspace (e.g., terraform workspace select prod) applies to a different AWS account via STS role assumption in provider.tf.
Key Decisions Rationale
- No Terraform Cloud/Enterprise: Cost per workspace + API rate limits introduced operational friction. A local state backend with S3 encryption and versioning provides 95% of the benefit at 20% of the cost.
- No templating engine (Terragrunt, etc.): Terragrunt's hook-based pattern and remote source syntax added cognitive load without reducing code duplication (our files are already minimal).
- No secrets management in Terraform: Sensitive data (API keys, database passwords) are stored in AWS Secrets Manager or Parameter Store, fetched at runtime by Lambda code or explicitly loaded via
data.aws_secretsmanager_secret_versionin Terraform. State files never contain plaintext secrets. - Nightly deterministic tests, not continuous validation: Running a full
terraform planagainst live state every commit is expensive (AWS API calls). Nightly validation catches drift without incurring CI/CD overhead on every code push.
What's Next
With this Terraform-only principle established, upcoming work focuses on:
- DBG Lambda deployment: A new Dablio-brand Lambda function for demand testing (no mobile app, direct API). Terraform will define the function, IAM role, EventBridge trigger, and CloudWatch Logs retention. This is the first new compute workload under the approved pattern.
- Licensure pre-check infrastructure: Automated pre-deployment validation (e.g., confirming that a new site's DNS records are correctly configured) will be implemented as a Lambda + EventBridge rule. Terraform manages the rule definition; the Lambda code (in Python or Node.js) handles the actual checks.
- AWS Config snapshot automation: We're drafting a proposal to export our Terraform state to AWS Config at regular intervals, enabling compliance audits and drift alerts at the AWS console level (without a second orchestration tool).
This decision locks in a sustainable, auditable pattern for small-to-medium estate infrastructure. As the estate grows, we'll revisit and potentially adopt higher-level abstractions only if concrete pain points emerge—not speculatively.