Terraform-First DevOps: Why We Rejected the Full Toolchain and Adopted Managed Services
\n\nWhat Was Done
\n\nAfter a comprehensive infrastructure review, the Jada Estate and Dablio teams made a deliberate architectural decision: reject a unified DevOps toolchain (like Terraform Cloud + Atlantis + custom CI/CD) in favor of a Terraform-principle-only approach paired with AWS managed services. This post covers the decision rationale, implementation strategy, and the immediate next steps in infrastructure automation.
\n\nThe Original Problem
\n\nThe estate infrastructure had grown organically across multiple AWS accounts (dev, staging, production), CloudFront distributions, Route53 zones, and S3 buckets. State management was inconsistent, deployment workflows were manual, and there was no single source of truth for infrastructure configuration. A full DevOps toolchain seemed like the right solution—until we ran the numbers.
\n\nWhy We Rejected a Full Toolchain
\n\nA traditional DevOps stack would have included:
\n\n- \n
- Terraform Cloud for state management and runs \n
- Atlantis for pull request-based workflow automation \n
- Custom CI/CD orchestration (GitHub Actions or similar) \n
- Compliance and audit logging layers \n
The blocker: operational overhead. These tools require dedicated maintenance, custom policy code, approval workflows, and substantial configuration complexity. For a team of 2-3 people managing infrastructure across 3-4 AWS accounts, this represented >40% operational tax before delivering any new infrastructure.
\n\nThe insight: AWS has already solved most of these problems. CloudFormation StackSets, Systems Manager Parameter Store, AWS Config, and Terraform's native AWS provider give us state management, change tracking, and drift detection—without the tooling overhead.
\n\nThe Terraform-Principle Approach
\n\nWe committed to three hard rules:
\n\n- \n
- All infrastructure is code. No console clicks. Every resource lives in Terraform modules under version control at `~/dablio/infrastructure/` (or estate-equivalent paths). \n
- State is remote and locked. S3 backend with DynamoDB locking (no local state). State files are encrypted at rest and in transit. \n
- Changes are peer-reviewed. Every Terraform apply requires a `terraform plan` output review in a pull request, signed approval from at least one other engineer, and a CI check that validates syntax and policy compliance. \n
This trades off some automation (no auto-merge workflows) for simplicity and safety. A PR-based approach scales better to our team size and reduces the blast radius of a misconfiguration.
\n\nInfrastructure Decisions
\n\nState Backend: We chose S3 + DynamoDB over Terraform Cloud for cost and control. The S3 bucket is named according to estate naming conventions (read-only outside the automation account), versioning is enabled, and MFA delete is configured. DynamoDB table uses on-demand billing to avoid over-provisioning.
\n\nAWS Config for Drift Detection: Terraform is our source of truth, but AWS Config watches for manual changes or drift. A nightly Lambda function (deployed via Terraform, naturally) queries Config and alerts if any managed resource has drifted from its desired state. This catches both accidents and permission creep.
\n\nSecrets Management: Sensitive values (API keys, database passwords, OAuth tokens) are stored in AWS Secrets Manager and referenced in Terraform via `data \"aws_secretsmanager_secret_version\"`. The Terraform code itself is open (in a private repo) but never contains credentials.
\n\nImmediate Deliverables: The AWS Config Dump Script
\n\nThe first concrete task from this decision is a nightly read-only script that audits infrastructure state. Located at `~/dablio/scripts/aws-config-dump.py`, this Python script:
\n\n- \n
- Runs via Lambda on a CloudWatch Events schedule (2 AM daily, UTC). \n
- Queries AWS Config's aggregator across all accounts to enumerate all managed resources. \n
- Compares Config's recorded state against Terraform's `.tfstate` for each resource type. \n
- Generates a human-readable report and stores it in S3 at `s3://jada-estate-audit-logs/config-dumps/YYYY-MM-DD.json`. \n
- Sends a CloudWatch Logs summary to a team Slack channel if drift is detected. \n
The script uses boto3 with AWS credentials provided via IAM roles (no hardcoded keys), and it's idempotent—safe to run multiple times without side effects.
\n\nWhat's Next
\n\nWith the Terraform foundation in place, we're unblocking three workstreams:
\n\n- \n
- DBG Lambda Deployment: Migrate DBG (a data processing service) from EC2 to Lambda functions, managed entirely via Terraform. This reduces operational overhead and scales to zero during off-peak hours. \n
- Compliance Checkpoints: Integrate Terraform Cloud's policy-as-code (or open-source Sentinel equivalent) to enforce tagging standards, encryption requirements, and account isolation rules automatically. \n
- Cross-Account Multi-Region State: Extend Terraform modules to support active-active deployments across us-east-1 and eu-west-1, with automatic failover via Route53 health checks. \n
The Terraform-principle approach is intentionally lightweight. It gives us 80% of the value of a full DevOps platform with 20% of the operational cost—and more importantly, it lets us move fast without choking on tooling.
\n\nKey Takeaway
\n\nDon't confuse tooling maturity with operational maturity. A powerful DevOps platform can actually slow you down if your team is small or if the underlying infrastructure is still in motion. Start with Terraform, S3 state, peer review, and AWS managed services. Add orchestration tools only when you've hit their limits, not before.
"}}