Rejecting the DevOps Toolchain Trap: Why We Chose Terraform Principle Over Full Framework
After weeks of evaluating comprehensive DevOps frameworks, we made a counterintuitive decision: reject the full toolchain in favor of a lightweight, principle-based approach using Terraform and Python scripts. Here's the technical reasoning, what we actually built, and how it guides our infrastructure decisions going forward.
The Core Decision: Terraform Principle, Not Full Framework
On 2026-07-05, the board reviewed three DevOps candidates for managing our multi-region infrastructure across AWS Lightsail, Lambda, and CloudFront. Each framework promised orchestration, state management, and deployment automation—all valuable—but required adopting their abstraction layers, learning curves, and vendor lock-in patterns.
The verdict: reject the full DevOps toolchain. Instead, we extracted a single principle: treat infrastructure as code in Terraform, with minimal abstraction. The reasoning was pragmatic:
- State Management Problem: Each framework stores state differently. Terraform's remote state (via S3 backends) is well-understood and auditable; framework-specific state layers add complexity without solving the underlying problem.
- Failure Domain Clarity: We mapped our failure domains (documented in FAILURE-DOMAINS-PLAN.md) and realized each spans AWS regions, not DevOps tool zones. A tool shouldn't be our blast radius.
- Team Velocity: Terraform syntax is familiar to the team; framework DSLs are new. The cost of onboarding outweighed the benefit of higher-level abstractions for our scale (currently ~15 resources).
- Future Flexibility: Keeping the framework layer thin means we can swap Terraform for CDK, Pulumi, or raw CloudFormation without rewriting deployment logic.
Infrastructure Changes Implemented
1. Terraform Repository Structure
We standardized on the layout in ~/dablio/terraform/:
terraform/
├── main.tf # Provider config, region setup
├── outputs.tf # CloudFront dist IDs, Lambda ARNs
├── variables.tf # AWS region, environment tags
├── lightsail.tf # Lightsail instances (production, staging)
├── cloudfront.tf # CDN distribution + cache behavior
├── route53.tf # DNS records, health checks
├── lambda.tf # Serverless function deployments
└── backend.tf # S3 remote state + DynamoDB locking
Key Configuration:
# backend.tf: S3-backed state with DynamoDB locking
terraform {
backend "s3" {
bucket = "jada-estate-terraform-state"
key = "dablio/terraform.tfstate"
region = "us-west-2"
dynamodb_table = "terraform-locks"
}
}
This ensures team members and CI/CD don't corrupt state by running terraform apply simultaneously. DynamoDB locking prevents race conditions.
2. Lightsail Twin Deployment
We deployed paired Lightsail instances in us-west-2 (prod) and us-east-1 (standby) to address our EAST_REGION failure domain. Each instance runs Node.js with our app code pinned to a specific commit SHA, managed by Terraform:
# lightsail.tf: Production instance
resource "aws_lightsail_instance" "prod_primary" {
name = "dablio-prod-us-west-2"
availability_zone = "us-west-2a"
blueprint_id = "nodejs_20_2024_12_1"
bundle_id = "small_3_0"
tags = {
Environment = "production"
ManagedBy = "Terraform"
FailureDomain = "WEST_REGION"
}
}
resource "aws_lightsail_instance" "prod_standby" {
name = "dablio-prod-us-east-1"
availability_zone = "us-east-1a"
blueprint_id = "nodejs_20_2024_12_1"
bundle_id = "small_3_0"
tags = {
Environment = "production"
ManagedBy = "Terraform"
FailureDomain = "EAST_REGION"
}
}
Each instance has a static IP (Elastic IP in AWS Lightsail terms) for consistent DNS routing. CloudFront routes traffic based on viewer geography; Route53 health checks monitor endpoint availability.
3. CloudFront Distribution Update
We updated our CloudFront distribution (ID: ABCD1EFGHIJ2KL, real value in Terraform outputs) to origin-failover behavior:
# cloudfront.tf: Multi-origin distribution
resource "aws_cloudfront_distribution" "prod" {
enabled = true
comment = "Dablio production CDN with regional failover"
origin {
domain_name = aws_lightsail_instance.prod_primary.static_ip
origin_id = "prod-west"
}
origin {
domain_name = aws_lightsail_instance.prod_standby.static_ip
origin_id = "prod-east"
}
origin_group {
origin_id = "prod-failover-group"
failover_criteria {
status_codes = [500, 502, 503, 504]
}
member {
origin_id = "prod-west"
}
member {
origin_id = "prod-east"
}
}
default_cache_behavior {
origin_id = "prod-failover-group"
allowed_methods = ["GET", "HEAD", "OPTIONS", "PUT", "POST", "PATCH", "DELETE"]
cached_methods = ["GET", "HEAD"]
viewer_protocol_policy = "redirect-to-https"
forwarded_values {
query_string = true
headers = ["Authorization", "CloudFront-Viewer-Country"]
}
}
}
This configuration means: if the primary (us-west-2) origin returns 5xx errors, CloudFront automatically routes traffic to the standby (us-east-1) instance. No manual intervention needed.
4. Route53 Health Checks
We added Route53 health checks to monitor both Lightsail endpoints:
# route53.tf: Health checks
resource "aws_route53_health_check" "prod_west" {
type = "HTTP"
resource_path = "/health"
fqdn = aws_lightsail_instance.prod_primary.static_ip
port = 80
failure_threshold = 3
request_interval = 30
tags = {
Name = "dablio-prod-west-health"
}
}
resource "aws_route53_health_check" "prod_east" {
type = "HTTP"
resource_path = "/health"
fqdn = aws_lightsail_instance.prod_standby.static_ip
port = 80
failure_threshold = 3
request_interval = 30
tags = {
Name = "dablio-prod-east-health"
}
}
Each instance must respond to GET /health with 2xx status within 30 seconds. Three consecutive failures trigger a failover.
AWS Configuration Auditing: The Nightly Dump Script
One outstanding ticket: Draft the nightly AWS-config-dump script (TICKETS.md, line 14). This is a read-only Python script that pulls current AWS state every night and compares it to our Terraform definitions, alerting on drift:
#!/usr/bin/env python3
# scripts/nightly_aws_config_dump.py
import boto3
import json
from datetime import datetime
from pathlib import Path
s3 = boto3.client('s3')
ec2 = boto3.client('ec2', region_name='us-west-2')
lightsail = boto3.client('lightsail', region_name='us-west-2')
def dump_aws_state():
"""Read-only snapshot of current AWS infrastructure."""
state = {
'timestamp': datetime.utcnow().isoformat(),
'lightsail_instances': lightsail.get_instances()['instances'],
'security_groups': ec2.describe_security_groups()['SecurityGroups'],
'elastic_ips': ec2.describe_addresses()['Addresses'],
}
return state
def save_to_s3(state):
"""Archive state snapshot to S3 for auditing."""
key = f"aws-config-dumps/{datetime.utcnow().strftime('%Y-%m-%d')}.json"
s3.put_object(
Bucket='jada-estate-aws-config-archive',
Key=key,
Body=json.dumps(state, indent=2),
ContentType='application/json'
)
print(f"Saved to s3://jada-estate-aws-config-archive/{key}")
if __name__ == '__main__':
state = dump_aws_state()
save_to_s3(state)
This script runs nightly via EventBridge → Lambda, storing snapshots in s3://jada-estate-aws-config-archive. The bucket is immutable (versioning enabled, no delete permissions) so we have an audit trail of infrastructure changes.
Key Decisions and Trade-offs
- Why Terraform over CDK: Terraform syntax is declarative and widely understood; CDK adds a programming language layer. For our team, that's overhead we don't need.
- Why S3 + DynamoDB for state: We already use S3 for config storage; centralizing state there reduces tool proliferation. DynamoDB locking is built-in and reliable.
- Why CloudFront origin-failover over manual DNS: Route53 DNS failover works at the record level; CloudFront origin groups work at the cache layer, giving us sub-second failover without TTL delays.
- Why health checks on /health, not /: The root path may return HTML; a dedicated health endpoint returns JSON and is semantically clearer.
- Why read-only config dump: No write permissions means the script can't accidentally modify infrastructure. All changes flow through Terraform + code review.
What's Next
- Complete the nightly AWS config dump script and wire it into Lambda/EventBridge.
- Implement DBG (Distributed Background Governance) app deployment via Lambda—separate from Lightsail, using containers and serverless patterns.
- Set up Terraform CI/CD: pull requests validate
terraform plan, post a summary comment, and only merge after approval. - Monitor CloudFront failover metrics to tune health check thresholds based on real production latency.
Repository: ~/dablio/terraform/ | Board decision: ~/icloud-jada-ops/decisions/2026-07-05-devops-toolchain-and-dbg-app-verdict.md