Rejecting the Full DevOps Toolchain: Why We Chose Terraform Principles Over Infrastructure-as-Code Frameworks
What Was Done
After a comprehensive evaluation of DevOps toolchain options (including Terraform, Pulumi, and CloudFormation), we made a deliberate architectural decision: reject a unified infrastructure-as-code framework in favor of a hybrid approach that preserves Terraform principles while using AWS-native services and Lambda-based automation for specific deployment patterns.
This decision directly impacts how we deploy the Dablio Business Group (DBG) application layer and manage AWS infrastructure at scale. Rather than standardizing on a single framework, we're adopting a targeted Lambda deployment model for ephemeral workloads combined with minimal Terraform for state management of persistent resources.
Why This Matters: The Decision Rationale
The full DevOps toolchain evaluation revealed critical tradeoffs:
- Framework Lock-In Risk: Adopting a comprehensive framework (Terraform Enterprise, Pulumi SaaS, CDK) would require migrating existing deployment logic and centralizing all infrastructure decisions through a single abstraction layer. This creates operational brittle points.
- Licensing & Cost: Enterprise tiers of most frameworks introduce per-resource or per-operation billing that doesn't scale linearly with infrastructure complexity in our use case.
- Debuggability at Production Scale: When infrastructure automation fails, framework abstractions can obscure the actual AWS API calls being made, making root-cause analysis difficult under time pressure.
- Team Velocity on Small Changes: For iterative deployments (hotfixes, compliance updates), invoking a full framework pipeline often takes longer than direct AWS API calls with proper guardrails.
What We're Building Instead
Terraform for Persistent State
We'll continue using Terraform for infrastructure that requires durable state management:
- S3 bucket configurations: Versioning, encryption policies, and lifecycle rules defined in
terraform/s3.tf - Route53 DNS records: CNAME and A records pointing to CloudFront distributions and ALB endpoints
- CloudFront distributions: Cache behaviors, origin configurations, and WAF rule associations
- IAM roles and policies: Least-privilege service roles for Lambda, ECS, and database access
Example Terraform pattern for S3 + CloudFront:
resource "aws_s3_bucket" "app_assets" {
bucket = "dablio-app-assets-prod"
acl = "private"
}
resource "aws_cloudfront_distribution" "app_cdn" {
origin {
domain_name = aws_s3_bucket.app_assets.bucket_regional_domain_name
s3_origin_config {
origin_access_identity = aws_cloudfront_origin_access_identity.oai.cloudfront_access_identity_path
}
}
default_cache_behavior {
allowed_methods = ["GET", "HEAD"]
cached_methods = ["GET", "HEAD"]
target_origin_id = "S3Origin"
compress = true
}
}
Lambda-Driven Deployments for DBG
For the Dablio Business Group application deployments, we're implementing a Lambda-based deployment orchestrator that invokes AWS services directly using boto3, rather than wrapping everything through Terraform:
- Deployment Function:
dbg-deploy-orchestratorLambda function (Python 3.11 runtime) - Input: Application artifact S3 path, environment target (staging/prod), feature flags
- Execution: Boto3 calls to CloudFormation, ECS task definition registration, SNS notifications
- State Storage: Minimal metadata stored in DynamoDB table
dbg-deployments-statefor audit trails
This approach gives us:
- Fine-grained control: The Lambda can implement custom validation (license compliance checks, dependency version locks) before invoking AWS APIs
- Fast iteration: Code changes to the orchestrator don't require Terraform plan/apply cycles
- Integration testing: Deployments can be tested against a real AWS account in a staging environment without mocking infrastructure
Infrastructure Guardrails via Lambda
Rather than relying on framework constraints, we're implementing infrastructure validation logic directly in the deployment function:
import boto3
from botocore.exceptions import ClientError
def validate_deployment_target(account_id, environment):
"""Verify target account matches environment policies"""
sts = boto3.client('sts')
identity = sts.get_caller_identity()
allowed_accounts = {
'prod': '123456789012',
'staging': '210987654321'
}
if identity['Account'] != allowed_accounts[environment]:
raise ValueError(f"Account mismatch for {environment}")
return True
def register_ecs_task_definition(app_version, environment):
"""Register new ECS task definition for DBG"""
ecs = boto3.client('ecs')
task_def = ecs.register_task_definition(
family=f'dbg-{environment}',
networkMode='awsvpc',
requiresCompatibilities=['FARGATE'],
cpu='512',
memory='1024',
containerDefinitions=[
{
'name': f'dbg-app-{environment}',
'image': f'123456789012.dkr.ecr.us-east-1.amazonaws.com/dbg:{app_version}',
'portMappings': [{'containerPort': 8080, 'protocol': 'tcp'}],
'logConfiguration': {
'logDriver': 'awslogs',
'options': {
'awslogs-group': f'/ecs/dbg-{environment}',
'awslogs-region': 'us-east-1',
'awslogs-stream-prefix': 'ecs'
}
}
}
]
)
return task_def['taskDefinition']['taskDefinitionArn']
Key Infrastructure Components
- ECS Cluster:
dablio-prod-cluster(Fargate, on-demand capacity) - RDS Database: PostgreSQL 14.x with automated backups, managed by Terraform
- CloudWatch Logs: DBG application logs stream to
/ecs/dbg-prodand/ecs/dbg-staginglog groups - DynamoDB Tables:
dbg-deployments-state— audit trail of all deploymentsdbg-license-compliance— tracks licensure requirements per deployment
- SNS Topics:
dbg-deployment-notificationsfor Slack/email alerts on deployment events
What's Next
This architecture unfolds in phases:
- Week of Jul 8: Finalize Terraform state backend configuration (S3 + DynamoDB for locking)
- Week of Jul 15: Deploy initial DBG Lambda orchestrator function and integration tests against staging account
- Week of Jul 22: Blue-green deployment test with real ECS service and load balancer failover
- Early August: Production rollout with monitoring dashboards in CloudWatch and compliance audit logging
The DBG application will be deployable via a single Lambda invocation (triggered by CI/CD or manual approval), with all infrastructure validation, compliance checks, and rollback logic embedded in the orchestration function rather than delegated to a framework.
For Engineers Reading This
If you're maintaining this system:
- Run
terraform planbefore infrastructure changes to review S3, CloudFront, and Route53 modifications - Check DynamoDB
dbg-deployments-statetable for a complete audit trail of all past deployments - Lambda function source code lives in
infrastructure/lambda/dbg-deploy-orchestrator/with unit tests intests/integration/ - Deployment logs appear in CloudWatch Logs under
/lambda/dbg-deploy-orchestrator— monitor this for any credential or API errors - Test a deployment to staging before touching production: use the same Lambda with
environment=stagingparameter