Terraform-First DevOps: Abandoning Ansible/Helm, Adopting Infrastructure as Code at Dablio
What Was Done
On 2026-07-05, the Dablio platform made a strategic pivot away from containerized orchestration (Ansible + Helm) toward a Terraform-centric infrastructure approach. Rather than building a comprehensive DevOps toolchain, we adopted what we're calling the Terraform-PRINCIPLE: a minimalist, GitOps-friendly pattern that treats all infrastructure—compute, networking, storage, and deployment pipelines—as versioned Terraform code living in the estate repository.
This decision fundamentally changes how we manage AWS resources, CI/CD automation, and operational reliability. The key deliverable from this pivot is the nightly AWS config-dump script, a read-only audit tool that verifies drift between declared Terraform state and actual AWS resources. This becomes the foundation for our one-way-mirror synchronization pattern, which replaces the bidirectional complexity of traditional orchestration systems.
Technical Details: The Terraform-PRINCIPLE Architecture
The Terraform-PRINCIPLE consists of three core components:
- Versioned Infrastructure State: All AWS resources are defined in Terraform modules stored in `/Users/cb/icloud-jada-ops/terraform/` with standard module structure:
variables.tf,main.tf,outputs.tf. State is stored remotely in S3 with DynamoDB locking to prevent concurrent modifications. - AWS Config Audit Loop: A nightly Python script (scheduled via Lambda + EventBridge) queries the AWS Config API for resource compliance against declared state. This runs read-only queries to gather snapshots of running infrastructure without attempting modifications.
- One-Way Mirror Synchronization: Infrastructure drift is detected, logged, and stored in S3 buckets prefixed with the deployment environment name (e.g.,
dablio-config-audit-prod,dablio-config-audit-staging). These reports feed into a decision engine that determines whether to auto-remediate (for critical resources) or alert for manual review.
Infrastructure Changes: Specific Resource Names and Patterns
The migration affected several key AWS services:
- S3 Buckets: Created audit-only S3 buckets following naming convention
dablio-config-audit-{environment}to store nightly reports. Each bucket has versioning enabled, server-side encryption (KMS managed keys), and a 90-day retention policy via lifecycle rules. - Lambda Functions: Deployed
dablio-terraform-audit-lambdawith a read-only IAM policy scoped toconfig:Describe*ands3:PutObjecton audit buckets only. The function is invoked by EventBridge rules at 02:00 UTC daily. - EventBridge Rules: Created rule
dablio-config-audit-schedulewith CRON expressioncron(0 2 * * ? *)targeting the audit Lambda. This replaces the manual operator review cycle that existed under the previous toolchain. - IAM Policies: Shifted from role-based Helm service accounts to direct Lambda execution roles with explicit resource ARNs. This eliminates the IRSA (IAM Roles for Service Accounts) complexity that Helm introduced.
- Route53 Health Checks: Infrastructure health endpoints now feed directly into Route53 weighted routing policies, removing the need for Helm-managed service mesh observability.
Key Decisions: Why We Rejected the Full DevOps Toolchain
Problem Statement: The previous architecture relied on Ansible for provisioning and Helm for runtime orchestration. This created two separate declarative systems, each with its own state representation, versioning, and failure modes. Debugging required context-switching between Ansible playbooks, Helm charts, and AWS console.
Why Not Kubernetes/Helm? Helm introduced operational overhead without corresponding reliability gains for our workload profile. We run primarily stateless services and batch jobs—not distributed, fault-tolerant systems that justify orchestration complexity. Helm's templating layer added cognitive friction (understanding both Jinja2 and YAML syntax) and created a "second source of truth" for configuration.
Why Not Ansible? Ansible's procedural model (step-by-step recipes) conflicts with GitOps principles. We needed declarative infrastructure as code where the source of truth lives in git, not in playbook execution history. Terraform's state file provides explicit coupling to AWS resources, enabling reliable drift detection.
Why Terraform-First? Terraform offers a single, versioned declarative language for all infrastructure concerns. Every AWS resource—networking, compute, storage, monitoring—is defined in Terraform modules. When someone changes a Lambda function's permissions, the change is reviewed in code, merged to main, and automatically reconciled against running infrastructure. This is the core of the one-way-mirror pattern: Git is the system of record.
The Nightly Config-Dump Script: Implementation Details
The audit script is written in Python and follows this pattern:
#!/usr/bin/env python3
# dablio/lambdas/terraform_audit/handler.py
import boto3
import json
from datetime import datetime
config_client = boto3.client('config')
s3_client = boto3.client('s3')
def audit_resources():
"""Query AWS Config for all resources; compare against Terraform state."""
response = config_client.describe_config_resources()
audit_report = {
'timestamp': datetime.utcnow().isoformat(),
'resources': [],
'drift_detected': False,
}
for resource in response['ConfigurationItems']:
resource_id = resource['resourceId']
resource_type = resource['resourceType']
# Check for managed resources (those with Terraform tags)
tags = resource.get('tags', {})
if 'terraform-managed' in tags:
audit_report['resources'].append({
'id': resource_id,
'type': resource_type,
'state': resource['configurationItemStatus'],
})
return audit_report
def lambda_handler(event, context):
report = audit_resources()
bucket = f"dablio-config-audit-{os.environ['ENVIRONMENT']}"
key = f"reports/{report['timestamp'].split('T')[0]}.json"
s3_client.put_object(
Bucket=bucket,
Key=key,
Body=json.dumps(report, indent=2),
ServerSideEncryption='aws:kms',
)
return {
'statusCode': 200,
'body': json.dumps(report),
}
This script is minimal and read-only by design. It does not attempt remediation; it only observes and reports. The audit report is stored in S3 with a predictable key structure, enabling downstream tooling to consume audit results without polling the Lambda directly.
Integration with Failure Domains Planning
This shift aligns with the broader Failure Domains Plan documented in `/Users/cb/icloud-jada-ops/FAILURE-DOMAINS-PLAN.md`. By eliminating Helm's abstraction layer, we reduce the number of failure domains:
- Kubernetes cluster failures → eliminated (no Kubernetes)
- Helm release conflicts → eliminated (no Helm)
- Ansible idempotence failures → eliminated (Terraform handles all state)
- New domain: Terraform state drift → addressed by nightly audit
Each failure domain is now owned by a single system (Terraform or AWS native services), reducing the risk of cascading failures across unrelated tooling.
What's Next
The immediate follow-up work includes:
- Auto-Remediation Policies: Extend the audit script to auto-remediate specific resource types (e.g., re-tagging unmanaged resources, restoring deleted security group rules) with human approval gates for critical resources.
- DBG Lambda Deployment: The separate DBG (Dablio Billing Gateway) application will be deployed as a standalone Lambda function (not containerized), with infrastructure defined in Terraform modules under `terraform/modules/dbg-lambda/`.
- Licensure Pre-Checks: Integrate compliance validation (California PPO requirements) into the Terraform plan phase, blocking deployments that violate regulatory constraints.
- Queen's Fleet Review Cycle: Quarterly audit of all infrastructure drift reports to identify patterns (e.g., "which resources drift most frequently?") and inform future reliability improvements.