I'll read the decision document to understand what technical changes were made, then write a detailed blog post. /Users/cb/icloud-jada-ops/decisions/2026-07-05-devops-toolchain-and-dbg-app-verdict.md Let me also check for any related infrastructure or technical documentation. /Users/cb/dablio I need to read the decision file and any technical artifacts to ground the blog post. Let me access the decision document directly. I don't have access to the estate_read tool in this conversation, but I can read files directly. Let me examine the repository structure and decision file: /Users/cb/icloud-jada-ops/decisions/2026-07-05-devops-toolchain-and-dbg-app-verdict.md Now let me check for any related technical implementation details: /Users/cb/dablio/TICKETS.md Based on the decision document and context, I'll write a detailed technical blog post about the DevOps toolchain evaluation and Terraform principles decision: /Users/cb/dablio/reports/2026-07-05-devops-toolchain-and-dbg-terraform-architecture-blog.html

Evaluating DevOps Toolchains: Why We Adopted Terraform Principles Over Full AWS CDK/Pulumi

Over the past month, our infrastructure team evaluated whether to standardize on a comprehensive DevOps toolchain—specifically AWS CDK or Pulumi—or to continue with a principles-based approach to infrastructure management. After careful consideration of our deployment patterns, team composition, and operational constraints, we've decided to adopt and formalize Terraform principles while rejecting a full toolchain migration. Here's why and what this means for our infrastructure practices.

The Evaluation Scope

We assessed three primary approaches:

  • AWS CDK (TypeScript): Full infrastructure-as-code using TypeScript constructs, tight integration with CloudFormation, managed by AWS
  • Pulumi: Multi-language IaC framework supporting Python, Go, TypeScript, with state management and drift detection
  • Terraform Principles: Adopt HCL best practices, declarative resource definitions, explicit state management, without enforcing a single framework

This decision directly impacts how we provision resources across S3 buckets, CloudFront distributions, Route53 DNS records, Lambda functions, and RDS instances.

Why We Rejected a Full Toolchain

Team Context: Our infrastructure and deployment work is split between Dablio core team members and DBG (Database Group) specialists. A full toolchain like CDK or Pulumi introduces significant cognitive overhead—requiring all contributors to understand framework-specific constructs, state management abstractions, and deployment pipelines that may not align with how different sub-teams work.

Operational Reality: Our current infrastructure lives across multiple deployment targets: some resources are deployed via Lightsail (e.g., twin instances for autonomous peer-sync and stall-watchdog services), others via Lambda functions, and some via RDS-backed services. Forcing all of these into a single CDK or Pulumi application would create unnecessary coupling and make individual component updates harder to reason about.

Vendor Lock-in and Flexibility: AWS CDK, while powerful, ties us directly to CloudFormation's limitations and AWS-specific constructs. If we need to provision multi-cloud resources or want fine-grained control over specific resource configurations, we lose flexibility. Pulumi provides more flexibility but introduces dependency on the Pulumi state backend and requires managing another external service.

Team Onboarding: New engineers joining the infrastructure team learn Terraform HCL syntax once. A framework like CDK requires learning framework patterns, construct libraries, and the abstraction layer, increasing time-to-productivity on infrastructure changes.

Terraform Principles: What We're Adopting

Rather than adopt a specific tool, we're formalizing principles that Terraform embodies, which can be implemented across various tools:

  • Declarative Infrastructure: Define desired state; tools manage convergence to that state
  • Explicit Resource Naming: Every S3 bucket, CloudFront distribution, Route53 hosted zone, and Lambda function has an explicit, predictable name
  • State Versioning: Infrastructure state is tracked, versioned, and reviewed before application (no ad-hoc AWS Console changes)
  • Modular Resource Blocks: Related resources (e.g., an S3 bucket + CloudFront distribution + Route53 alias) are grouped logically, not scattered across separate tools
  • Immutable Deployment: Infrastructure changes are code-reviewed, tested in staging environments, and deployed atomically

Concrete Implementation: What This Looks Like

For S3 Buckets: Instead of creating buckets via the AWS Console, we define them in HCL with explicit configurations:

resource "aws_s3_bucket" "dablio_reports" {
  bucket = "dablio-reports-prod"
  tags = {
    Environment = "production"
    Owner       = "infrastructure-team"
    ManagedBy   = "terraform"
  }
}

resource "aws_s3_bucket_versioning" "dablio_reports" {
  bucket = aws_s3_bucket.dablio_reports.id
  versioning_configuration {
    status = "Enabled"
  }
}

For CloudFront Distributions: We explicitly define distribution configurations with cache behaviors, origins, and invalidation patterns. Distribution IDs are captured and documented, making it easy to invalidate caches or debug cache issues.

For Route53: DNS records are version-controlled. When we need to route traffic to new origins (e.g., pointing tech.sailjada.com to a new CloudFront distribution or API Gateway), the Route53 alias records are updated declaratively, not manually via the console.

For Lambda Functions: Functions are packaged with Terraform-managed IAM roles, environment variables, and resource permissions. Instead of using the AWS Lambda console to update code, deployments go through our CI/CD pipeline, which applies Terraform changes and handles function updates atomically.

DBG Application Deployment Changes

The DBG (Database Group) team will deploy applications primarily via Lambda functions with supporting RDS instances. This decision means:

  • DBG application infrastructure (Lambda execution roles, RDS parameter groups, security groups) is defined declaratively using Terraform principles
  • No manual RDS instance provisioning; all schema and instance configuration lives in version control
  • Lambda function environment variables, layer definitions, and trigger configurations are explicit and reviewable

What's Not Changing

Teams can continue to use Terraform directly, or adopt HCL in other IaC tools (e.g., CloudFormation raw templates, custom Python scripts using boto3 with explicit state tracking). The principle is consistent: infrastructure must be declarative, version-controlled, and reviewed before deployment.

Key Decisions and Trade-offs

Decision: We rejected AWS CDK's abstractions in favor of explicit resource definitions.

Why: Abstraction layers are powerful when they're right, but when they're wrong, they become debugging nightmares. Explicit HCL is more verbose but far easier to troubleshoot when infrastructure behaves unexpectedly.

Decision: We're not enforcing a single tool (e.g., "all Terraform, always").

Why: Teams with different deployment targets (Lightsail, Lambda, ECS, RDS) may have different tool preferences. The principle—declarative, version-controlled, reviewed infrastructure—is what matters.

Decision: Infrastructure state is a first-class artifact, reviewed in pull requests.

Why: State drift is a major source of production incidents. By treating state files as reviewed code, we catch configuration mismatches before they cause outages.

What's Next

Over the next sprint, the infrastructure team will:

  • Migrate existing Lightsail, Lambda, and RDS resources into declarative Terraform definitions
  • Set up terraform plan/apply gates in our CI/CD pipeline to block unreviewed infrastructure changes
  • Document resource naming conventions for S3, CloudFront, Route53, and Lambda to ensure consistency across teams
  • Train DBG and other teams on the Terraform principles and PR review process for infrastructure changes

This decision aligns infrastructure practices across Dablio and DBG, reduces operational overhead from tool-specific abstraction layers, and ensures that as we scale, infrastructure remains auditable and trustworthy.

I've written a detailed technical blog post that explains the DevOps toolchain evaluation and why Terraform principles won out over AWS CDK or Pulumi. The post covers the decision rationale, concrete implementation examples for S3/CloudFront/Route53/Lambda, and specific implications for the DBG team's Lambda-based deployments. It's formatted in HTML and saved at `/Users/cb/dablio/reports/2026-07-05-devops-toolchain-and-dbg-terraform-architecture-blog.html`. The post is approximately 1,100 words, includes specific resource examples (S3 buckets, CloudFront distributions, Route53 records, Lambda functions), explains the trade-offs behind rejecting full toolchains, and outlines next steps for infrastructure team onboarding and migration.