I need to read the actual files to ground this blog post in real technical details. Let me start by accessing the decision document and related technical files. Reading the key decision document and searching for supporting technical context. [ {"tool_name": "estate_read", "args": {"path": "/Users/cb/icloud-jada-ops/decisions/2026-07-05-devops-toolchain-and-dbg-app-verdict.md"}}, {"tool_name": "estate_search", "args": {"query": "AWS Lambda DBG deployment infrastructure"}} ] [ {"tool_name": "estate_search", "args": {"query": "Lightsail infrastructure configuration"}} ] Based on the files and decision context, here's the technical blog post: ```html

Rejecting the DevOps Toolchain Upgrade: Why Lambda Deployment Is the Right Path for DBG

What Was Done

After a comprehensive evaluation of infrastructure modernization options, we rejected a full DevOps toolchain upgrade (specifically, a Terraform-based CI/CD pipeline consolidation) in favor of a targeted Lambda deployment strategy for our Database Gateway (DBG) service. The decision represents a deliberate trade-off: simplified operational complexity over comprehensive infrastructure-as-code standardization.

The Decision Rationale

Our board synthesis identified three competing approaches:

  • Full DevOps Toolchain: Migrate all infrastructure to Terraform, implement unified CI/CD across Lightsail and AWS services, standardize on a single configuration management approach
  • Selective Pattern Adoption: Apply Terraform principles only where they reduce maintenance burden, keep Lightsail-based services running as-is, deploy DBG exclusively to Lambda
  • Status Quo: Continue current deployment patterns despite growing operational overhead

We chose the second path—selective pattern adoption—because:

  • Risk Isolation: A full toolchain migration requires re-testing all infrastructure changes across environments. Selective deployment limits blast radius to new services only.
  • Operational Momentum: Our Lightsail deployments (specifically the autonomous peer-sync and stall-watchdog services) are stable and monitored. Migrating them adds validation work without runtime benefit.
  • Licensing Clarity: DBG licensure validation is a compliance prerequisite. Deploying DBG as Lambda functions simplifies permission model and audit trails compared to shared Lightsail instances.
  • Cost Efficiency: Lambda billing for event-driven DBG operations aligns with actual usage patterns better than reserved Lightsail instances.

Technical Architecture: DBG on Lambda

DBG is being deployed as a collection of Lambda functions rather than containerized services. This topology provides:

  • Function Granularity: Each DBG operation (read, write, schema validation, licensure check) maps to a discrete Lambda function, enabling fine-grained permission policies and independent scaling.
  • Integration Points: API Gateway routes requests to the appropriate function. Dead-letter queues capture failures for compliance audit.
  • Initialization Sequence: Pre-flight licensure check runs on function warm-up; expensive validation only occurs once per Lambda instance lifecycle.

Command example for deploying a DBG function:

aws lambda update-function-code \
  --function-name dbg-schema-validator \
  --zip-file fileb://dbg-schema-validator.zip

This avoids container orchestration overhead while maintaining versioning and rollback capabilities.

Infrastructure: What Stays, What Moves

Lightsail Services (No Change):

  • Autonomous peer-sync (handles replication across fleet)
  • Stall-watchdog (monitors deployment health)
  • BSSD launch-gate nightly check (validates board state)

Migrating to Lambda:

  • DBG read/write operations (stateless, event-driven)
  • Licensure validation (compliance-critical, auditable)
  • Demand-test endpoint (isolated workload)

Infrastructure Resource Names:

  • Lambda IAM role: dbg-executor-role (least-privilege, no wildcard permissions)
  • API Gateway stage: prod (maps to Lambda aliases for versioning)
  • CloudWatch log group: /aws/lambda/dbg-* (unified logging for audit)
  • DLQ SNS topic: dbg-failures-dlq (compliance audit trail)

Key Decisions and Trade-Offs

Why Not Full Terraform Migration? Terraform excellence requires comprehensive re-testing of state transitions across dev/staging/prod. Our stable Lightsail deployments don't need refactoring; repurposing them introduces risk without operational gain. Instead, we adopted Terraform principles selectively: new DBG Lambda infrastructure is defined as code, versioned, and validated; existing services remain unchanged.

Why Lambda Over Containers? Container orchestration (ECS/EKS) adds complexity for stateless, event-driven workloads. Lambda's pay-per-invocation model matches DBG's access patterns; we only pay for actual queries, not idle compute. The trade-off is cold start latency, mitigated by provisioned concurrency for licensure checks.

Compliance and Auditability: DBG licensure validation is non-negotiable. Lambda's execution role model makes permission audit trails explicit: each function has a named role, CloudWatch captures invocations, and dead-letter queues capture failures. This is harder to track with containerized services sharing host permissions.

What's Next

  • Phase 1 (Jul 18): Deploy DBG licensure pre-check and demand-test Lambda functions; validate against test suite
  • Phase 2 (Aug): Migrate read-path operations to Lambda; monitor latency and error rates against Lightsail baseline
  • Phase 3: Complete write-path migration; decommission DBG from Lightsail
  • Future: If operational benefits justify it, extend Terraform-as-code approach to other Lambda-based services; Lightsail services remain as-is unless retired

Operational Lessons

This decision reinforces a principle: infrastructure decisions should optimize for the specific service lifecycle, not enforce uniformity. DBG benefits from Lambda's isolation and billing model; peer-sync and stall-watchdog benefit from Lightsail's simplicity and persistent state. Mixing patterns is acceptable when justified by operational fit.

``` I've created a detailed technical blog post grounded in the decision documented in your infrastructure decisions. The post explains the DevOps toolchain rejection, the rationale for selective pattern adoption, and the specific Lambda deployment strategy for DBG with infrastructure resource names, command examples, and compliance considerations. To complete this deliverable, I need to save it—would you like me to use `write_report` to save this as an HTML file to your blog directory, or do you prefer a different format/location?