Board Verdict on DevOps Toolchain: Adopting Terraform Principle Instead of Full Platform
What Was Done
On 2026-07-05, the board rejected a comprehensive DevOps toolchain proposal but approved adoption of the Terraform Principle—a lightweight, infrastructure-as-code approach focused on version-controlled resource definitions, declarative configuration, and repeatable deployments. This decision fundamentally shaped our infrastructure strategy while maintaining operational simplicity and avoiding vendor lock-in.
The verdict eliminated complexity from a proposed all-in-one CI/CD platform and instead created a focused technical roadmap: implementing read-only AWS config dumps, establishing Lambda-based deployment patterns for DBG, and operationalizing our Lightsail dual-instance infrastructure with autonomous peer-sync and stall-watchdog mechanisms.
Why the Toolchain Was Rejected
The comprehensive DevOps toolchain introduced several risks that misaligned with our scale and team size:
- Complexity Overhead: Full platform toolchains require operational burden (upgrades, security patches, monitoring) that scales with feature count, not deployment frequency.
- Vendor Coupling: Platform-specific DSLs and abstractions create switching costs if operational priorities shift.
- Team Fit: We have experienced engineers who prefer direct AWS API and Terraform control over abstraction layers that hide the underlying infrastructure.
- Cost: Platform licensing and management overhead for occasional deployments is economically inefficient.
Decision Logic: By adopting Terraform as our principle rather than adding a platform, we preserve the flexibility to evolve our deployment patterns without re-architecture. Infrastructure state lives in version control, and changes flow through Git rather than proprietary UIs.
The Terraform Principle in Practice
The Terraform Principle means all cloud resources are:
- Declared in code: Resource definitions in HCL are committed to git, creating an audit trail and enabling code review of infrastructure changes.
- State-managed explicitly:
terraform.tfstatefiles track deployed resource state. We use remote state with versioning to prevent drift. - Modular and composable: Infrastructure modules map to business domains—separate modules for networking, compute (Lightsail instances), Lambda functions, and storage (S3 buckets).
- Tested before deployment:
terraform planoutputs are reviewed as part of the merge request process. Drift detection runs nightly.
Infrastructure Implementation
Lightsail Twin Deployment
Our primary compute runs on two Lightsail instances in different availability zones, deployed via Terraform under the resource names:
aws_lightsail_instance.primary
aws_lightsail_instance.secondary
These instances host the core Dablio workload. Terraform manages instance creation, security group rules, static IP allocation, and disk snapshots. DNS records point to a Route53 alias that routes traffic using application-level health checks.
Autonomous Peer-Sync
The Lightsail instances run a peer-sync service that automatically replicates configuration and state between primary and secondary. This service:
- Periodically compares local state with the peer instance via authenticated SSH.
- Syncs critical files (application configuration, certificates, database schemas) using rsync with exclusion patterns for logs and temporary data.
- Logs sync operations to CloudWatch for observability.
- Runs idempotently so repeated syncs don't corrupt state.
Stall-Watchdog Mechanism
A stall-watchdog Lambda function (deployed via Terraform as aws_lambda_function.stall_watchdog) runs every 5 minutes to detect hung peer-sync processes:
- Checks the timestamp of the last successful sync on both instances.
- If either instance has not synced in 15+ minutes, it sends a CloudWatch alarm to the ops channel.
- For critical stalls (30+ minutes), it triggers an automated incident response: stop the stalled sync process, reset the state file, and retry.
- All actions are logged with context (instance ID, sync duration, error message) for later analysis.
AWS Config Dump Script
Per the board verdict, we're implementing a nightly read-only AWS config dump as a Terraform-principle resource audit. The script:
- Runs as a Lambda function on a scheduled CloudWatch Events rule (cron:
0 2 * * *UTC). - Calls AWS Config API to enumerate all deployed resources, their properties, and relationships.
- Exports the dump to an S3 bucket (
s3://config-state-dumps/) with a timestamp-based key:dumps/YYYY-MM-DD-HH-mm-ss-config.json. - Compares the new dump against the previous day's dump to detect untracked changes (drift from Terraform state).
- Publishes a drift report to
s3://config-state-dumps/reports/with checksums of modified resources.
Why read-only? The dump is an audit mechanism, not a remediation tool. It reveals when humans or automation have changed infrastructure outside Terraform, enabling us to either update Terraform to match or revert the change. This prevents silent drift.
DBG Deployment Strategy
DBG (a separate domain service) uses a Lambda-based deployment pattern:
- No persistent app containers: DBG logic runs as ephemeral Lambda functions rather than long-running processes, reducing operational surface area.
- Deployment target: Lambda functions are deployed via Terraform with CloudFormation or SAM templates committed to git.
- Triggering: API Gateway endpoints or EventBridge rules invoke Lambda functions for licensure checks, demand testing, and domain-specific operations.
- State storage: Persistent state (licensure records, test results) lives in DynamoDB tables, also defined in Terraform.
This pattern avoids maintaining application infrastructure while preserving the ability to run scheduled or event-driven workloads. Licensure pre-checks run as Lambda functions with access to DBG's certification database.
Key Decisions and Trade-offs
- Terraform state backend: State is stored in an S3 bucket with versioning and MFA delete enabled. DynamoDB provides locking to prevent concurrent modifications. This requires careful credential management (AWS credentials stored in 1Password, rotated quarterly).
- No full CI/CD pipeline: We deploy manually via
terraform applyafter code review, not through a push-button CI system. This maintains control and reduces deployment velocity risk for our small team. - Monitoring and alerting: CloudWatch metrics and logs are the primary observability layer. Custom dashboards track Lightsail CPU/memory, Lambda error rates, and config dump drift.
- Cost optimization: Lightsail instances are right-sized for our load profile. Lambda functions scale to zero. S3 buckets use lifecycle policies to transition old dumps to Glacier after 90 days.
What's Next
- Drift detection automation: Expand the config dump script to auto-remediate certain drift patterns (e.g., security group rule changes) with approval gates.
- Multi-region failover: Plan secondary Lightsail instances in a different AWS region with cross-region RDS read replicas.
- Compliance auditing: Integrate AWS Config rules to enforce security policies (encryption, tagging, public access restrictions) declaratively in Terraform.
- Team onboarding: Document Terraform workflows and create runbooks for common operations (scaling, incident response, config updates).
Conclusion
By adopting the Terraform Principle and rejecting a heavyweight platform, we've created a lean, version-controlled infrastructure that scales with our team's expertise and operational needs. The combination of Lightsail dual instances, automated peer-sync, Lambda-based DBG deployments, and read-only config audits gives us both reliability and visibility without the complexity tax of a full DevOps platform.