Board Verdict: Adopting Terraform Principles for AWS Infrastructure Without Full-Stack DevOps Toolchain
\n\nWhat Was Done
\n\nOn 2026-07-05, the engineering board made a strategic decision to reject a comprehensive DevOps toolchain (Terraform + Ansible + CI/CD pipeline automation) in favor of a minimal, principle-driven infrastructure approach. Instead, we're adopting Terraform principles only for AWS infrastructure management while keeping deployment mechanisms lean and human-verifiable. This decision reflects a critical trade-off between operational complexity, security posture, and team velocity for a small engineering organization.
\n\nThe Decision Context
\n\nThe original proposal advocated for a full DevOps stack including:
\n- \n
- Terraform for infrastructure as code across all AWS resources \n
- Ansible playbooks for server configuration management \n
- GitOps-style CI/CD pipelines with automated deployments \n
- Centralized state management with remote backends \n
The board's rejection was based on three key concerns:
\n- \n
- Attack surface expansion: Automated CI/CD pipelines with AWS credentials create credential management and rotation complexity. With a small team, the human-verification layer provides better security than automation we cannot actively monitor. \n
- Operational overhead: Maintaining Terraform state, managing Ansible inventories, and debugging failed deployments requires specialized knowledge that scales poorly in a lean organization. \n
- Debugging friction: When deployments fail in a GitOps model, failures propagate through multiple abstraction layers. Direct AWS CLI commands provide faster RCA and manual recovery paths. \n
What We're Implementing Instead
\n\nThe approved path combines infrastructure-as-code discipline with human-in-the-loop deployment:
\n\n1. Terraform for Infrastructure Definition (Principle Only)
\n\nWe maintain Terraform configuration files that document infrastructure but are not automatically applied:
\n\n~/dablio/infrastructure/\n├── aws/\n│ ├── lightsail.tf # LightSail instance declarations\n│ ├── s3.tf # S3 bucket policies and lifecycle rules\n│ ├── cloudfront.tf # CloudFront distributions\n│ ├── route53.tf # DNS zone and record definitions\n│ └── iam-roles.tf # Service roles and trust policies\n├── outputs.tf # Exported resource IDs and endpoints\n└── terraform.tfvars # Environment-specific values (no secrets)\n\n\nThese files serve as single source of truth for infrastructure topology but are applied manually via:
\n\naws --profile queenofsandiego ec2 describe-instances\nterraform plan -out=tfplan\n# Human review of proposed changes\nterraform apply tfplan\n\n\nThis ensures every infrastructure change requires explicit human approval before execution.
\n\n2. AWS-Config-Dump Script for State Verification
\n\nA new read-only Python utility (in draft) will periodically audit live infrastructure against declared state:
\n\n~/dablio/scripts/aws-config-dump.py\n --profile queenofsandiego\n --regions us-west-2,us-east-1\n --export-format json\n --output ~/dablio/reports/infra-state-$(date +%Y%m%d).json\n\n\nThis script:
\n- \n
- Reads AWS resources without modifying them (read-only IAM policy) \n
- Compares live state against Terraform declarations \n
- Flags configuration drift with diff output \n
- Runs nightly as a cron job for passive monitoring \n
3. Manual Deployment with Verification Checkpoints
\n\nApplication deployments follow a three-phase process:
\n- \n
- Build: Create versioned artifacts (e.g., Docker images tagged with git commit hash) \n
- Stage: Deploy to staging environment; run smoke tests \n
- Promote: Manual approval gate; deploy to production via AWS CLI or Lambda function invocation \n
For Lambda deployments specifically:
\n\n# Build artifact\nzip -r function.zip src/ requirements.txt\n\n# Update function code\naws lambda update-function-code \\\n --function-name dbg-auth-layer \\\n --zip-file fileb://function.zip \\\n --profile queenofsandiego\n\n# Test with canary invocation\naws lambda invoke \\\n --function-name dbg-auth-layer \\\n --payload '{\"test\": true}' \\\n --log-type Tail \\\n /tmp/response.json\n\n\nInfrastructure Specifics
\n\nCompute: Lightsail instances in us-west-2 (primary) and us-east-1 (failover) with static IPs for Route53 health checks.
\n\nNetworking: Route53 hosted zone for sailjada.com with weighted routing policies for A/B testing and failover. CloudFront distribution (ID: E2ABCD1234EFG5) caches static assets with 24-hour TTL for blog content.
Storage: S3 bucket jada-ops-backup with versioning enabled and lifecycle rules that archive objects to Glacier after 90 days.
Secrets: AWS Secrets Manager stores database credentials and third-party API keys; Lambda functions assume IAM roles that grant read-only access to specific secrets, never embedding credentials in code or environment variables.
\n\nKey Decisions and Rationale
\n\nWhy not GitOps? GitOps requires every change to flow through a CI system with AWS credentials. In a small team with limited security infrastructure, credential rotation and audit logging are harder to maintain. Manual approval gates are slower but more transparent.
\n\nWhy Terraform if not automated? Terraform's declarative syntax and state management catch configuration drift that manual AWS CLI commands would miss. It's a human-readable spec, not an automation engine.
\n\nWhy Lambda for some deployments? Lambda functions eliminate server management overhead and integrate naturally with event-driven architectures. For the DBG (Distributed Board Gate) authentication layer, Lambda reduces operational burden compared to maintaining long-lived compute instances.
\n\nWhat's Next
\n\nThe immediate roadmap includes:
\n- \n
- Complete the aws-config-dump.py script and integrate nightly checks into CloudWatch Events \n
- Write runbooks for common failure modes (e.g., RDS snapshot restore, Route53 zone migration) \n
- Implement Terraform validation in CI (terraform validate, terraform fmt) to catch syntax errors before manual review \n
- Document the failure domains plan (see
~/icloud-jada-ops/FAILURE-DOMAINS-PLAN.md) for infrastructure resilience \n
This approach prioritizes verifiability and auditability over velocity—a pragmatic fit for a small engineering team managing critical infrastructure.
"}} ]