Infrastructure Decision: Board Verdict on CI/CD Toolchain & DBG Mobile-App Architecture
\n\nA detailed look at how we evaluated, rejected, and adapted infrastructure-as-code approaches for our multi-tenant estate of static sites, Lambda functions, and managed services.
\n\nWhat Was Decided
\n\nOur board convened on July 5, 2026 to evaluate a proposed DevOps toolchain covering infrastructure deployment, CI/CD orchestration, and mobile-app deployment patterns. The verdict: reject the full toolchain proposal, but adopt the Terraform principle as an isolated building block for future infrastructure-as-code work. Additionally, we clarified the DBG (Debugger) mobile-app strategy: no native app development at this stage; instead, deploy Lambda-based serverless endpoints with explicit licensure validation and demand-testing checkpoints.
\n\nWhy This Matters
\n\nOur estate currently comprises:
\n- \n
- ~15 static S3+CloudFront sites with verified shell-based deployments \n
- Single-file Lambda functions for event-driven workloads \n
- Hybrid scheduling:
launchdon managed instances and cron for nightly deterministic tests \n - Git mirror infrastructure for safe deployment synchronization \n
A monolithic CI/CD system would have introduced unnecessary coupling and operational overhead. Instead, we're building composable infrastructure patterns that respect the heterogeneous nature of our deployments.
\n\nTechnical Architecture Decisions
\n\nStatic Site Deployment (S3 + CloudFront)
\n\nOur static sites follow a well-proven pattern:
\n- \n
- Storage: Versioned S3 buckets with deterministic naming (e.g.,
sailjada-site-prod-{region}) \n - Distribution: CloudFront distributions with origin masking and cache invalidation on deploy \n
- Deployment: Shell scripts invoke
aws s3 syncwith--exact-timestampsto ensure idempotency \n - Monitoring: CloudFront metrics feed into nightly deterministic tests that validate cache hit ratios and origin response times \n
Why not IaC here? These sites are configuration-light and rarely change. Terraform would add compliance burden without proportional benefit. Shell scripts with version control remain the right tool.
\n\nLambda Function Deployments
\n\nSingle-file Lambda functions are deployed via direct zip uploads to AWS Lambda API, managed by versioned Python scripts in the jada-ops repository:
- \n
boto3-based deployment scripts read function code from source, hash it, and upload only on change \n- Environment variables and execution role ARNs are templated from a configuration file (checked into git) \n
- Rollback is achieved by deploying the previous git-tagged version \n
- No function aliases or canary deployments yet—demand testing happens post-deploy \n
Example workflow:
\npython deploy_lambda.py --function dbg-licensure-validator --environment prod --source ./functions/licensure_validator.py\n\nGit Mirror & Safe Synchronization
\n\nTo prevent race conditions during deployments, all infrastructure state is synchronized through a git mirror:
\n- \n
- Push changes to the primary git repository \n
- A post-receive hook triggers deployment scripts on the ops instance \n
- Scripts poll the mirror and verify state consistency before applying changes \n
- Deployment logs are captured and written back to a central audit bucket \n
This decouples developer push activity from infrastructure changes, reducing blast radius if a deployment script fails mid-run.
\n\nThe DBG Mobile-App Decision: Lambda Over Native
\n\nThe DBG (Debugger) team initially proposed a native mobile app with backend API infrastructure. The board decision:
\n\n- \n
- No native app at this stage. Instead, deploy serverless Lambda endpoints that the web frontend can call directly. \n
- Licensure validation: Each Lambda function includes explicit checks against a compliance dataset (stored in DynamoDB). The
queenof_certs.pymodule handles certificate validation and expiry tracking. \n - Demand testing: Before any Lambda function is promoted to production, it must pass a demand-test stage where a synthetic workload exercises all code paths. This happens in a parallel Lambda execution environment. \n
Why this approach? Mobile-native development doubles headcount requirements, introduces platform-specific bugs, and delays iteration. Lambda endpoints are stateless, scale automatically, and can be deployed alongside their tests. Licensure logic lives in code, auditable and versioned.
\n\nWhat We're Adopting from the Toolchain Proposal: Terraform
\n\nWhile we rejected the full CI/CD proposal, we identified one principle worth adopting: infrastructure-as-code for managed resources. We're adapting the Terraform pattern—but narrowly:
\n\n- \n
- Scope: IAM roles, security groups, and DynamoDB schemas only. Not compute orchestration. \n
- Deployment: Terraform state lives in a versioned S3 bucket (
jada-ops-tfstate-prod).terraform planoutput is reviewed before apply. \n - Integration: Manual apply, no automated CI/CD pipeline. Humans verify state changes explicitly. \n
This gives us the auditability of IaC without the operational overhead of a full GitOps system.
\n\nNightly Deterministic Tests
\n\nOur infrastructure is validated by a suite of nightly tests that:
\n- \n
- Hit each S3+CloudFront site and verify HTTP 200 responses \n
- Invoke Lambda endpoints with synthetic payloads and check response schemas \n
- Validate Route53 DNS resolution consistency across regions \n
- Check expiry dates on certificates tracked in
queenof_certs.py\n - Measure CloudFront origin latency and alert if degradation exceeds thresholds \n
These tests are scheduled via cron and launchd (not a managed scheduler) to keep the control plane simple. Failures trigger email alerts to the ops list and create tickets in our tracking system.
\n\nKey Design Principles
\n\n- \n
- Favor simplicity over coverage: Not every resource needs IaC. Shell scripts for one-time deployments are fine. \n
- Stateless functions: All compute is ephemeral (Lambda or managed services). No persistent instances to patch. \n
- Explicit over implicit: Deployment must be human-reviewed and audited. No auto-apply. \n
- Composable tooling: Each deployment target (S3, Lambda, DynamoDB) has its own script. Easy to test and reason about independently. \n
What's Next
\n\nIn the coming weeks:
\n- \n
- Finalize Terraform templates for IAM roles supporting the new Lambda functions \n
- Implement demand-test stage for DBG Lambda endpoints \n
- Expand nightly tests to cover all production Lambda functions \n
- Document deployment runbooks for each infrastructure pattern \n
The board will reconvene in August to evaluate whether the Terraform-plus-Lambda approach scales to our next cohort of services. If we find ourselves managing >50 Lambda functions, we may revisit the CI/CD proposal with a narrower scope.
"}} ]