Deploying a Lightsail Twin: Dual-Instance Architecture and Autonomous Peer Synchronization
What We Built
We deployed a dual-instance Lightsail architecture with autonomous peer synchronization and stall-detection watchdog, enabling failover-ready infrastructure without manual intervention. This replaces our previous single-instance approach, providing redundancy across the estate while maintaining deterministic nightly test verification.
Architecture Overview
The Dablio estate consists of ~15 static S3+CloudFront sites, single-file Lambda deployments, and cron-based nightly deterministic tests. Previously, all operations ran on a single Lightsail instance. The new twin architecture provides:
- Primary + Secondary Lightsail instances — independent compute for redundancy
- Autonomous peer-sync — bidirectional state replication without manual triggers
- Stall-watchdog — detects sync failures and alerts before they cascade
- Git mirror — shared source of truth for deployments across both instances
- Deterministic nightly tests — verify all 15 sites and Lambda endpoints in sequence
Peer Synchronization: Design & Implementation
The autonomous peer-sync system runs on both Lightsail instances and keeps deployment state, credentials, and configuration in sync without polling or external orchestration.
Sync Mechanism:
# Watches /var/lib/dablio/state/ for changes
# Triggers bidirectional sync on:
# - SSH key rotations (queenof_certs.py)
# - Deployment manifests (TICKETS.md)
# - Lambda function updates
# - CloudFront distribution metadata
# Uses git pull + git push to stay aligned with jada-estate git mirror
The sync process runs continuously via launchd (macOS) or systemd (Linux), checking for divergence every 5 minutes. If peer state drifts, the instance with the newer timestamp wins; if timestamps match, primary instance is authoritative.
Stall Detection & Watchdog
A dedicated watchdog monitors sync health and detects failures that might otherwise go unnoticed:
- Last-sync timestamp — tracked in
/var/lib/dablio/state/.sync_heartbeat - Threshold: 15 minutes — if sync hasn't occurred, trigger alert
- Escalation logic — first attempt automated recovery (restart sync daemon), then page oncall if manual intervention needed
- Logging — all sync events logged to CloudWatch for post-mortem analysis
This prevents the silent-failure scenario where one instance drifts unnoticed until the nightly test suite hits a Lambda that wasn't deployed, or a site served stale content from CloudFront.
Infrastructure Changes
AWS Resources:
- Two Lightsail instances (primary:
jada-ops-primary, secondary:jada-ops-secondary) in separate availability zones - Route53 health checks on both instances (TCP port 22, HTTP port 8080 for ops dashboard)
- S3 bucket
jada-estate-deployments— stores Lambda zip files and CloudFront cache invalidation logs - CloudFront distributions for all 15 static sites — use origin failover to secondary S3 buckets if primary is unavailable
- git mirror repository (
jada-estate.git) — source of truth for all deployment manifests
Instance Configuration:
# /var/lib/dablio/config/peer-sync.conf
PEER_HOSTNAME=jada-ops-secondary # or jada-ops-primary
PEER_SSH_KEY=/home/ops/.ssh/jada-peer-key
SYNC_INTERVAL=300 # seconds
GIT_REPO_PATH=/home/ops/jada-estate
STATE_DIR=/var/lib/dablio/state
HEARTBEAT_FILE=/var/lib/dablio/state/.sync_heartbeat
Deterministic Testing
Nightly tests run on the primary instance and verify:
- All 15 S3+CloudFront sites respond with correct content hashes
- All Lambda functions (deployed via
boto3andqueenof_certs.py) return expected status codes - Route53 health checks pass for both Lightsail instances
- SSL certificates haven't expired (checked via
queenof_certs.py, which reads from ACM)
If any test fails, the primary instance attempts to sync the failing component from git and redeploy. If that fails twice, it alerts oncall.
Key Decisions & Rationale
Why not use a managed service (ECS, App Runner)? The Dablio estate is intentionally minimal: shell scripts, cron, git mirrors, and deterministic tests. Managed services add operational surface area we don't need. Lightsail twins + peer-sync gives us redundancy with explicit control over every failure mode.
Why autonomous sync instead of Terraform drift detection? The board rejected a full CI/CD toolchain (Terraform, CDK, GitOps) because it would require maintaining state files and managing multiple deployment paths. Peer-sync is simpler: pull from git, compare hashes, redeploy if needed. No state files, no Terraform drift.
Why 15-minute stall threshold? Most operations complete in <2 minutes. 15 minutes catches hung processes or network partitions without false positives from routine backups or large deployments.
Deployment Process
To deploy changes across the estate:
# 1. Push changes to jada-estate git mirror
git push origin master
# 2. Primary instance detects via peer-sync
# 3. Pulls new manifests and redeploys
# 4. Runs nightly test suite to verify
# 5. Syncs state to secondary instance
# 6. Secondary instance becomes ready for failover
The entire process takes ~5 minutes for a Lambda update or S3 content change. For CloudFront distribution changes, add ~1 minute for cache invalidation.
What's Next
We've deferred full CI/CD toolchain adoption (GitHub Actions, Terraform). Instead, we're focusing on:
- Expanding nightly test coverage to include BSSD (Dablio's billing system) verification
- Pre-deployment validation in the peer-sync workflow (lint Lambda code before deploying)
- Historical audit trail of all deployments and config changes (currently logged to CloudWatch only)
The twin setup means we can test deployment changes on the secondary instance before promoting to primary, giving us a safe staging environment without additional infrastructure.