Deploying Redundant Lightsail Infrastructure with Autonomous Peer-Sync Architecture
\n\nWhat Was Done
\n\nDeployed a twin Lightsail instance configuration with autonomous peer-to-peer synchronization for the Dablio platform, enabling stateless failover and load distribution across geographically-separated compute nodes. The deployment introduced a stall-watchdog monitoring system to detect and recover from peer desynchronization events, coupled with a Terraform-driven infrastructure-as-code principle that standardizes resource provisioning across environments.
\n\nTechnical Architecture
\n\nDual Lightsail Deployment Model:
\n- \n
- Primary and secondary Lightsail instances deployed as peer nodes rather than master-replica topology \n
- Each instance runs identical application code but maintains separate state directories for peer identity and sync metadata \n
- State is coordinated through a shared S3 backend bucket configured with versioning and server-side encryption \n
- Route53 health checks validate application responsiveness on both nodes independently \n
Peer-Sync Protocol:
\n- \n
- Autonomous synchronization leverages file system watching on designated sync directories (typically
/var/lib/dablio/sync/) \n - Peer identity established through instance metadata tags:
peer-id,peer-role,sync-generation\n - On boot, each instance queries S3 for the peer manifest at
s3://dablio-peer-state/peer-manifest.jsonto discover active peer nodes \n - Changes to application state trigger atomic writes to S3 with conditional headers (
If-Match/If-None-Match) to prevent lost updates \n - Peer nodes poll the S3 state bucket every 5 seconds (configurable via
PEER_SYNC_INTERVALenvironment variable) and merge remote changes into local state \n
Stall-Watchdog Subsystem
\n\nProblem Being Solved:
\nIn peer-to-peer architectures, silent failures where a node stops syncing but remains running are difficult to detect. The stall-watchdog monitors bidirectional sync health:
\n\n- \n
- Each peer writes a heartbeat object to S3 at
s3://dablio-peer-state/heartbeats/{peer-id}every 10 seconds, including a monotonic counter and latest sync timestamp \n - The local watchdog daemon reads peer heartbeats and triggers remediation if any peer hasn't updated for 30+ seconds \n
- Remediation steps:\n
- \n
- First attempt: restart the peer's sync worker process via SSH (requires VPC security group rule permitting inter-peer SSH) \n
- Second attempt (if peer remains stalled): terminate the instance and trigger ASG replacement \n
\n - Watchdog logs all decisions to CloudWatch Logs group
/aws/lightsail/dablio-peer-watchdogfor audit and debugging \n
Implementation Details:
\n- \n
- Watchdog runs as a systemd service:
/etc/systemd/system/dablio-peer-watchdog.service\n - Configuration file at
/etc/dablio/watchdog-config.yamldefines heartbeat thresholds and remediation strategies \n - Peer instance terminates with exit code logged to CloudWatch to distinguish planned vs. unplanned restarts \n
Infrastructure Setup with Terraform
\n\nTerraform Structure:
\n- \n
- Configuration root:
infra/terraform/with modules for Lightsail instances, S3 state backend, and Route53 health checks \n - Lightsail module provisions two instances:
dablio-peer-primaryanddablio-peer-secondary\n - Blueprint used: Ubuntu 24.04 LTS (standard Lightsail blueprint)\n
- \n
- Instance type:
medium_3_0(3 vCPU, 1 GB RAM, 60 GB SSD) \n - Each instance includes user-data script (
templates/peer-init.sh) that:\n- \n
- Installs runtime dependencies: Node.js 20.x, AWS CLI v2, jq \n
- Configures IAM instance profile with S3 read/write permissions to
dablio-peer-statebucket \n - Deploys application code from S3:
s3://dablio-releases/app-{version}.tar.gz\n - Starts peer-sync daemon and stall-watchdog service \n
\n
\n - Instance type:
S3 State Backend:
\n- \n
- Bucket:
dablio-peer-state(environment-tagged:Environment=production) \n - Object key structure:\n
\npeer-manifest.json # List of active peers\nheartbeats/{peer-id} # Per-peer heartbeat objects\nsync/{component}/{version} # Application state snapshots\narchive/{date-partition}/ # Archived states for compliance retention\n - Versioning enabled for audit trail; lifecycle policy archives to Glacier after 90 days \n
- Server-side encryption with KMS key
arn:aws:kms:us-east-1:XXXX:key/dablio-peer-state\n
Route53 Health Checks:
\n- \n
- Two health checks created: one per Lightsail instance, each probing
http://{instance-ip}:8080/healthz\n - Health check interval: 30 seconds, failure threshold: 2 consecutive failures \n
- A-record for
peer.dablio.localconfigured with health-check evaluation; automatically points to healthy instance (failover within ~1 minute) \n
Deployment Commands and Validation
\n\nInitial Deployment:
\ncd infra/terraform/\nterraform init -backend-config=\"bucket=dablio-terraform-state\" -backend-config=\"key=peer/prod.tfstate\"\nterraform plan -var=\"lightsail_region=us-east-1\" -var=\"peer_count=2\"\nterraform apply -auto-approve\n\n\nPost-Deployment Validation:
\n- \n
- SSH to each instance and verify peer-sync daemon status:
systemctl status dablio-peer-sync\n - Check S3 state:
aws s3 ls s3://dablio-peer-state --recursive\n - Verify heartbeats:
aws s3 ls s3://dablio-peer-state/heartbeats/\n - Monitor watchdog logs:
aws logs tail /aws/lightsail/dablio-peer-watchdog --follow\n - Trigger a manual sync test by modifying a watched file on the primary instance and confirming propagation to secondary within 5 seconds \n
Key Architectural Decisions
\n\nPeer-to-Peer vs. Master-Replica: We chose peer-to-peer topology to eliminate single points of failure during promotion events. Master-replica requires failover logic and can cause brief unavailability; peer-to-peer allows traffic to route around either node independently.
\n\nS3 as Sync Coordinator: Instead of a dedicated etcd or Consul cluster, S3 serves as the authoritative state store. This reduces operational overhead (no new stateful services) and leverages IAM/encryption we already had in place. Trade-off: slightly higher latency (~100ms round-trip) compared to local cluster coordinators.
\n\nTerraform as IaC Principle: Standardized all infrastructure provisioning through Terraform to enable reproducible environments and audit trails. State file stored in versioned S3 backend with DynamoDB lock table to prevent concurrent modifications.
\n\nMonitoring and Observability
\n\n- \n
- CloudWatch custom metrics published by peer-sync daemon:\n
- \n
DablioPeerSync:OutOfSyncObjects— count of objects diverged between peers \n DablioPeerSync:SyncLatencyMs— time between change detection and S3 write confirmation \n DablioPeerWatchdog:HeartbeatsMissed— total remediation events triggered \n
\n - CloudWatch Alarms configured to notify ops channel when stall-watchdog triggers more than 5 remediation events in 1 hour \n
- Application logs streamed to
/aws/lightsail/dablio-applog group for centralized troubleshooting \n
What's Next
\n\nFuture work includes: (1) introducing eventual consistency verification tests to catch divergence bugs before they reach production, (2) adding capacity planning for sync traffic as state volumes grow, and (3) evaluating DynamoDB as a potential replacement for S3 coordination to reduce latency for high-frequency sync workloads.
\n"}} ]