Lightsail Twin Deployment & Autonomous Infrastructure Management for Dablio
\n\nWhat Was Done
\nWe completed a dual Lightsail instance deployment paired with autonomous peer-sync and stall-watchdog mechanisms to provide redundant, self-healing infrastructure for Dablio's core services. This work involved establishing a primary and secondary Lightsail instance, implementing automated synchronization between instances, and deploying monitoring/recovery logic to detect and remediate service failures without manual intervention.
\n\nInfrastructure Architecture
\nThe deployment consists of:
\n- \n
- Lightsail Primary & Secondary Instances: Two independent compute instances deployed in separate availability zones, each running containerized Dablio services \n
- Autonomous Peer-Sync: A background process that continuously synchronizes configuration, secrets, and deployment state between primary and secondary instances using secure rsync over SSH with cryptographic verification \n
- Stall-Watchdog: Health monitoring agent that detects service degradation (stalled processes, memory exhaustion, network timeout) and triggers automatic failover or restart procedures \n
- Route53 Failover Routing: DNS-level failover configured to automatically route traffic to the secondary instance if the primary becomes unhealthy, detected via Route53 health checks \n
- CloudFront Distribution: Static assets and API responses cached at edge locations with TTL tuning optimized for deployment frequency \n
Key Technical Decisions
\n\nWhy Lightsail Over EC2
\nLightsail provides a managed, predictable networking layer with built-in firewall rules and public IP provisioning. For a bootstrapping infrastructure, this eliminates VPC/subnet complexity while providing ample compute for our workload. We avoided EC2 to reduce ops overhead during this scaling phase.
\n\nPeer-Sync Over Traditional Replication
\nRather than relying on database replication or container registries, peer-sync leverages file-level synchronization with cryptographic checksums. This approach:
\n- \n
- Works with stateful application code (not just data) \n
- Survives complete instance loss without external state recovery \n
- Tolerates partial network failures through incremental delta sync \n
- Provides transparent state visibility for debugging \n
Watchdog Over CloudWatch Alarms
\nWe deployed an in-host stall-watchdog agent rather than relying solely on CloudWatch metrics. The watchdog:
\n- \n
- Detects application-layer stalls (e.g., a worker process hung in a system call) not visible to CloudWatch \n
- Responds in <2 seconds vs. CloudWatch's 60-second+ metric aggregation window \n
- Triggers local remediation (process restart, cache flush) before escalating to Route53 failover \n
Implementation Details
\n\nPeer-Sync Configuration
\nThe peer-sync service runs as a systemd timer triggering every 30 seconds:
\nrsync -av --delete --checksum --exclude='.git' --exclude='node_modules' \\\n /opt/dablio/app/ root@SECONDARY_INSTANCE_IP:/opt/dablio/app/\n\nChecksums ensure only modified files transfer, reducing bandwidth. The --delete flag maintains true state parity, preventing orphaned files.
Stall-Watchdog Logic
\nThe watchdog monitors three categories of stalls:
\n- \n
- Process Stalls: Child processes in WCHAN state for >10 seconds trigger a SIGTERM and restart \n
- Memory Pressure: If free memory drops below 100MB, the watchdog flushes application caches and forcefully terminates idle connections \n
- Network Timeouts: TCP connections stuck in ESTABLISHED state for >30 seconds are forcibly closed; the service layer automatically reconnects \n
All recovery actions are logged to /var/log/dablio-watchdog.log with timestamps and recovery reason codes for post-incident analysis.
Route53 Health Check Configuration
\nPrimary health checks every 10 seconds over HTTP GET to /api/health. Secondary is marked failover target. A failure threshold of 3 consecutive failed checks triggers DNS routing change, propagating globally within 60 seconds.
CloudFront Distribution Setup
\nStatic assets (JS, CSS, images) cached with 7-day TTL in s3://dablio-assets-prod/ distribution. API responses cached with 60-second TTL at /api/cached/* paths. Origin headers automatically purged on deployment via Lambda trigger on S3 object updates.
Infrastructure as Code
\nAll Lightsail, Route53, and CloudFront resources are defined in Terraform under terraform/prod/:
- \n
terraform/prod/lightsail.tf— Primary and secondary instance specs, SSH key pairs, firewall rules \nterraform/prod/route53.tf— Failover record sets, health checks, TTL configuration \nterraform/prod/cloudfront.tf— Distribution origins, cache behaviors, OAI (Origin Access Identity) for S3 \n
To apply changes:
\ncd terraform/prod\nterraform plan -out=tfplan\nterraform apply tfplan\n\n\nMonitoring & Observability
\nKey metrics exported to CloudWatch every 60 seconds:
\n- \n
DablioApp/PeerSyncLag— Delta bytes pending sync (target: <100KB) \nDablioApp/WatchdogRestarts— Count of automatic restarts in 5-minute window (alert >5) \nDablioApp/FailoverEvents— Route53 routing changes (audit trail) \nDablioApp/CacheHitRatio— CloudFront cache effectiveness (target: >85%) \n
Logs aggregated to CloudWatch Logs under /aws/lightsail/dablio/primary and /aws/lightsail/dablio/secondary.
What's Next
\nImmediate follow-up work includes:
\n- \n
- Automated failback logic — once primary recovers, rebalance traffic back gradually over 10 minutes \n
- Cross-region replication — extend dual-instance pattern to a second AWS region for geographic redundancy \n
- Chaos engineering tests — deliberately trigger stalls and network partitions to validate recovery behavior \n
- Cost optimization — evaluate reserved Lightsail capacity discounts for the secondary instance \n
The foundation is now in place for true autonomous infrastructure management; future work focuses on expanding redundancy scope and validating failure scenarios under load.
"}}]