Deploying Autonomous Peer-Sync Infrastructure on AWS Lightsail: Architecture and Operational Decisions
What We Built
We deployed a dual Lightsail instance architecture for autonomous peer synchronization with stall detection and watchdog monitoring. This infrastructure powers the peer sync layer that coordinates across distributed team operations, with automated failover and operational observability built in from the ground up.
Architecture Overview
The deployment consists of:
- Primary Lightsail Instance: Hosts the peer sync daemon with autonomous operation capability
- Secondary Lightsail Instance (Twin): Standby replica for high-availability failover, synchronized via peer-sync protocol
- Stall Watchdog: Independent monitoring daemon that detects synchronization stalls and triggers operational alerts
- CloudFront Distribution: Caches peer metadata and health checks, reducing direct instance load
- Route53 Health Checks: Monitor primary instance health and automatically promote replica on failure
Infrastructure Configuration
Lightsail Instance Specifications
Both instances use consistent configuration for predictable failover:
- Region: AWS Lightsail deployed in us-east-1 (primary), with replica in us-east-1b
- Instance Plan: 2GB RAM, 1 vCPU baseline with burstable performance enabled
- Storage: 60GB SSD per instance with automated snapshots every 6 hours
- Static IP Assignment: Fixed elastic IPs for consistent DNS resolution
- Security Groups: Ingress rules on 443 (HTTPS peer sync), 22 (SSH ops only from hardened bastion), 8080 (internal health checks)
Networking and DNS
Route53 configuration enforces strict consistency:
- Primary Record:
peer-sync.sailjada.com→ Lightsail primary elastic IP with TTL 30s (tight coupling during active operation) - Health Check Configuration: HTTPS endpoint at
https://peer-sync.sailjada.com/healthchecked every 10 seconds with 3-consecutive-failure threshold before failover - Failover Logic: CloudWatch alarm triggers Route53 weighted routing shift to replica on health check failure
- CloudFront Distribution ID:
d-peer-metadata-cacheconfigured for /metadata/* paths, origin pull-through from primary with 60s cache TTL
Peer Sync Protocol Implementation
Autonomous Operation Mode
The peer sync daemon operates autonomously with minimal manual intervention:
// Core sync loop pseudocode
for {
state := readLocalState()
remote := fetchPeerState(remoteInstance)
merged := mergeStates(state, remote)
if hasConflict(merged) {
// Autonomous resolution: prefer highest lamport clock
resolved := resolveByTimestamp(merged)
applyState(resolved)
broadcastResolution()
}
time.Sleep(5 * time.Second)
}
Key behaviors:
- Lamport Clock Versioning: Each state change increments a logical timestamp; ties resolved by instance ID (primary always wins)
- Idempotent Merges: All state operations are designed to be replayable without side effects
- No Manual Intervention Required: The system converges to consistency automatically within the sync interval (5 seconds default)
Stall Detection and Watchdog
A separate watchdog process monitors sync health:
- Stall Threshold: If primary hasn't updated remote state in 30 seconds, mark stall
- Watchdog Action: Emit CloudWatch metric
peer-sync/stall, trigger SNS notification to ops - Automatic Recovery: After 2 consecutive stall detections, initiate controlled failover (replica promoted)
- Log Location: All stall events logged to
/var/log/peer-sync-watchdog.logwith full state dumps on timeout
Deployment Process and Automation
Infrastructure as Code (Terraform)
All Lightsail resources defined in:
terraform/aws/peer-sync/
├── main.tf # Primary instance definition
├── replica.tf # Secondary instance (twin)
├── networking.tf # Route53 + CloudFront setup
├── security.tf # Security groups and IAM roles
└── variables.tf # Environment-specific overrides
Key Terraform resources:
aws_lightsail_instance.peer_sync_primary: Primary instance with user data script for daemon bootstrapaws_lightsail_instance.peer_sync_replica: Replica with identical configuration, tagged for automated syncaws_route53_record.peer_sync: Weighted routing policy (100% primary, 0% replica until failover)aws_cloudwatch_alarm.peer_stall: Triggers on stall metric > 0, notification to SNS topicarn:aws:sns:us-east-1:ACCOUNT:peer-ops-alerts
Deployment Commands
# Plan infrastructure changes
terraform plan -var-file=prod.tfvars -out=tfplan
# Apply with automatic approval (gated by code review)
terraform apply tfplan
# Verify peer sync is operational on both instances
ssh ubuntu@LIGHTSAIL_IP "systemctl status peer-sync-daemon"
# Tail live sync logs
ssh ubuntu@LIGHTSAIL_IP "tail -f /var/log/peer-sync.log | grep -v DEBUG"
Key Architectural Decisions
Why Lightsail for Peer Sync?
Decision: Use AWS Lightsail instead of EC2 or ECS for peer sync infrastructure.
Rationale: Lightsail provides fixed, predictable pricing with integrated static IPs and snapshots—critical for maintaining peer identity across reboots. ECS complexity would add operational burden (service definitions, task placement) without benefit for a single daemon per instance. The peer sync protocol requires stable IP addresses and direct SSH access for troubleshooting, both native to Lightsail's model.
Autonomous Operation Without Consensus Protocols
Decision: Implement autonomous merge-by-timestamp rather than RAFT/Paxos consensus.
Rationale: Peer sync state is operational metadata (team status, task assignments), not critical financial data. A simple Lamport clock merge with deterministic tie-breaking (primary instance ID) provides fast convergence and eliminates leader election complexity. If a state loss occurs, it's recoverable from the source system (task queue, member database). This trades off consistency guarantees for operational simplicity and reduced latency.
Stall Watchdog as Separate Process
Decision: Run stall detection in a separate daemon rather than in-process health checks.
Rationale: In-process checks are blind to application deadlocks—if the peer sync daemon hangs, it won't report itself as stalled. A separate watchdog process monitors from outside, can be restarted independently, and provides an external signal to automated failover logic. This is a proven pattern in systems like etcd and Consul.
Operational Monitoring and Alerts
- Primary Metric:
peer-sync/state-update-latency(p95 target: <500ms) - Health Dashboard: CloudWatch dashboard at
dashboards/peer-sync-operationswith replica lag, stall count, failover events - Alert Channels: SNS to ops-team Slack, email fallback for critical stalls (2+ consecutive)
- Audit Trail: All state merges logged with operator/source for compliance; retention 90 days in CloudWatch Logs group
/aws/lightsail/peer-sync
What's Next
The autonomous peer-sync system is now operational with dual-instance redundancy and stall detection. Upcoming work includes: (1) adding multi-region replica support (us-west-2 standby), (2) implementing state compression to reduce network overhead, and (3) integrating with Queen's Fleet operations for distributed team coordination across multiple deployments.
Last deployed: July 5, 2026. Repository: /Users/cb/dablio, branch master. Tickets tracking this work: See TICKETS.md for active items and blockers.