24/7 Operational Heartbeat: Building Continuous Observability into Dablio's Core Architecture
What Was Done
On July 3rd, 2026, the Dablio engineering team completed a significant architectural milestone: deploying a production-ready heartbeat and observability system designed to run 24/7, unattended. This work implements the core infrastructure for continuous operational awareness—moving from day-shift event logging to always-on system monitoring. The solution combines launchd-based process management with a unified state tracking mechanism that serves as the foundation for autonomous agent operations.
Technical Architecture
Process Management with launchd
Rather than relying on traditional cron or manual process invocation, we leveraged macOS's native launchd daemon framework to ensure the heartbeat runs continuously without operator intervention. This decision trades the simplicity of ad-hoc commands for reliability and automatic recovery:
- Plist Configuration: Created
~/Library/LaunchAgents/com.dablio.heartbeat.plistwith standard launchd properties:Label,Program,RunAtLoad,StartInterval, andStandardOutPath/StandardErrorPathfor log capture. - Auto-Recovery: launchd automatically restarts the process if it exits unexpectedly, critical for overnight operations where manual intervention isn't possible.
- Logging: All heartbeat output directed to
~/Library/Logs/dablio-heartbeat.log, enabling post-hoc analysis of overnight behavior.
Symbolic Linking for Portability
The heartbeat script was exposed in ~/bin/dablio via symlink, creating a canonical entrypoint that remains valid regardless of the repository or installation location. This pattern allows the system to survive directory renames or migrations without updating launchd configuration:
ln -s /Users/cb/dablio/scripts/heartbeat.py ~/bin/dablio
# launchd plist references ~/bin/dablio directly, decoupled from repo path
Why this approach: Absolute paths in launchd plists are brittle. A symlink at a stable location lets us update the actual script location without touching the daemon configuration.
Infrastructure & State Management
Unified State Tracking
The heartbeat feeds into /Users/cb/dablio/state/, a directory-based state store serving as the source of truth for operational status. Rather than distributed logs or external metrics systems, we use the filesystem as our operational backbone:
state/heartbeat.json— Timestamp and health snapshot updated every 5 minutesstate/sessions.json— Running and recently-completed sessions, refreshed on each heartbeatstate/errors.jsonl— Append-only error log captured during heartbeat execution
This filesystem-first approach integrates cleanly with the existing estate architecture and avoids adding external dependencies (no Redis, no database). The tradeoff: scale limitations if we grow beyond single-host operations, but appropriate for current scope.
Night-Shift Ticket Pipeline
Completed sessions are logged to briefings/YYYY-MM-DD-day-log.md, a markdown manifest that pairs with the heartbeat state. Each entry contains:
- Session ID and timestamp
- Claimed artifacts:
report:,deploy:,engine:,site: - Status and summary notes
The day-log serves dual purposes: human-readable briefing for the next operational session, and machine-verifiable manifest for artifact auditing (enabling cross-check workflows).
Key Architectural Decisions
Why launchd over Docker/Kubernetes
For a 24/7 agent running on a single development machine, launchd provides native, zero-ops process management. Docker adds abstraction burden; Kubernetes is overkill. launchd is already installed, integrates with macOS logging, and requires only a plist file.
Why Filesystem State over External Metrics
We prioritized operational simplicity and auditability. All state lives in version-controllable directories. This makes it easy to audit history, snapshot state for analysis, or replay sessions. External metrics systems (Datadog, Prometheus) offer scale but add operational overhead and vendor lock-in—not justified for current throughput.
Why Symlinks for Script Entrypoints
The symlink at ~/bin/dablio decouples deployment location from daemon configuration. This is critical for development: we can move the repo, update the symlink once, and launchd continues working without redeployment. Production benefit: migrations and backups don't require daemon reconfiguration.
Deployment & Verification
Verification steps for night-shift readiness:
# Load plist into launchd
launchctl load ~/Library/LaunchAgents/com.dablio.heartbeat.plist
# Verify it's running
launchctl list | grep dablio
# Tail the log for live heartbeat output
tail -f ~/Library/Logs/dablio-heartbeat.log
# Check state files were created
ls -la ~/Library/Logs/dablio-heartbeat.log
cat state/heartbeat.json # Should have recent timestamp
All artifacts verified to exist:
- ✓ launchd plist configuration file created and loaded
- ✓ Symlink at
~/bin/dabliopointing to heartbeat script - ✓ State directory (
/Users/cb/dablio/state/) initialized with heartbeat output - ✓ Log file at
~/Library/Logs/dablio-heartbeat.logreceiving heartbeat events - ✓ Day-log briefing file (
briefings/2026-07-03-day-log.md) created with session manifest
What's Next
With the heartbeat running 24/7, the next phases are:
- Autonomous Actions: Wire heartbeat signals to automatic remediation (restart services, alert on errors, escalate stalled sessions).
- Multi-Host Coordination: Extend state tracking to support distributed agents (multiple machines, redundancy).
- Metrics Ingestion: Build a metrics aggregation layer that can consume heartbeat state and feed dashboards for long-term trend analysis.
Summary: The 24/7 heartbeat infrastructure is now in production, using launchd for reliable daemon management and filesystem-based state tracking for operational observability. All artifacts are verified and logged in the day-log briefing; the system is ready for unattended overnight operations.