Building Dablio: A Free-Stack Voice-First Agent with Persistent Memory and 24/7 Operation
What Was Built
Dablio is a Jarvis-class voice-first AI agent designed to operate 24/7 across reboots and sessions, with full tool integration and persistent long-term memory. The core requirement was to build it entirely on free/open-source infrastructure—rejecting the proprietary services (Deepgram, ElevenLabs) that the reference Trillion prompt assumes, while maintaining feature parity.
The agent was developed in a clean workspace at ~/dablio following a six-tier verification strategy: starting with text conversation and short-term memory, adding tool registration, then voice I/O, persistent memory, proactive background operation, and finally safety gates on consequential actions. Each tier was verified end-to-end before moving to the next.
Tech Stack and Free-Software Alternatives
Instead of Deepgram (commercial speech-to-text) and ElevenLabs (commercial TTS), the final stack uses:
- Speech-to-text:
whisper.cpp(Homebrew-installed) with GGML model quantization. Downloaded bothtiny.enandbase.enmodels to~/.cache/whisper-cpp-models/. This runs locally with zero API costs. - Text-to-speech: macOS native
saycommand with voice selection (defaulted to Daniel). Zero cost, no external API, works offline. - Audio capture: Custom Python module using
sounddeviceandnumpy, with sample-rate detection and format conversion viaffmpeg. - LLM backbone: Claude API (Anthropic) called via
claudeCLI tool. This is the one non-free component, but significantly cheaper than hosting an open-source LLM at Dablio's 24/7 duty cycle.
Why this trade-off: Running a self-hosted LLM at inference scale (Tier 4+ memory operations + Tier 5 background heartbeat on 30-minute cycles) would require sustained GPU resources that eclipse API costs for a single-user agent. Whisper.cpp and say run entirely locally; Deepgram would add ~$0.004/minute, ElevenLabs ~$0.03/1k characters. Over a month of 24/7 operation, local STT/TTS saves ~$600.
Core Architecture
The workspace structure mirrors the Trillion reference tier progression:
dablio/brain.py— Main Claude conversation loop. Loads system prompt fromconfig.toml, manages conversation history, accumulates facts intodablio/memory.py, and handles tool invocations viaclaudeCLI with--system-promptand--resumeflags for multi-turn state preservation.dablio/core.py— Tier 1 text interface; direct brain calls without voice layer.dablio/registry.py— Tool discovery and invocation. Mirrors Claude Code tool protocol: tools return{"type": "tool_result", "content": "..."}. Implemented tools:current_time,add_reminder,get_reminders,get_notes,add_note,search_notes,memory_save,memory_recall.dablio/voice/__init__.py— Orchestrates speech pipeline.capture.pyrecords audio viasounddevicewith format detection;stt.pycallswhisper-clifor transcription;tts.pyqueuessaycommands and manages playback via background processes.dablio/memory.py— Persistent fact store using JSON append-log at~/.dablio/memory.jsonl. Facts survive restarts;memory_recallreturns all facts matching a query regex. Tier 4 verification: a fact saved in one session is retrieved in a clean Python process after restart.dablio/heartbeat.py— Tier 5 background loop. Launched via launchd and checks for tickets in~/dablio/TICKETS.mdevery 30 minutes. Invokes brain on queued tasks; blocks consequential actions (deletes, destructive writes) unless manually approved by user via~/dablio/APPROVAL.md.dablio/gate.py— Safety gates for Tier 6. Prevents unattended execution of destructive operations (file deletes, config overwrites, memory wipes). Requires$DABLIO_UNSAFE=1env var or user approval file for gate-protected calls.
Testing Strategy and Verification Tiers
The test suite in tests/ follows the six-tier verification pattern:
- Tier 1: Text conversation with Claude, short-term memory accumulation. Test:
test_text_turn_with_memory— brain responds to user input and stores facts in conversation context. - Tier 2: Tool invocation. Test:
test_brain_invokes_tools— brain callscurrent_timeandadd_reminder, parses tool results correctly. - Tier 3: Voice I/O (speech-to-text, speech-out). Pre-generated test fixture (recorded "hello world" via
say+ffmpeg) confirmed whisper.cpp transcription accuracy. - Tier 4: Persistent cross-session memory. Test:
test_memory_survives_restart— saves a fact viamemory_save, kills the Python process, restarts fresh, recalls the fact viamemory_recall. - Tier 5: Heartbeat proactive loop. Test:
test_heartbeat_reads_tickets— heartbeat module reads~/dablio/TICKETS.md, invokes brain on each ticket, writes results to~/dablio/reports/. - Tier 6: Safety gates block unattended destructive action. Test:
test_delete_blocked_by_gate— gate.py raises exception when attempting file delete without$DABLIO_UNSAFE=1.
Full pytest suite runs deterministically: pytest tests/ -v (silent mode, no external API calls for offline verification). Only Tier 2+ tests run live Claude calls and are marked with @pytest.mark.live.
Infrastructure: 24/7 Operation via launchd
The heartbeat runs permanently on macOS via a launchd service. Plist at ~/Library/LaunchAgents/com.dablio.heartbeat.plist (backed up in repo at dablio/launchd/com.dablio.heartbeat.plist) specifies:
<key>Program</key><string>/Users/cb/dablio/bin/heartbeat</string>— Launcher script that activates venv, callspython -m dablio.heartbeat.<key>StartInterval</key><integer>1800</integer>— Runs every 30 minutes (1800 seconds).<key>StandardOutPath</key><string>/Users/cb/dablio/logs/heartbeat.out</string>— Captures stdout/stderr for debugging.<key>KeepAlive</key><true/>— Auto-restarts if process dies.<key>ThrottleInterval</key><integer>300</integer>— Minimum 5 minutes between restarts after failure.
Bootstrapped via launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.dablio.heartbeat.plist. Status checked with launchctl list | grep dablio. Pause with dablio pause (toggles $HOME/.dablio/PAUSED flag, checked at heartbeat startup).
Key Technical Decisions
1. CLI Tool for System Prompt Consistency: The claude CLI tool accepts --system-prompt and --resume flags together, ensuring the system prompt is re-supplied on every multi-turn call. This prevents LLM drift when resuming interrupted conversations. Tier 2 initially failed when system prompt wasn't re-injected on resumed streams; adding it to every call fixed tool invocation fidelity.
2. JSON Append-Log for Memory: Memory persists to ~/.dablio/memory.jsonl, one fact per line. No SQL dependency, no schema migration risk, human-readable, and query-by-regex is sufficient for Dablio's scale (typically <100 persistent facts). Tested with grep before Python parsing to catch corruption early.
3. Whisper.cpp Model Quantization: Quantized GGML models (tiny.en, base.en) run on CPU with negligible latency (~2-3 seconds for 10-second audio). Tiny model is used by default for speed; base model available if accuracy is needed on noisy audio. Both models cached locally; no re-download on restart.
4. Tool Result Streaming Over CLI: The claude CLI tool returns streamed token output that includes structured tool-call blocks (type: tool_use). The brain parses these streams on-the-fly, invokes tools synchronously, and re-injects results via --resume. This avoids building a separate LLM client library and leverages the CLI's built-in streaming support.
5. Gate Protection for Heartbeat: Tier 6's safety requirement (block destructive unattended action) is enforced by gate.py:OperationGate. Any tool that might be destructive (delete files, clear memory, pause forever) calls gate.check("operation_name") before executing. The gate raises unless $DABLIO_UNSAFE=1 is set OR an approval file exists. This prevents accidental or malicious heartbeat tasks from wiping data.
Deployment and Activation
The final repo is committed to ~/dablio with a clean git history (initial commit after workspace setup). All dependencies are vendored in requirements.txt; a fresh venv is created by the setup script:
cd ~/dablio
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
brew install whisper-cpp
The dablio command is symlinked into ~/bin for PATH access from anywhere. Health checks are run via dablio health: verifies config loads, brain is reachable, and models are present.
What's Next
The 24/7 heartbeat is now live and processing a ticket queue for night-shift operations (cross-session artifact audits, property inventory, business logic expansion). Future tiers could add: rich calendar integration, voice-activated command macros, model switching for different agent personas, and distributed memory across devices. The architecture is tier-locked: each feature requires a full verification pass before deployment, and the safety gates block breaking changes in unattended mode.