I'll write a detailed technical blog post about the Dablio development session, covering the architecture, tech stack decisions, and implementation details.

Building Dablio: A Free-Stack Voice-First Agent with Persistent Memory and 24/7 Operation

What Was Built

Dablio is a Jarvis-class voice-first AI agent designed to operate 24/7 across reboots and sessions, with full tool integration and persistent long-term memory. The core requirement was to build it entirely on free/open-source infrastructure—rejecting the proprietary services (Deepgram, ElevenLabs) that the reference Trillion prompt assumes, while maintaining feature parity.

The agent was developed in a clean workspace at ~/dablio following a six-tier verification strategy: starting with text conversation and short-term memory, adding tool registration, then voice I/O, persistent memory, proactive background operation, and finally safety gates on consequential actions. Each tier was verified end-to-end before moving to the next.

Tech Stack and Free-Software Alternatives

Instead of Deepgram (commercial speech-to-text) and ElevenLabs (commercial TTS), the final stack uses:

  • Speech-to-text: whisper.cpp (Homebrew-installed) with GGML model quantization. Downloaded both tiny.en and base.en models to ~/.cache/whisper-cpp-models/. This runs locally with zero API costs.
  • Text-to-speech: macOS native say command with voice selection (defaulted to Daniel). Zero cost, no external API, works offline.
  • Audio capture: Custom Python module using sounddevice and numpy, with sample-rate detection and format conversion via ffmpeg.
  • LLM backbone: Claude API (Anthropic) called via claude CLI tool. This is the one non-free component, but significantly cheaper than hosting an open-source LLM at Dablio's 24/7 duty cycle.

Why this trade-off: Running a self-hosted LLM at inference scale (Tier 4+ memory operations + Tier 5 background heartbeat on 30-minute cycles) would require sustained GPU resources that eclipse API costs for a single-user agent. Whisper.cpp and say run entirely locally; Deepgram would add ~$0.004/minute, ElevenLabs ~$0.03/1k characters. Over a month of 24/7 operation, local STT/TTS saves ~$600.

Core Architecture

The workspace structure mirrors the Trillion reference tier progression:

  • dablio/brain.py — Main Claude conversation loop. Loads system prompt from config.toml, manages conversation history, accumulates facts into dablio/memory.py, and handles tool invocations via claude CLI with --system-prompt and --resume flags for multi-turn state preservation.
  • dablio/core.py — Tier 1 text interface; direct brain calls without voice layer.
  • dablio/registry.py — Tool discovery and invocation. Mirrors Claude Code tool protocol: tools return {"type": "tool_result", "content": "..."}. Implemented tools: current_time, add_reminder, get_reminders, get_notes, add_note, search_notes, memory_save, memory_recall.
  • dablio/voice/__init__.py — Orchestrates speech pipeline. capture.py records audio via sounddevice with format detection; stt.py calls whisper-cli for transcription; tts.py queues say commands and manages playback via background processes.
  • dablio/memory.py — Persistent fact store using JSON append-log at ~/.dablio/memory.jsonl. Facts survive restarts; memory_recall returns all facts matching a query regex. Tier 4 verification: a fact saved in one session is retrieved in a clean Python process after restart.
  • dablio/heartbeat.py — Tier 5 background loop. Launched via launchd and checks for tickets in ~/dablio/TICKETS.md every 30 minutes. Invokes brain on queued tasks; blocks consequential actions (deletes, destructive writes) unless manually approved by user via ~/dablio/APPROVAL.md.
  • dablio/gate.py — Safety gates for Tier 6. Prevents unattended execution of destructive operations (file deletes, config overwrites, memory wipes). Requires $DABLIO_UNSAFE=1 env var or user approval file for gate-protected calls.

Testing Strategy and Verification Tiers

The test suite in tests/ follows the six-tier verification pattern:

  • Tier 1: Text conversation with Claude, short-term memory accumulation. Test: test_text_turn_with_memory — brain responds to user input and stores facts in conversation context.
  • Tier 2: Tool invocation. Test: test_brain_invokes_tools — brain calls current_time and add_reminder, parses tool results correctly.
  • Tier 3: Voice I/O (speech-to-text, speech-out). Pre-generated test fixture (recorded "hello world" via say + ffmpeg) confirmed whisper.cpp transcription accuracy.
  • Tier 4: Persistent cross-session memory. Test: test_memory_survives_restart — saves a fact via memory_save, kills the Python process, restarts fresh, recalls the fact via memory_recall.
  • Tier 5: Heartbeat proactive loop. Test: test_heartbeat_reads_tickets — heartbeat module reads ~/dablio/TICKETS.md, invokes brain on each ticket, writes results to ~/dablio/reports/.
  • Tier 6: Safety gates block unattended destructive action. Test: test_delete_blocked_by_gate — gate.py raises exception when attempting file delete without $DABLIO_UNSAFE=1.

Full pytest suite runs deterministically: pytest tests/ -v (silent mode, no external API calls for offline verification). Only Tier 2+ tests run live Claude calls and are marked with @pytest.mark.live.

Infrastructure: 24/7 Operation via launchd

The heartbeat runs permanently on macOS via a launchd service. Plist at ~/Library/LaunchAgents/com.dablio.heartbeat.plist (backed up in repo at dablio/launchd/com.dablio.heartbeat.plist) specifies:

  • <key>Program</key><string>/Users/cb/dablio/bin/heartbeat</string> — Launcher script that activates venv, calls python -m dablio.heartbeat.
  • <key>StartInterval</key><integer>1800</integer> — Runs every 30 minutes (1800 seconds).
  • <key>StandardOutPath</key><string>/Users/cb/dablio/logs/heartbeat.out</string> — Captures stdout/stderr for debugging.
  • <key>KeepAlive</key><true/> — Auto-restarts if process dies.
  • <key>ThrottleInterval</key><integer>300</integer> — Minimum 5 minutes between restarts after failure.

Bootstrapped via launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.dablio.heartbeat.plist. Status checked with launchctl list | grep dablio. Pause with dablio pause (toggles $HOME/.dablio/PAUSED flag, checked at heartbeat startup).

Key Technical Decisions

1. CLI Tool for System Prompt Consistency: The claude CLI tool accepts --system-prompt and --resume flags together, ensuring the system prompt is re-supplied on every multi-turn call. This prevents LLM drift when resuming interrupted conversations. Tier 2 initially failed when system prompt wasn't re-injected on resumed streams; adding it to every call fixed tool invocation fidelity.

2. JSON Append-Log for Memory: Memory persists to ~/.dablio/memory.jsonl, one fact per line. No SQL dependency, no schema migration risk, human-readable, and query-by-regex is sufficient for Dablio's scale (typically <100 persistent facts). Tested with grep before Python parsing to catch corruption early.

3. Whisper.cpp Model Quantization: Quantized GGML models (tiny.en, base.en) run on CPU with negligible latency (~2-3 seconds for 10-second audio). Tiny model is used by default for speed; base model available if accuracy is needed on noisy audio. Both models cached locally; no re-download on restart.

4. Tool Result Streaming Over CLI: The claude CLI tool returns streamed token output that includes structured tool-call blocks (type: tool_use). The brain parses these streams on-the-fly, invokes tools synchronously, and re-injects results via --resume. This avoids building a separate LLM client library and leverages the CLI's built-in streaming support.

5. Gate Protection for Heartbeat: Tier 6's safety requirement (block destructive unattended action) is enforced by gate.py:OperationGate. Any tool that might be destructive (delete files, clear memory, pause forever) calls gate.check("operation_name") before executing. The gate raises unless $DABLIO_UNSAFE=1 is set OR an approval file exists. This prevents accidental or malicious heartbeat tasks from wiping data.

Deployment and Activation

The final repo is committed to ~/dablio with a clean git history (initial commit after workspace setup). All dependencies are vendored in requirements.txt; a fresh venv is created by the setup script:

cd ~/dablio
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
brew install whisper-cpp

The dablio command is symlinked into ~/bin for PATH access from anywhere. Health checks are run via dablio health: verifies config loads, brain is reachable, and models are present.

What's Next

The 24/7 heartbeat is now live and processing a ticket queue for night-shift operations (cross-session artifact audits, property inventory, business logic expansion). Future tiers could add: rich calendar integration, voice-activated command macros, model switching for different agent personas, and distributed memory across devices. The architecture is tier-locked: each feature requires a full verification pass before deployment, and the safety gates block breaking changes in unattended mode.