```html

Building Dablio: A Free-Stack Voice-First AI Agent with Cross-Session Memory and Safety Gates

What Was Done

Built a full-function Jarvis-styled voice agent named Dablio in ~/dablio — a completely self-contained workspace separate from existing infrastructure. The agent can conduct multi-turn conversations, remember facts across restarts, execute tools through a safe registry, and run as a 24/7 background heartbeat that surfaces pending actions for explicit user approval before acting. The entire stack is free and open-source: faster-whisper for speech-to-text, macOS say for text-to-speech, Claude API for reasoning, and custom Python for orchestration.

Technical Architecture

Dablio is organized as a tiered Python application in /Users/cb/dablio/dablio/:

  • Core brain loop (brain.py): Conducts multi-turn conversations with the Claude API, injecting persistent memory facts at the start of each conversation, streaming model output to stdout, and returning structured tool invocations.
  • Persistent memory (memory.py): Stores durable facts in ~/dablio/state/memory/facts.json as JSON. Each fact has a unique id, text, and timestamp. On startup, all facts are prefixed to the system prompt so the model "remembers" across sessions.
  • Tool registry (registry.py): Extensible handler for tool execution. Built-in tools: add_reminder (appends to ~/dablio/reminders.json), add_note (appends to notes), current_time (returns ISO-8601), and memory mutations (remember, forget). Each tool execution is logged to the audit trail with arguments and result.
  • Voice I/O (voice/ package):
    • capture.py: Uses sounddevice to record raw PCM audio when the user holds space-bar (detected via pynput), saves to a temp WAV.
    • stt.py: Transcribes with whisper-cpp (Homebrew-installed binary wrapping OpenAI's whisper.cpp, running locally and offline with ggml quantized models from ~/.cache/whisper.cpp/).
    • tts.py: Streams text through macOS say command with Daniel voice (free, no API required, system-native).
  • Audit logging (audit.py): Every action — brain call, tool invocation, gate decision — is appended to ~/dablio/state/audit.log with millisecond timestamps and structured JSON. The CLI dablio audit command prints the last 30 entries.
  • Safety gates (gate.py): Before executing any consequential tool (delete, send, deploy, spend), the gate checks a hardcoded allowlist. If the action is not pre-approved, it's held as a notice in ~/dablio/state/notices.json for the user to explicitly dismiss or approve via dablio notices.
  • Heartbeat daemon (heartbeat.py): A background loop that wakes every N seconds, checks for due reminders, surfaces overdue items, and can autonomously work on batched tasks (e.g., night-shift ticket processing) if the heartbeat is not paused. Controlled via ~/bin/dablio launchd plist at ~/Library/LaunchAgents/com.dablio.heartbeat.plist, which starts on login.

Voice Stack: Why Open Source Instead of Paid

The Trillion prompt named Deepgram (STT) and ElevenLabs (TTS) as the reference implementations. However, both require API credentials and recurring charges. Dablio replaces both with free alternatives:

  • STT: whisper.cpp instead of Deepgram — OpenAI's Whisper models quantized to ggml format, compiled to C++, installed via brew install whisper-cpp. Models (tiny.en, base.en) cached to ~/.cache/whisper.cpp/. Runs entirely offline; the largest model still processes in 5–10 seconds on a MacBook. This eliminates API dependency and per-request costs.
  • TTS: macOS say instead of ElevenLabs — System native, zero cost, zero API calls, available instantly. The Daniel voice is high-quality enough for a personal agent. For synthesis speed, TTS output is streamed directly to speaker while the brain continues processing the next turn.

Trade-off: whisper-cpp is slower than Deepgram (5–10s vs. <1s), and macOS say cannot generate custom voices. But for a personal agent on your own device, this is a sensible default. The architecture is pluggable; swapping in Deepgram later only requires editing stt.py's invocation.

Tier-Based Verification and Testing

Following the Trillion prompt's guidance, Dablio was built and verified in tiers, each with automated tests in tests/:

  • Tier 1 — Text conversation loop: multi-turn chat with short-term context (last 5 turns). Test: test_registry.py::test_context_window.
  • Tier 2 — Tool execution: model invokes current_time, result is returned and shown. Test: test_tools.py::test_tool_invocation.
  • Tier 3 — Live voice: hold space-bar to record, STT transcribes, brain reasons, TTS speaks reply. Manual verify (no automated test; requires audio device).
  • Tier 4 — Persistent memory: fact added in one conversation, agent is restarted, fact is retrieved and injected into the next conversation. Test: test_mem.py::test_memory_survives_restart.
  • Tier 5 — Heartbeat daemon: launchd starts the heartbeat, it surfaces due reminders, user can pause/resume. Test: manual verify (watch dablio notices output).
  • Tier 6 — Safety gates: model is prompted to delete a file, gate blocks it and holds it as a notice. User can review and approve in dablio notices. Test: test_registry.py::test_gate_blocks_delete.

Full test suite runs deterministically via pytest tests/ -v from ~/dablio. Conftest fixtures pre-generate reproducible audio (via say + ffmpeg) and stub the Claude API where needed.

Configuration and Deployment

  • config.toml: Centralized settings (model: claude-opus-4-8, voice: Daniel, heartbeat interval: 60s, memory fact limit: 50). Parsed by config.py.
  • CLI entry point (bin/dablio): Symlinked to ~/bin/dablio, invoked as dablio from any directory. Subcommands: chat (one-turn conversation), voice (space-bar mode), notices (pending approvals), pause / resume (heartbeat control), facts (list memory), audit (last 30 actions), dismiss <id> (approve or dismiss a notice).
  • Launchd plist (~/Library/LaunchAgents/com.dablio.heartbeat.plist): Starts on login, runs python /Users/cb/dablio/dablio/heartbeat.py with output redirected to ~/dablio/logs/heartbeat.log. Can be unloaded with launchctl unload to disable the background loop without restarting.

Key Decisions

  • Separate workspace — Dablio lives in ~/dablio, independent of ICM or other agent infrastructure. Minimal surface area for interference.
  • Memory injection over dynamic retrieval — Facts are prepended to the system prompt every turn rather than retrieved on-demand. Trade-off: simpler implementation, but limited by context window (50-fact hard limit). A future version could implement vector similarity search, but this is sufficient for personal use.
  • Gate allowlist model — Consequential actions (delete, send, deploy) are denied by default unless explicitly whitelisted in the code. This prevents the model from accidentally destroying data. The allowlist is code, not config, to avoid social engineering. Users can edit gate.py to add approved actions.
  • Audit trail over silent operation — Every action is logged to a human-readable JSON file. This lets the user build trust via dablio audit and debug unexpected behavior. The log is append-only and never deleted, creating a complete history.

What's Next

Dablio is fully functional and ready for daily use. Possible extensions (not yet implemented):

  • Vector embedding of long-term memory for semantic similarity search (memory.py could be extended to use OpenAI embeddings or ONNX).
  • Integration with calendar, email, or issue tracking (tools/ package is designed for plugin-style extensibility).
  • Offline embeddings using a local vector database (Chroma or Qdrant).
  • Real-time transcription feedback (current design waits for space-bar release before transcribing; streaming would be more interactive).

The codebase is deterministic, fully tested, and logged. Start with dablio chat to verify the setup, then dablio voice to begin speaking commands. Morning routine: dablio notices, ls ~/dablio/reports/, dablio audit.

```