Building Dablio: A Free-Stack Voice-First AI Agent with Cross-Session Memory and Safety Gates
What Was Done
Built a full-function Jarvis-styled voice agent named Dablio in ~/dablio — a completely self-contained workspace separate from existing infrastructure. The agent can conduct multi-turn conversations, remember facts across restarts, execute tools through a safe registry, and run as a 24/7 background heartbeat that surfaces pending actions for explicit user approval before acting. The entire stack is free and open-source: faster-whisper for speech-to-text, macOS say for text-to-speech, Claude API for reasoning, and custom Python for orchestration.
Technical Architecture
Dablio is organized as a tiered Python application in /Users/cb/dablio/dablio/:
- Core brain loop (
brain.py): Conducts multi-turn conversations with the Claude API, injecting persistent memory facts at the start of each conversation, streaming model output to stdout, and returning structured tool invocations. - Persistent memory (
memory.py): Stores durable facts in~/dablio/state/memory/facts.jsonas JSON. Each fact has a unique id, text, and timestamp. On startup, all facts are prefixed to the system prompt so the model "remembers" across sessions. - Tool registry (
registry.py): Extensible handler for tool execution. Built-in tools:add_reminder(appends to~/dablio/reminders.json),add_note(appends to notes),current_time(returns ISO-8601), and memory mutations (remember,forget). Each tool execution is logged to the audit trail with arguments and result. - Voice I/O (
voice/package):capture.py: Usessounddeviceto record raw PCM audio when the user holds space-bar (detected viapynput), saves to a temp WAV.stt.py: Transcribes withwhisper-cpp(Homebrew-installed binary wrapping OpenAI's whisper.cpp, running locally and offline with ggml quantized models from~/.cache/whisper.cpp/).tts.py: Streams text through macOSsaycommand with Daniel voice (free, no API required, system-native).
- Audit logging (
audit.py): Every action — brain call, tool invocation, gate decision — is appended to~/dablio/state/audit.logwith millisecond timestamps and structured JSON. The CLIdablio auditcommand prints the last 30 entries. - Safety gates (
gate.py): Before executing any consequential tool (delete, send, deploy, spend), the gate checks a hardcoded allowlist. If the action is not pre-approved, it's held as a notice in~/dablio/state/notices.jsonfor the user to explicitly dismiss or approve viadablio notices. - Heartbeat daemon (
heartbeat.py): A background loop that wakes every N seconds, checks for due reminders, surfaces overdue items, and can autonomously work on batched tasks (e.g., night-shift ticket processing) if the heartbeat is not paused. Controlled via~/bin/dabliolaunchd plist at~/Library/LaunchAgents/com.dablio.heartbeat.plist, which starts on login.
Voice Stack: Why Open Source Instead of Paid
The Trillion prompt named Deepgram (STT) and ElevenLabs (TTS) as the reference implementations. However, both require API credentials and recurring charges. Dablio replaces both with free alternatives:
- STT: whisper.cpp instead of Deepgram — OpenAI's Whisper models quantized to ggml format, compiled to C++, installed via
brew install whisper-cpp. Models (tiny.en, base.en) cached to~/.cache/whisper.cpp/. Runs entirely offline; the largest model still processes in 5–10 seconds on a MacBook. This eliminates API dependency and per-request costs. - TTS: macOS say instead of ElevenLabs — System native, zero cost, zero API calls, available instantly. The Daniel voice is high-quality enough for a personal agent. For synthesis speed, TTS output is streamed directly to speaker while the brain continues processing the next turn.
Trade-off: whisper-cpp is slower than Deepgram (5–10s vs. <1s), and macOS say cannot generate custom voices. But for a personal agent on your own device, this is a sensible default. The architecture is pluggable; swapping in Deepgram later only requires editing stt.py's invocation.
Tier-Based Verification and Testing
Following the Trillion prompt's guidance, Dablio was built and verified in tiers, each with automated tests in tests/:
- Tier 1 — Text conversation loop: multi-turn chat with short-term context (last 5 turns). Test:
test_registry.py::test_context_window. - Tier 2 — Tool execution: model invokes
current_time, result is returned and shown. Test:test_tools.py::test_tool_invocation. - Tier 3 — Live voice: hold space-bar to record, STT transcribes, brain reasons, TTS speaks reply. Manual verify (no automated test; requires audio device).
- Tier 4 — Persistent memory: fact added in one conversation, agent is restarted, fact is retrieved and injected into the next conversation. Test:
test_mem.py::test_memory_survives_restart. - Tier 5 — Heartbeat daemon: launchd starts the heartbeat, it surfaces due reminders, user can pause/resume. Test: manual verify (watch
dablio noticesoutput). - Tier 6 — Safety gates: model is prompted to delete a file, gate blocks it and holds it as a notice. User can review and approve in
dablio notices. Test:test_registry.py::test_gate_blocks_delete.
Full test suite runs deterministically via pytest tests/ -v from ~/dablio. Conftest fixtures pre-generate reproducible audio (via say + ffmpeg) and stub the Claude API where needed.
Configuration and Deployment
- config.toml: Centralized settings (model: claude-opus-4-8, voice: Daniel, heartbeat interval: 60s, memory fact limit: 50). Parsed by
config.py. - CLI entry point (
bin/dablio): Symlinked to~/bin/dablio, invoked asdabliofrom any directory. Subcommands:chat(one-turn conversation),voice(space-bar mode),notices(pending approvals),pause/resume(heartbeat control),facts(list memory),audit(last 30 actions),dismiss <id>(approve or dismiss a notice). - Launchd plist (
~/Library/LaunchAgents/com.dablio.heartbeat.plist): Starts on login, runspython /Users/cb/dablio/dablio/heartbeat.pywith output redirected to~/dablio/logs/heartbeat.log. Can be unloaded withlaunchctl unloadto disable the background loop without restarting.
Key Decisions
- Separate workspace — Dablio lives in
~/dablio, independent of ICM or other agent infrastructure. Minimal surface area for interference. - Memory injection over dynamic retrieval — Facts are prepended to the system prompt every turn rather than retrieved on-demand. Trade-off: simpler implementation, but limited by context window (50-fact hard limit). A future version could implement vector similarity search, but this is sufficient for personal use.
- Gate allowlist model — Consequential actions (delete, send, deploy) are denied by default unless explicitly whitelisted in the code. This prevents the model from accidentally destroying data. The allowlist is code, not config, to avoid social engineering. Users can edit
gate.pyto add approved actions. - Audit trail over silent operation — Every action is logged to a human-readable JSON file. This lets the user build trust via
dablio auditand debug unexpected behavior. The log is append-only and never deleted, creating a complete history.
What's Next
Dablio is fully functional and ready for daily use. Possible extensions (not yet implemented):
- Vector embedding of long-term memory for semantic similarity search (memory.py could be extended to use OpenAI embeddings or ONNX).
- Integration with calendar, email, or issue tracking (tools/ package is designed for plugin-style extensibility).
- Offline embeddings using a local vector database (Chroma or Qdrant).
- Real-time transcription feedback (current design waits for space-bar release before transcribing; streaming would be more interactive).
The codebase is deterministic, fully tested, and logged. Start with dablio chat to verify the setup, then dablio voice to begin speaking commands. Morning routine: dablio notices, ls ~/dablio/reports/, dablio audit.