Building Dablio: A Free, Open-Source Voice-First AI Agent with Cross-Session Memory
What Was Done
We built Dablio, a Jarvis-class voice-first AI agent that can maintain context across sessions, execute tools, remember facts, and run an always-on background loop — all using free, open-source components. The agent runs locally on macOS with Claude API as the brain, processes voice input/output entirely offline, and persists memory and state across restarts with safety gates preventing unattended consequential actions.
Architecture Overview
Dablio consists of six functional tiers, each verified before the next begins:
- Tier 1: Short-term conversation memory — Text-in, Claude-out with a rolling context buffer stored in
~/.claude/sessions/ - Tier 2: Tool registry and execution — Reminders, notes, info queries all registered in
dablio/registry.py, with streaming tool calls captured in the audit log - Tier 3: Voice I/O — Speech-to-text via
whisper.cpp(homebrew install), text-to-speech via macOSsaycommand, both offline and local - Tier 4: Persistent cross-session memory — Facts saved to
~/dablio/notes/as markdown files, automatically injected into the system prompt on each turn - Tier 5: Always-on heartbeat — A launchd plist at
~/Library/LaunchAgents/com.jada.dablio.heartbeat.plistwakes the agent every 30 minutes to check for background tasks - Tier 6: Safety gates for consequential actions — Delete, email, and financial operations require human approval in
~/dablio/NOTICES.mdbefore execution; gated indablio/gate.py
Technical Details
Voice Stack (Tier 3 Decision)
The Trillion prompt suggested Deepgram (paid STT) + ElevenLabs (paid TTS). We rejected both:
- Speech-to-Text: Installed
whisper.cppvia Homebrew (brew install whisper-cpp) and downloaded two GGML quantized models:
The tiny model handles real-time voice capture; the base model runs for verification. Invoked viawhisper-cpp/models/ggml-tiny.en.bin (75 MB) whisper-cpp/models/ggml-base.en.bin (140 MB)whisper-cliindablio/voice/stt.py. - Text-to-Speech: Native macOS
saycommand (/usr/bin/say) with a configurable voice (default: Daniel) and 1.0x speed. No API calls, no rate limits, fully offline. Wired indablio/voice/tts.py.
Configuration (Tier 1-5)
Central config file at config.toml controls:
[agent]
name = "Dablio"
personality = "A Jarvis-class voice-first intelligence"
claude_model = "claude-opus-4-8"
system_prompt_file = "AGENT.md"
[voice]
stt_engine = "whisper-cli"
tts_voice = "Daniel"
tts_rate = 1.0
[memory]
notes_dir = "notes/"
max_context_lines = 200
Brain and Tool Registry (Tier 2)
The agent's core loop is in dablio/core.py:Brain.turn():
- Load conversation history from
~/.claude/sessions/current.jsonl - Inject persistent memory facts from
notes/as preamble to system prompt (every turn, to avoid stale context across restarts) - Send user message + memory + all available tools to Claude Opus 4.8
- Stream tool calls through
registry.py(execute_tool method) - Log all actions, tool results, and final response to
~/.claude/sessions/audit.log - Save complete turn record to the session JSONL file
Tools are registered in dablio/tools/:
reminders.py—add_reminder(text, minutes_delay)saves toTICKETS.mdnotes.py—save_note(key, value)persists facts tonotes/{key}.mdinfo.py—current_time(),get_note(key)for read-only queriesmemory_tools.py—remember(fact),recall(topic)for agent introspection
Each tool is declared with a human-friendly description, and the registry rejects calls that don't match the tool's registered signature, logged to audit.
Persistent Memory (Tier 4)
Facts are stored as individual markdown files in ~/dablio/notes/. On each new turn, memory.py:load_memory() reads all .md files and injects them into the system prompt as a "Known Facts" section. This survives full process restarts and ensures the agent never forgets cross-session state. Memory is append-only; corrections are appended as new notes to preserve decision history.
Always-On Heartbeat (Tier 5)
A launchd plist runs the agent every 30 minutes in the background:
~/Library/LaunchAgents/com.jada.dablio.heartbeat.plist
The plist invokes:
bin/dablio --headless --check-tickets --output-format=json
On each heartbeat:
- Read
~/dablio/TICKETS.mdfor pending tasks (one task per line) - Process the first uncompleted task in a single brain turn
- Log results and any notices to
~/dablio/reports/{timestamp}.json - If a gate prevents execution, write a notice to
~/dablio/NOTICES.md(human review required) - Exit cleanly; launchd re-triggers in 30 minutes
Safety Gates (Tier 6)
Consequential actions (delete, email, financial transfers) are gated in dablio/gate.py:should_allow_action():
- If action is detected and user is present (interactive mode): prompt for approval
- If action is detected and running unattended (heartbeat/background): log notice to
NOTICES.md, skip execution, alert user - Safe reads (time, reminders, memory queries) never gated
CLI and Voice Interface
Entry point: bin/dablio (made executable, installed to ~/bin/dablio via symlink)
dablio— Interactive text input, Claude responds aloud viasaydablio voice— Space-bar mode: hold space to record, release to transcribe + responddablio notices— Show all gated actions awaiting approvaldablio audit— Tail the audit logdablio pause— Kill the heartbeat (stops background processing)dablio --headless— Non-interactive, used by launchd
Infrastructure
Everything runs locally, no servers. Persistent state lives in:
~/.claude/sessions/ # Conversation history & audit log
~/dablio/notes/ # Cross-session memory facts
~/dablio/reports/ # Background task results
~/dablio/TICKETS.md # Task queue
~/dablio/NOTICES.md # Gated actions awaiting approval
Models stored in:
~/.cache/whisper.cpp/models/ggml-tiny.en.bin
~/.cache/whisper.cpp/models/ggml-base.en.bin
Key Decisions
- No paid APIs for voice:
whisper.cpp+sayare free, offline, and good enough for local voice work. Deepgram and ElevenLabs cost money and add network latency. - System prompt injected every turn: Rather than caching the system prompt at session start, we rebuild it on every call to include latest memory. This is safe because Claude's API caches prompts; the repeated injection costs nothing after the first call and guarantees memory is always current.
- Launchd over a daemon: macOS launchd is simpler, standard, and integrates with system permissions. No custom daemon code needed.
- Markdown files for memory, not a database: Markdown is human-readable, git-friendly, and requires no migrations. For Dablio's scale, the simplicity wins over query performance.
- Gating at execution time, not approval time: Rather than asking for approval before making a plan, we let Claude reason normally, then refuse execution of gated actions. This keeps the agent's decision-making clear and human review focused on actual consequences, not hypotheticals.
Testing
A deterministic test suite in tests/ verifies all six tiers:
test_registry.py— Tool registration and invocationtest_tools.py— Individual tool behavior (reminders, notes, info queries)test_mem*— Memory persistence and cross-session recall- Voice tests use pre-recorded audio fixtures (captured via
say + ffmpeg), not real microphone input, so they run offline and deterministically in CI
Run with:
cd ~/dablio && python -m pytest tests/ -v
What's Next
Dablio is fully live with a 24/7 heartbeat and night-shift processing. The next moves are:
- Accumulate task results and memory facts from extended background operation
- Refine the system prompt based on real multi-turn conversations
- Add tool plugins as use cases emerge (calendar, email, web search)
- Monitor audit logs for edge cases in gate logic