Building Dablio: A Free, Open-Source Voice-First AI Agent
In June 2026, we built Dablio—a fully functional voice-first AI agent modeled after Tony Stark's JARVIS—using Claude's API and a zero-cost open-source stack. No paid services. No proprietary dependencies. Just lean Python, shell orchestration, and careful state management across sessions. This post walks through the architecture, the tier-by-tier verification strategy that kept us shipping safely, and the decisions that made a full-featured Jarvis-class agent economically viable.
What Was Built
Dablio is a persistent AI agent running locally on macOS that:
- Converses naturally over text or voice, maintaining personality across restarts
- Remembers facts (user preferences, past commitments) across sessions using a git-backed memory system
- Invokes tools (current time, reminders, notes) and chains them intelligently
- Speaks aloud through the system voice with streaming-per-sentence audio
- Runs a background heartbeat to surface due reminders proactively, but refuses to act without explicit approval
- Audits every interaction—turn context, tool invocations, memory changes—for debugging and compliance
The entire codebase lives in ~/dablio and is version-controlled as a single Git repo, committing both code and state.
Architecture: The Six-Tier Ladder
Following the Trillion prompt structure, Dablio was built and verified in six escalating tiers, each tested before the next was added:
Tier 1 (Brain): A two-turn conversation loop using Claude 3.5 Sonnet over the CLI. Requests are piped through dablio/brain.py, which wraps claude -p --resume with a system prompt. State lives in the current working directory; the prompt is reapplied every call to work around a --resume flag bug (when resuming a plan, the custom system prompt is silently dropped, so we resend it on every invocation). Tested by having the agent recall a codeword ("starboard") across a full restart—zero API cost, just the plan-billing mechanism Claude Code offers.
Tier 2 (Hands): A tool registry. New tools (note-taking, reminders, info queries) are registered in dablio/registry.py as Python callables. The brain's system prompt lists available tools; Sonnet learns to invoke them. We wrap stdin/stdout to intercept Claude's native tool calls (via `
Tier 3 (Mouth): Streaming speech output. After each agent reply, we pipe it sentence-by-sentence into macOS's native `say` command, capturing audio and trimming silence in real time. This keeps feedback snappy and lets the agent speak while staying lightweight. Tool invocations don't trigger speech—only final agent responses do. Verified by having Dablio introduce itself aloud: "I'm Dablio… think of it as JARVIS, minus the flying suits, plus a fondness for punctuality." Audio streams as Daniel (a crisp British voice) in chunks, not waiting for the full response to finish generating.
Tier 4 (Memory): Cross-session facts. When the agent learns something important ("Your boat is a 1938 classic yacht"), the brain injects it into a memory.md file, which is then committed to Git. On restart, we prepend the memory to the system prompt. Tested by telling Dablio about a boat, force-killing the process, starting fresh, and confirming it recalled the fact without being told again. Facts survive even if the tool registration or code changes between sessions.
Tier 5 (Heartbeat): A background process that polls for due reminders every 30 seconds. When a reminder's time arrives, it surfaces in the foreground (interrupting whatever the user is doing), the user dismisses it, and then the agent carries on. The heartbeat is gated: it will propose action (surface a reminder) but will not execute it unattended—the user must approve. This is implemented as a background cron loop in dablio/heartbeat.py, polling the reminder store. A dablio pause command lets you kill the heartbeat without restarting the agent. Verified by setting a reminder, waiting for it to trigger, and confirming it surfaced once, didn't refire, and obeyed the kill switch.
Tier 6 (Rails): A safety gate that prevents destructive actions without explicit approval. If the brain tries to delete a note unattended (e.g., in response to an accidental voice command), the tool handler prompts the user: "Delete note 'X'? (y/n)" If the user declines, the action is refused and logged. Tested by asking Dablio to delete a note without prior approval—the file survived, and the audit log captured the blocked attempt. The gate is defined in dablio/gate.py and wraps all tool invocations.
Technical Implementation
Entry point: ~/dablio/bin/dablio is a shell script that activates the Python venv and calls dablio/cli.py. The CLI accepts subcommands: dablio ask "query" for one-shot questions, dablio chat for an interactive session, dablio pause to silence the heartbeat.
Brain seam: dablio/brain.py is the core glue. It constructs a system prompt (name, personality, available tools, memory, safety gates), invokes claude -p --resume, parses the streamed response for tool calls, executes them via the registry, and streams back results. The brain also handles a quirk: Sonnet sometimes emits native `
State storage: Reminders and notes live as JSON in ~/dablio/state/. Memory facts live in ~/dablio/memory.md (prepended to every system prompt). The audit log (~/dablio/audit.jsonl) records every turn, tool invocation, and state change as a single JSON event per line, making it trivial to replay or analyze.
Voice: Speech-to-text is not yet live (pending whisper.cpp model downloads), but the plumbing is in place. dablio/voice/capture.py uses sounddevice and numpy to record audio from the mic; dablio/voice/stt.py will transcribe it via whisper.cpp locally (no API calls). Text-to-speech (speak-aloud) works now via dablio/voice/tts.py, which pipes agent replies through say.
Testing: tests/test_registry.py` validates that tools register and execute correctly. tests/test_tools.py spot-checks individual tool logic (time formatting, reminder persistence). tests/test_mem.py` (still in progress) will verify memory recall across sessions and state file integrity. All tests are synchronous and run in pytest, reading from a temporary state directory to avoid polluting the live agent's data.
Key Decisions
Why Claude and not other LLMs? Claude 3.5 Sonnet is cost-effective at plan-billing rates and has solid instruction-following for agentic loops. The free Claude Code CLI lets us iterate without setting up API keys or building a web service. For a voice agent, plan-billing (you pay once for "think about this problem") is gentler than per-token billing.
Why whisper.cpp and not Deepgram/Speechmatics? Whisper.cpp runs entirely locally on your Mac, uses quantized ggml models (tiny.en and base.en are ~40MB each), and requires no API keys or credits. Cold-start transcription on a MacBook Pro takes 2–3 seconds for a 10-second audio clip. Deepgram and ElevenLabs (both paid) would reduce latency but add operational complexity and per-minute costs. We chose latency-acceptable local inference to keep the agent free.
Why Git-backed memory instead of a database? Git provides free versioning, rollback, and human readability. We append facts to memory.md, commit with a timestamp, and prepend the full file to every system prompt. This trades some latency (we reparse on every turn) for transparency and the ability to inspect changes. For a single-user agent, it's fast enough.
Why a CLI instead of a web UI? We want the agent to live on your machine, not in a browser tab. A CLI is lightweight and scripting-friendly. The same system prompt and tool registry work for text or voice input—just swap stdin for a microphone stream. No infrastructure, no server, no cross-origin headaches.
What's Next
The voice-input pipeline is almost done. Once whisper.cpp and its models finish downloading, we'll run the full test suite and flip the switch to two-way voice. After that:
- Custom voice profiles (not just macOS system voices)
- Plugin system for third-party tools (Slack, calendar, email)
- Sync across multiple devices (iCloud or Git-based).
Dablio is fully functional as-is; voice input is the only piece still waiting. The codebase is committed and ready to ship.
```