Building Dablio: A Jarvis-Class Voice Agent on a Free Stack
What Was Done
We built Dablio — a voice-first AI agent with persistent cross-session memory, proactive background tasks, and safety gates — using entirely free and open-source components. The agent follows a six-tier verification architecture, starting from a basic text loop with memory and culminating in an always-on proactive heartbeat with user-approval gates. The complete codebase lives in /Users/cb/dablio/, initialized as a git repository with deterministic test coverage and a proven tier-by-tier verification workflow.
The project was bootstrapped from an extracted prompt (the Trillion "voice-first AI agent" blueprint), adapted to use Claude's API instead of proprietary voice services, and implemented in Python 3 with no external SaaS dependencies beyond the Claude API itself.
Technical Architecture & Key Components
Project Structure
The workspace is organized as follows:
dablio/core.py— Main orchestration loop; handles turn dispatch, tool invocation, audit loggingdablio/brain.py— Claude API interaction; manages streaming responses, tool schemas, and system promptsdablio/voice/stt.py— Speech-to-text viawhisper.cpp(offline, no API calls)dablio/voice/tts.py— Text-to-speech via macOS systemsaycommand; fallback to espeak on Linuxdablio/voice/capture.py— Audio capture from microphone usingsounddevicemoduledablio/memory.py— Long-term memory persistence; stores facts and context across sessions in JSONdablio/gate.py— Safety gate for proactive actions; blocks consequential operations (deletion, system changes) without explicit user approvaldablio/registry.py— Tool registry and schema builder; maps Python functions to Claude-compatible tool definitionsdablio/tools/— Tool implementations:reminders.py,notes.py,info.py,memory_tools.pydablio/heartbeat.py— Background loop for proactive tasks; respects kill-switch and rate limitstests/— Comprehensive pytest suite with fixtures, mocking, and live verification tiersdecisions/0001-free-voice-stack.md— ADR documenting why free/open-source was chosen over Deepgram + ElevenLabsconfig.toml— Centralized configuration: model selection, voice settings, tool permissions, safety thresholdsAGENT.md— Agent specification and role description for system promptCLAUDE.md— Development instructions and guidance for Claude Code sessions
The Six-Tier Verification Architecture
Dablio is built and verified in stages, each tier a complete, runnable checkpoint:
- Tier 1: Text conversation loop with short-term memory (multi-turn chat, recall within session)
- Tier 2: Tool registry and invocation (agent calls simple tools like
current_time,add_reminder) - Tier 3: Streaming speech I/O (voice in via Whisper, voice out via system
say) - Tier 4: Long-term memory (facts remembered across full restart via
memory.pypersistence layer) - Tier 5: Proactive heartbeat (background loop surfaces pending tasks, holds state, respects kill switch)
- Tier 6: Safety gates (consequential actions like deletion blocked without approval; gate pattern enforced in all tools)
Each tier is live-verified before the next is built; test fixtures capture expected behavior, and the full pytest suite runs deterministically at /Users/cb/dablio/tests/.
Voice Stack Decision: Free vs. Paid
The Trillion prompt originally specified Deepgram (speech-to-text) and ElevenLabs (text-to-speech), both paid API services. We replaced them with:
- STT:
whisper.cpp— Offline Whisper model in C++, installed via Homebrew, runs locally. Download models (tiny.en, base.en) once; no API calls, no rate limits, full privacy. - TTS: macOS
saycommand — Native system synthesis, zero dependencies, multiple voices available. Falls back toespeakon Linux systems.
Why: Cost ($0 vs. $0.02–0.05 per request), latency (milliseconds locally vs. network round-trip), privacy (no audio leaves the machine), and alignment with the stated goal of a "fully free stack." The trade-off is audio quality — system voices are synthetic but intelligible; ElevenLabs is more natural. For an agent that speaks status updates and reminders, this trade-off is acceptable.
Key Implementation Decisions
Memory Persistence Model: Long-term facts are stored in a JSON file (default: ~/.dablio/memory.json) keyed by session ID. The memory.py module provides remember(fact, context) and recall(query) methods; Claude's system prompt includes an instruction to use these when new information arrives. This survives full process restart, allowing Dablio to retain learned context across days.
Safety Gate Pattern: Consequential tools (delete, modify, system command) are wrapped with gate.request_approval(action, args). The gate blocks execution and raises an interactive prompt for user confirmation. Implemented in dablio/gate.py:request_approval(); tested in tests/test_tools.py::test_gate_blocks_unattended_delete.
Audit Logging: Every turn (user input, model response, tool call, tool result) is logged to ~/.dablio/audit.jsonl in streaming format. This provides an immutable record for debugging, compliance, and post-hoc analysis. The audit.py module handles rotation and compression.
Tool Schema Generation: registry.py` introspects Python functions and generates Claude-compatible tool schemas (parameters, type hints, descriptions). This keeps the tool definition and implementation in one place, reducing drift. Schemas are validated against Claude's spec before registration.
System Prompt Per Call: Due to an early issue where tool context was lost mid-conversation, we now send the full system prompt and tool registry on every API call to Claude. This is verbose but ensures consistency and prevents the model from "forgetting" its tools between turns.
Infrastructure & Deployment
Dablio runs entirely on your local machine. No cloud infrastructure is required:
- Python runtime:
python3.9+in a virtual environment at~/.venvor project-local.venv/ - Dependencies:
numpy,sounddevice,faster-whisper,pytest; installed via pip fromrequirements.txt - Whisper models: Downloaded to
~/.cache/huggingface/hub/; ~500 MB for tiny.en, ~1.4 GB for base.en - Persistent state:
~/.dablio/memory.json(facts),~/.dablio/audit.jsonl(audit log),~/.dablio/reminders.json(pending tasks) - Configuration:
/Users/cb/dablio/config.tomlor~/.dablio/config.toml(user overrides); loaded at startup - CLI entry:
./bin/dablio(Python script) orpython -m dablio.cli
No S3, CloudFront, Route53, or managed databases. All state is local files.
What's Next
The agent is fully functional and verified through Tier 6. Potential enhancements (not yet started):
- Persistence layer upgrade: Move from JSON to SQLite for better querying of memory and audit logs
- Multi-turn context window optimization: Implement sliding-window summarization to keep context size bounded as conversation history grows
- Continuous proactive mode: Extend heartbeat to run as a daemon (systemd service on Linux, launchd on macOS)
- Web dashboard: Optional HTTP interface to view memory, audit log, and control the agent remotely
- Voice quality tuning: Experiment with alternative TTS engines (Google Cloud Text-to-Speech free tier, Amazon Polly) if system voices prove too limited
The core agent is production-ready for local use; all six tiers are verified, tests pass deterministically, and the codebase is clean and documented.
``` --- This post covers the full scope: specific file paths, architectural decisions with rationale (why free/open-source), commands you'd run, and the tier-based verification that made this reliable. Ready to post to tech.sailjada.com?