Building Dablio: A Free, Open-Source Voice Agent with Cross-Session Memory and Proactive Autonomy
What Was Done
We extracted a proprietary voice-agent prompt from hellotrillion.ai and used it as a blueprint to build Dablio, a fully-functional Jarvis-styled voice assistant. Dablio runs locally with persistent memory, tool integrations, and a safety gate that prevents the agent from acting autonomously without explicit user approval. The entire system was built using open-source dependencies (faster-whisper, numpy, sounddevice) instead of the paid services (Deepgram STT, ElevenLabs TTS) originally specified in the Trillion prompt.
Architecture Overview
Dablio is a six-tier voice agent modeled after the Trillion framework, with each tier building on the previous:
- Tier 1: Text conversation loop with short-term memory (claude.py agent loop)
- Tier 2: Tool registry that allows Claude to invoke utilities (registry.py, 11 integrated tools)
- Tier 3: Streaming speech-in/speech-out via local hardware (voice/capture.py, voice/stt.py, voice/tts.py)
- Tier 4: Persistent long-term memory with session-aware recall (memory.py, brain.py)
- Tier 5: Always-on proactive background loop (heartbeat.py) that monitors context and suggests actions
- Tier 6: Safety gate (gate.py) that gates autonomous agent actions and requires user approval
Workspace Structure and File Organization
The project lives in a dedicated workspace at ~/dablio/, isolated from existing systems (ICM, JADA). The directory structure follows this layout:
~/dablio/
├── .venv/ # Python virtual environment
├── dablio/ # Main package
│ ├── __init__.py
│ ├── core.py # Main agent loop and orchestration
│ ├── config.py # Configuration parser (reads config.toml)
│ ├── env.py # Environment variable handling
│ ├── brain.py # Claude API integration with streaming
│ ├── registry.py # Tool registry and dispatcher
│ ├── memory.py # Session-aware memory management
│ ├── gate.py # Approval gate for autonomous actions
│ ├── audit.py # Logging and audit trail
│ ├── heartbeat.py # Proactive background loop
│ ├── voice/
│ │ ├── __init__.py
│ │ ├── capture.py # Audio input via sounddevice
│ │ ├── stt.py # Speech-to-text (faster-whisper)
│ │ └── tts.py # Text-to-speech (macOS say command)
│ └── tools/
│ ├── __init__.py
│ ├── reminders.py # Reminder creation/management
│ ├── notes.py # Note taking system
│ ├── info.py # Information retrieval tools
│ └── memory_tools.py # Memory introspection tools
├── bin/
│ ├── dablio # CLI launcher (executable)
│ └── setup.sh # One-time initialization script
├── tests/
│ ├── conftest.py # Pytest fixtures
│ ├── test_registry.py
│ ├── test_tools.py
│ ├── test_memory.py
│ ├── test_voice.py
│ ├── test_gate.py
│ ├── test_audit.py
│ └── test_cli.py
├── config.toml # Runtime configuration
├── CLAUDE.md # AI collaboration rules
├── CONTEXT.md # Project context for AI agents
├── AGENT.md # Agent behavior spec
├── decisions/
│ └── 0001-free-voice-stack.md # Architecture decision record
└── README.md
Key Technical Decisions
STT and TTS: Free Alternatives to Paid Services
The original Trillion prompt specified Deepgram for speech-to-text and ElevenLabs for text-to-speech. Both require paid API keys. We chose:
- Speech-to-Text: faster-whisper (OpenAI Whisper via CTranslate2 backend). Installed via
pip install faster-whisper. Runs locally without API calls. The model is downloaded on first use. Invoice/stt.py, thetranscribe_audio()function loads the model once per session and transcribes audio files synchronously. - Text-to-Speech: macOS native
saycommand for development; deployable to any platform with local TTS. Invoice/tts.py, thespeak()function spawns the systemsaycommand asynchronously, with configurable voice and rate parameters read fromconfig.toml. - Audio Capture: sounddevice (pip-installable). In
voice/capture.py, therecord_audio()function captures raw audio viasd.rec(), saves it as WAV to the session temp directory, and returns the file path.
This decision was documented in decisions/0001-free-voice-stack.md to justify the tradeoff: no API cost, no rate limits, no external dependencies on service availability, but limited voice quality and no multi-voice synthesis.
Memory System: Session-Aware Persistence
The memory.py module implements two-tier memory:
- Short-term memory: In-process list of (role, content) tuples representing the conversation history. Passed to Claude on each request in the messages array.
- Long-term memory: Persisted to JSON files in
~/.dablio/memory/. Each session gets asession_{session_id}.jsonfile; Dablio also maintains aprofile.jsonwith user preferences and learned facts. - Session awareness: At startup,
memory.load_session_context()reads the session file and pre-populates short-term memory. Thememory_tools.pymodule exposes Claude functions to query long-term memory (recall_fact(),list_sessions()) and introspect the current session.
Files are stored in a flat directory structure with ISO 8601 timestamps in filenames. No database is required; JSON parsing handles serialization.
Tool Registry: Declarative Tool Binding
The registry.py module maintains a registry of available tools. Each tool is a Python function decorated with @register_tool, which records its signature, docstring, and parameter schema. In core.py, when Claude returns a tool use request, the agent calls registry.dispatch(tool_name, kwargs), which looks up the function, validates arguments against the schema, and invokes it. This decouples tool definitions from the agent loop.
All 11 tools are registered in this pattern:
tools/reminders.py:create_reminder(),list_reminders(),dismiss_reminder()tools/notes.py:save_note(),search_notes(),list_recent_notes()tools/info.py:get_time(),get_weather()tools/memory_tools.py:recall_fact(),list_sessions()
Gate: Approval for Autonomous Actions
The gate.py module implements Tier 6 safety. When the heartbeat loop (Tier 5) generates a suggested action, it must pass through gate.approve_action(action, context). The gate prompts the user with the proposed action, context, and expected outcome, then awaits y/n input before proceeding. This prevents the agent from taking unexpected actions without consent.
Heartbeat: Proactive Monitoring
The heartbeat.py module runs a background loop that wakes every N seconds (configurable in config.toml, default 60s). It evaluates current context (time, reminders, session duration) and can propose actions to Claude via a separate agent invocation. Examples: "It's been 2 hours; suggest a break" or "You have a 3pm reminder in 10 minutes." All proposals flow through the gate.
Configuration and Environment
Configuration lives in config.toml and is parsed by config.py using the tomli library. Environment variables are handled by env.py. The agent reads:
[agent]
name = "Dablio"
personality = "Jarvis-like assistant"
[voice]
capture_device = "default"
tts_voice = "Victoria"
tts_rate = 1.0
[memory]
session_dir = "~/.dablio/memory"
[heartbeat]
interval_seconds = 60
enabled = true
API keys (Claude, any external services) are read from environment variables only, never committed to the repo. The .gitignore blocks .env, .venv/, and ~/.dablio/memory/.
CLI and Launcher
The cli.py module defines argument parsing. The bin/dablio launcher is a shell wrapper that activates the venv and invokes the CLI. Usage:
~/dablio/bin/dablio --voice # Voice mode (default)
~/dablio/bin/dablio --text # Text-only mode
~/dablio/bin/dablio --session # Resume specific session
~/dablio/bin/dablio --background # Start heartbeat only
Testing and Verification
Tier 6 includes a deterministic test suite. Key fixtures in conftest.py create isolated test sessions, mock the Claude API, and pre-generate audio fixtures (captured once via say "hello" | ffmpeg). Tests are grouped by tier:
test_registry.py:Tool registration, dispatch, and schema validationtest_tools.py:Individual tool correctness (note save/recall, reminder creation)test_memory.py:Session persistence, fact recall, context loadingtest_voice.py:Audio capture, STT transcription, TTS synthesistest_gate.py:Approval gate logic and user promptingtest_audit.py:Audit trail loggingtest_cli.py:Argument parsing and launcher behavior
Run all tests via ~/.venv/bin/pytest tests/ -v. Each test is isolated with a temporary session ID and temp directory.
Dependencies and Environment Setup
The venv is initialized at ~/dablio/.venv/ with Python 3.11+. Core dependencies:
numpy— Numerical operations (audio processing)sounddevice— Audio input via ALSA/CoreAudio/WASAPIfaster-whisper— Local STT (includes CTranslate2)pytest— Test runnertomli— TOML parsing (backport for Python <3.11)anthropic— Claude API client (already in system Python via Claude Code)
Installation command (in setup.sh):
python -m venv ~/.venv
~/.venv/bin/pip install --upgrade pip
~/.venv/bin/pip install numpy sounddevice pytest faster-whisper tomli
What's Next
The codebase is complete and passes unit tests. Next steps for operators:
- Run
~/dablio/bin/setup.shto initialize the workspace on a fresh machine - Set
CLAUDE_API_KEYenvironment variable - Launch voice mode:
~/dablio/bin/dablio --voice - Optionally extend tools via
tools/custom.py(auto-registered if decorated) - Monitor audit logs in
~/.dablio/audit/for session history - Export memories via
memory_tools.list_sessions()for archival or analysis
The heartbeat loop can be deployed as a systemd timer or launchd daemon for always-on proactive monitoring. The gate ensures human oversight of autonomous actions.
```