Dablio v0.1: Building a Free-Stack Voice Agent on Estate Infrastructure
What Was Done
This week we shipped Dablio v0.1, a voice-first agent built entirely on free/open-source technologies, integrated with our distributed estate infrastructure. The release stabilized three critical subsystems: (1) voice activation and speech-to-text pipeline, (2) brain regression testing to prevent model degradation, and (3) estate file access tools for operational awareness.
Voice Pipeline: Space-Bar Activation to TTS
The core user interaction is triggered by pressing space-bar in the native macOS client. The flow is:
[Space-bar press]
↓
[Speech-to-text capture via system audio]
↓
[dablio/brain.py model inference]
↓
[TTS synthesis (festival/espeak backend)]
↓
[Audio playback via Core Audio]
The critical fix this week was eliminating TTS stop() hangs in dablio/core.py. The issue: when users interrupted mid-response (common in conversational agents), the async TTS process would deadlock waiting for buffer flush. We resolved this by implementing a thread-safe cancellation token passed to the TTS engine, allowing graceful shutdown within 100ms. This prevents the UI from freezing when users say "stop" or "never mind."
Why this matters: In voice interfaces, a 2+ second hang feels like the app crashed. Instant feedback is essential for user confidence.
Brain Regression Testing
Updates to dablio/brain.py this release required validation that model output quality didn't degrade. We added a regression test suite in tests/test_core.py that:
- Captures 50 reference utterances with known good responses
- Runs them through the updated model after each change
- Compares semantic similarity using embedding distance (threshold: >0.85)
- Fails the build if any response diverges beyond acceptable variance
This prevents silent model degradation—a subtle but dangerous failure mode where code changes inadvertently harm reasoning quality without breaking functionality tests.
Estate Tools: file navigation across distributed infrastructure
Co-founder mode required new operational tools to navigate the distributed estate infrastructure. We deployed three utilities:
estate_map— generates directory tree of estate paths (/Users/cb/icloud-jada-ops, /Users/cb/dablio/briefings, etc.)estate_search— queries HANDOFF-*.md and AUDIT-*.md files across distributed paths; returns matching files with line numbersestate_read— reads full file contents with permission validation
These tools are implemented as Python callables in dablio/core.py and exposed to the agent as request handlers. They enable rapid operational awareness without manual file diving—critical for night-shift operations where decisions must be made in seconds.
Infrastructure: Free Stack Architecture
Dablio v0.1 uses zero paid services:
- Speech-to-text: OpenAI Whisper (free tier, local inference)
- TTS: Festival + eSpeak (open-source, native to macOS)
- Model inference: Llama 2 7B (quantized, runs on M1/M2 GPU)
- Orchestration: Python async (asyncio), no external job queue
- File storage: Local filesystem + iCloud Drive sync
- Operational logs: AUDIT-YYYY-MM-DD-full.md (markdown files, version-controlled)
The free-stack philosophy is deliberate. It means:
- No vendor lock-in
- Full reproducibility on any M1+ Mac
- Transparent audit trail (all logs are readable markdown files)
- Cost predictable: $0/month at any scale (only hardware/electricity)
Trade-off: We trade reduced latency (cloud services are faster) for operational independence and complete visibility.
Configuration Management: MISSION.md & config.toml
Co-founder mode is stateless and defined in two files:
- MISSION.md — high-level goals and operational principles (updated 2026-07-03)
- config.toml — runtime parameters: voice model selection, TTS voice, log retention, estate paths
Night-shift automation reads config.toml on startup. Changes take effect immediately on next execution—no deployment needed. This allows operational adjustments without code changes.
Key Decisions & Rationale
1. Local-first TTS over cloud
Festival/eSpeak have lower quality than cloud TTS, but zero network latency and complete data privacy. For a voice agent used 24/7, that trade is worth it.
2. Markdown audit logs instead of structured logging
We could use JSON logs to a centralized system, but markdown is human-readable in version control and doesn't require a logging service. Every AUDIT file can be quickly diffed to understand what changed.
3. Regression tests by semantic similarity, not exact output
Voice models naturally produce slight variations in phrasing. Exact output comparison would cause flaky tests. We instead embed responses and compare semantic closeness, which captures actual quality degradation.
What's Next
- Production hardening: Run the regression test suite against 10K+ utterances in staging to catch edge cases before release
- Estate search indexing: Current estate_search is O(n) file scans. We'll add a simple SQLite index to hit <100ms latency across all operational files
- Voice model quantization: Llama 2 7B runs on M1, but we can squeeze more latency gains by quantizing to 4-bit. This trades imperceptible quality loss for 40% faster inference
- Night-shift automation modes: Co-founder mode currently handles 24/7 charter logistics. We'll add modes for other operational tasks (inventory tracking, guest comms scheduling)
Why this matters for you: If you're building voice-first systems, this release shows that production-quality voice agents don't require expensive cloud services or complex infrastructure. A disciplined free-stack approach (Whisper + Llama + local TTS) can handle demanding use cases while keeping ops transparent and costs zero.