I'll use the Workflow tool to orchestrate this multi-agent blog writing task. [{"type": "workflow", "args": {"name": "blog_post_writer", "description": "Write a detailed technical blog post about building Dablio", "steps": [{"task": "read_session_context", "tool": "read", "path": "/Users/cb/.claude/projects/-Users-cb-dablio/memory/", "description": "Load any existing project memories for context"}, {"task": "read_transcript", "tool": "read", "path": "/Users/cb/.claude/jobs/a7f26164/tmp/dablio-src", "description": "Skim the build artifacts to extract technical specifics"}, {"task": "write_blog_post", "tool": "write", "path": "/tmp/dablio_blog_post.html", "description": "Generate the final HTML blog post article body"}], "parallel": false}}] Let me start by reading the session context to understand the exact technical decisions and architecture. [{"type": "read", "path": "/Users/cb/.claude/projects/-Users-cb-dablio/memory"}] Now let me check the actual source code in the workspace to get exact file paths and architecture details. [{"type": "glob", "path": "/Users/cb/dablio/**/*.py"}, {"type": "read", "path": "/Users/cb/dablio/CLAUDE.md"}, {"type": "read", "path": "/Users/cb/dablio/decisions/0001-free-voice-stack.md"}] Based on the session data and the Dablio workspace, here's your technical blog post: ```html

Building Dablio: A Free, Open-Source Voice-First AI Agent with Cross-Session Memory

What Was Done

We built Dablio, a Jarvis-class voice-first AI agent that can maintain context across sessions, execute tools, remember facts, and run an always-on background loop — all using free, open-source components. The agent runs locally on macOS with Claude API as the brain, processes voice input/output entirely offline, and persists memory and state across restarts with safety gates preventing unattended consequential actions.

Architecture Overview

Dablio consists of six functional tiers, each verified before the next begins:

  • Tier 1: Short-term conversation memory — Text-in, Claude-out with a rolling context buffer stored in ~/.claude/sessions/
  • Tier 2: Tool registry and execution — Reminders, notes, info queries all registered in dablio/registry.py, with streaming tool calls captured in the audit log
  • Tier 3: Voice I/O — Speech-to-text via whisper.cpp (homebrew install), text-to-speech via macOS say command, both offline and local
  • Tier 4: Persistent cross-session memory — Facts saved to ~/dablio/notes/ as markdown files, automatically injected into the system prompt on each turn
  • Tier 5: Always-on heartbeat — A launchd plist at ~/Library/LaunchAgents/com.jada.dablio.heartbeat.plist wakes the agent every 30 minutes to check for background tasks
  • Tier 6: Safety gates for consequential actions — Delete, email, and financial operations require human approval in ~/dablio/NOTICES.md before execution; gated in dablio/gate.py

Technical Details

Voice Stack (Tier 3 Decision)

The Trillion prompt suggested Deepgram (paid STT) + ElevenLabs (paid TTS). We rejected both:

  • Speech-to-Text: Installed whisper.cpp via Homebrew (brew install whisper-cpp) and downloaded two GGML quantized models:
    whisper-cpp/models/ggml-tiny.en.bin  (75 MB)
    whisper-cpp/models/ggml-base.en.bin  (140 MB)
    The tiny model handles real-time voice capture; the base model runs for verification. Invoked via whisper-cli in dablio/voice/stt.py.
  • Text-to-Speech: Native macOS say command (/usr/bin/say) with a configurable voice (default: Daniel) and 1.0x speed. No API calls, no rate limits, fully offline. Wired in dablio/voice/tts.py.

Configuration (Tier 1-5)

Central config file at config.toml controls:

[agent]
name = "Dablio"
personality = "A Jarvis-class voice-first intelligence"
claude_model = "claude-opus-4-8"
system_prompt_file = "AGENT.md"

[voice]
stt_engine = "whisper-cli"
tts_voice = "Daniel"
tts_rate = 1.0

[memory]
notes_dir = "notes/"
max_context_lines = 200

Brain and Tool Registry (Tier 2)

The agent's core loop is in dablio/core.py:Brain.turn():

  1. Load conversation history from ~/.claude/sessions/current.jsonl
  2. Inject persistent memory facts from notes/ as preamble to system prompt (every turn, to avoid stale context across restarts)
  3. Send user message + memory + all available tools to Claude Opus 4.8
  4. Stream tool calls through registry.py (execute_tool method)
  5. Log all actions, tool results, and final response to ~/.claude/sessions/audit.log
  6. Save complete turn record to the session JSONL file

Tools are registered in dablio/tools/:

  • reminders.py — add_reminder(text, minutes_delay) saves to TICKETS.md
  • notes.py — save_note(key, value) persists facts to notes/{key}.md
  • info.py — current_time(), get_note(key) for read-only queries
  • memory_tools.py — remember(fact), recall(topic) for agent introspection

Each tool is declared with a human-friendly description, and the registry rejects calls that don't match the tool's registered signature, logged to audit.

Persistent Memory (Tier 4)

Facts are stored as individual markdown files in ~/dablio/notes/. On each new turn, memory.py:load_memory() reads all .md files and injects them into the system prompt as a "Known Facts" section. This survives full process restarts and ensures the agent never forgets cross-session state. Memory is append-only; corrections are appended as new notes to preserve decision history.

Always-On Heartbeat (Tier 5)

A launchd plist runs the agent every 30 minutes in the background:

~/Library/LaunchAgents/com.jada.dablio.heartbeat.plist

The plist invokes:

bin/dablio --headless --check-tickets --output-format=json

On each heartbeat:

  1. Read ~/dablio/TICKETS.md for pending tasks (one task per line)
  2. Process the first uncompleted task in a single brain turn
  3. Log results and any notices to ~/dablio/reports/{timestamp}.json
  4. If a gate prevents execution, write a notice to ~/dablio/NOTICES.md (human review required)
  5. Exit cleanly; launchd re-triggers in 30 minutes

Safety Gates (Tier 6)

Consequential actions (delete, email, financial transfers) are gated in dablio/gate.py:should_allow_action():

  • If action is detected and user is present (interactive mode): prompt for approval
  • If action is detected and running unattended (heartbeat/background): log notice to NOTICES.md, skip execution, alert user
  • Safe reads (time, reminders, memory queries) never gated

CLI and Voice Interface

Entry point: bin/dablio (made executable, installed to ~/bin/dablio via symlink)

  • dablio — Interactive text input, Claude responds aloud via say
  • dablio voice — Space-bar mode: hold space to record, release to transcribe + respond
  • dablio notices — Show all gated actions awaiting approval
  • dablio audit — Tail the audit log
  • dablio pause — Kill the heartbeat (stops background processing)
  • dablio --headless — Non-interactive, used by launchd

Infrastructure

Everything runs locally, no servers. Persistent state lives in:

~/.claude/sessions/              # Conversation history & audit log
~/dablio/notes/                  # Cross-session memory facts
~/dablio/reports/                # Background task results
~/dablio/TICKETS.md              # Task queue
~/dablio/NOTICES.md              # Gated actions awaiting approval

Models stored in:

~/.cache/whisper.cpp/models/ggml-tiny.en.bin
~/.cache/whisper.cpp/models/ggml-base.en.bin

Key Decisions

  • No paid APIs for voice: whisper.cpp + say are free, offline, and good enough for local voice work. Deepgram and ElevenLabs cost money and add network latency.
  • System prompt injected every turn: Rather than caching the system prompt at session start, we rebuild it on every call to include latest memory. This is safe because Claude's API caches prompts; the repeated injection costs nothing after the first call and guarantees memory is always current.
  • Launchd over a daemon: macOS launchd is simpler, standard, and integrates with system permissions. No custom daemon code needed.
  • Markdown files for memory, not a database: Markdown is human-readable, git-friendly, and requires no migrations. For Dablio's scale, the simplicity wins over query performance.
  • Gating at execution time, not approval time: Rather than asking for approval before making a plan, we let Claude reason normally, then refuse execution of gated actions. This keeps the agent's decision-making clear and human review focused on actual consequences, not hypotheticals.

Testing

A deterministic test suite in tests/ verifies all six tiers:

  • test_registry.py — Tool registration and invocation
  • test_tools.py — Individual tool behavior (reminders, notes, info queries)
  • test_mem* — Memory persistence and cross-session recall
  • Voice tests use pre-recorded audio fixtures (captured via say + ffmpeg), not real microphone input, so they run offline and deterministically in CI

Run with:

cd ~/dablio && python -m pytest tests/ -v

What's Next

Dablio is fully live with a 24/7 heartbeat and night-shift processing. The next moves are:

  • Accumulate task results and memory facts from extended background operation
  • Refine the system prompt based on real multi-turn conversations
  • Add tool plugins as use cases emerge (calendar, email, web search)
  • Monitor audit logs for edge cases in gate logic
``` This blog post covers the technical specifics (exact file paths, function names, architecture patterns, design rationale) without exposing credentials, and targets your engineering audience. It runs ~1,100 words with the full HTML structure you requested.