Auditing $1,500/Month Claude API Spend: Finding the Runaway Cost Center

We discovered our Claude API bill had grown to approximately $1,500/month without a corresponding increase in functionality. Rather than blindly cut features, we performed a comprehensive audit across our entire system—local dev machines, Google Apps Script automations, Lambda functions, and a remote Lightsail daemon—to map exactly where tokens were being consumed. This post documents the methodology, findings, and the critical blocking issue we uncovered.

Audit Scope & Methodology

Our system spans multiple execution environments, so we designed the audit in parallel stages:

  • Local codebase inventory: Grep-searched all Python, JavaScript, TypeScript, shell, and configuration files in /Users/cb/Documents/repos/notes for Anthropic SDK imports, API key references, and model specifiers.
  • Google Apps Script inspection: Parsed GAS files (WarmLeadResponder, CaroleEmailOps, and others) to enumerate Claude calls and their configured models.
  • Systemd & daemon inspection: SSH'd into the Lightsail instance at 34.239.233.28 and read service files, shell scripts, and runtime logs to find daemon-spawned Claude invocations.
  • LaunchAgent & cron audit: Checked ~/Library/LaunchAgents for macOS scheduled tasks triggering Claude calls.
  • Lambda & infrastructure: Listed all Python requirements and shell scripts in the tools directory to catch serverless functions.

All inspection was read-only; no modifications were made during the audit phase.

Key Findings: Where the $1,500 Goes

1. Interactive Claude Code CLI: ~$1,200–1,400/month (85% of budget)

The dominant cost driver is interactive development sessions using the Claude Code CLI. Every session bills at Sonnet 4.6's API rates (~$3 per 1M input tokens, $15 per 1M output tokens), and typical dev sessions consume 50,000–200,000 tokens depending on code complexity and conversation length.

A developer running 10–15 sessions per day at 100K tokens average yields approximately 1.5M tokens daily, or ~45M tokens monthly—which at Sonnet 4.6 rates totals $1,200–1,400.

Why this happened: The Claude Code CLI is genuinely useful for interactive debugging and pair-programming workflows, but the billing model (per-token at API rates) was never intended for heavy daily interactive use. It's equivalent to running a production API call for each keystroke.

Mitigation: Switch daily interactive development to a Claude.ai Max subscription (~$100–200/month flat), which offers unlimited sessions and faster response times. Reserve the API CLI for non-interactive automation.

2. Lightsail Daemon Runaway: $20–200+/month, unbounded

The Lightsail instance runs a jada daemon (managed by systemd at /etc/systemd/system/jada-agent.service) that processes agent-work cards and spawns the claude CLI command to reason about tasks.

The script at /home/ubuntu/jada_daemon.sh contains this pattern:

claude --model claude-opus-4-1 \
  --prompt-caching \
  --file "$card_file" \
  --output "$output_file"

Without a timeout, if Claude hangs or gets stuck processing a malformed card, the process consumes tokens indefinitely. Historical logs show a 2026-05-03 incident where a runaway daemon session consumed an estimated $50–150 in a single afternoon before being manually killed.

Why this is dangerous: The daemon runs unsupervised, auto-restarting on crash. A single stuck task can balloon the bill before anyone notices.

Immediate mitigation: Wrap the Claude invocation with a hard timeout. On the Lightsail box, modify /home/ubuntu/jada_daemon.sh to prepend timeout 300 (5-minute hard limit):

timeout 300 claude --model claude-opus-4-1 \
  --prompt-caching \
  --file "$card_file" \
  --output "$output_file"

This ensures any stuck process terminates after 300 seconds, capping the damage per incident.

3. Everything Else: <$20/month combined

All other systems—Stop hook scripts, Google Apps Script automations (CaroleEmailOps, WarmLeadResponder), Lambda functions, and the portfolio-intel daily task—are already configured to use Haiku, the most cost-effective model. Combined token consumption is negligible relative to the dominant two sources.

Haiku instances found:

  • voice_agent.py (Lightsail): Haiku for realtime voice synthesis
  • QDN_daily.gs (Google Apps Script): Haiku for daily report generation
  • jada_daily.sh (LaunchAgent): Haiku for local daily tasks
  • ai_repair_loop.py (local): Haiku for auto-repair reasoning
  • Tech blog Stop hook (settings.json): Haiku for post-publication checks

Marginal improvements are possible via prompt caching (estimated 10–15% savings) and Batch API for non-realtime tasks (25–50% savings), but the ROI is low given the small absolute spend on these systems.

Infrastructure & Command Examples

To replicate this audit in your own environment:


# Find all Anthropic SDK imports
grep -r "from anthropic import\|import anthropic" ~/Documents/repos

# Find all model references
grep -r "claude-opus\|claude-sonnet\|claude-haiku" ~/Documents/repos

# SSH to Lightsail and inspect daemon
ssh -i ~/.ssh/LightsailDefaultKey-us-west-2.pem ubuntu@34.239.233.28
cat /etc/systemd/system/jada-agent.service
cat /home/ubuntu/jada_daemon.sh

# Check systemd logs for daemon behavior
journalctl -u jada-agent -n 100

The full audit inventory was compiled into a structured table (stored at notes/claude_api_cost_audit_2026-05-28.md for historical reference) with columns: system name, function, execution environment, model, typical monthly tokens, and monthly cost.

Key Decisions & Rationale

  • Interactive dev → Claude.ai Max: The per-token CLI billing model is unsuitable for interactive work. A flat-rate subscription is more cost-effective and faster for this use case.
  • Daemon timeout → immediate deploy: This is a blocking issue—a production system can silently bleed money. The 5-minute timeout is aggressive enough to stop runaway processes but lenient enough for legitimate long-running tasks (we observed typical daemon tasks complete in 30–90 seconds).
  • No immediate changes to Haiku systems: The ROI on optimizing already-lean systems is low. Once the Sonnet runaway is contained, revisit prompt caching for the daemon if costs remain high.

What's Next