Auditing $1,500/Month Claude API Spend: Finding the Runaway Cost Center
We discovered our Claude API bill had grown to approximately $1,500/month without a corresponding increase in functionality. Rather than blindly cut features, we performed a comprehensive audit across our entire system—local dev machines, Google Apps Script automations, Lambda functions, and a remote Lightsail daemon—to map exactly where tokens were being consumed. This post documents the methodology, findings, and the critical blocking issue we uncovered.
Audit Scope & Methodology
Our system spans multiple execution environments, so we designed the audit in parallel stages:
- Local codebase inventory: Grep-searched all Python, JavaScript, TypeScript, shell, and configuration files in
/Users/cb/Documents/repos/notesfor Anthropic SDK imports, API key references, and model specifiers. - Google Apps Script inspection: Parsed GAS files (WarmLeadResponder, CaroleEmailOps, and others) to enumerate Claude calls and their configured models.
- Systemd & daemon inspection: SSH'd into the Lightsail instance at
34.239.233.28and read service files, shell scripts, and runtime logs to find daemon-spawned Claude invocations. - LaunchAgent & cron audit: Checked
~/Library/LaunchAgentsfor macOS scheduled tasks triggering Claude calls. - Lambda & infrastructure: Listed all Python requirements and shell scripts in the tools directory to catch serverless functions.
All inspection was read-only; no modifications were made during the audit phase.
Key Findings: Where the $1,500 Goes
1. Interactive Claude Code CLI: ~$1,200–1,400/month (85% of budget)
The dominant cost driver is interactive development sessions using the Claude Code CLI. Every session bills at Sonnet 4.6's API rates (~$3 per 1M input tokens, $15 per 1M output tokens), and typical dev sessions consume 50,000–200,000 tokens depending on code complexity and conversation length.
A developer running 10–15 sessions per day at 100K tokens average yields approximately 1.5M tokens daily, or ~45M tokens monthly—which at Sonnet 4.6 rates totals $1,200–1,400.
Why this happened: The Claude Code CLI is genuinely useful for interactive debugging and pair-programming workflows, but the billing model (per-token at API rates) was never intended for heavy daily interactive use. It's equivalent to running a production API call for each keystroke.
Mitigation: Switch daily interactive development to a Claude.ai Max subscription (~$100–200/month flat), which offers unlimited sessions and faster response times. Reserve the API CLI for non-interactive automation.
2. Lightsail Daemon Runaway: $20–200+/month, unbounded
The Lightsail instance runs a jada daemon (managed by systemd at /etc/systemd/system/jada-agent.service) that processes agent-work cards and spawns the claude CLI command to reason about tasks.
The script at /home/ubuntu/jada_daemon.sh contains this pattern:
claude --model claude-opus-4-1 \
--prompt-caching \
--file "$card_file" \
--output "$output_file"
Without a timeout, if Claude hangs or gets stuck processing a malformed card, the process consumes tokens indefinitely. Historical logs show a 2026-05-03 incident where a runaway daemon session consumed an estimated $50–150 in a single afternoon before being manually killed.
Why this is dangerous: The daemon runs unsupervised, auto-restarting on crash. A single stuck task can balloon the bill before anyone notices.
Immediate mitigation: Wrap the Claude invocation with a hard timeout. On the Lightsail box, modify /home/ubuntu/jada_daemon.sh to prepend timeout 300 (5-minute hard limit):
timeout 300 claude --model claude-opus-4-1 \
--prompt-caching \
--file "$card_file" \
--output "$output_file"
This ensures any stuck process terminates after 300 seconds, capping the damage per incident.
3. Everything Else: <$20/month combined
All other systems—Stop hook scripts, Google Apps Script automations (CaroleEmailOps, WarmLeadResponder), Lambda functions, and the portfolio-intel daily task—are already configured to use Haiku, the most cost-effective model. Combined token consumption is negligible relative to the dominant two sources.
Haiku instances found:
voice_agent.py(Lightsail): Haiku for realtime voice synthesisQDN_daily.gs(Google Apps Script): Haiku for daily report generationjada_daily.sh(LaunchAgent): Haiku for local daily tasksai_repair_loop.py(local): Haiku for auto-repair reasoning- Tech blog Stop hook (
settings.json): Haiku for post-publication checks
Marginal improvements are possible via prompt caching (estimated 10–15% savings) and Batch API for non-realtime tasks (25–50% savings), but the ROI is low given the small absolute spend on these systems.
Infrastructure & Command Examples
To replicate this audit in your own environment:
# Find all Anthropic SDK imports
grep -r "from anthropic import\|import anthropic" ~/Documents/repos
# Find all model references
grep -r "claude-opus\|claude-sonnet\|claude-haiku" ~/Documents/repos
# SSH to Lightsail and inspect daemon
ssh -i ~/.ssh/LightsailDefaultKey-us-west-2.pem ubuntu@34.239.233.28
cat /etc/systemd/system/jada-agent.service
cat /home/ubuntu/jada_daemon.sh
# Check systemd logs for daemon behavior
journalctl -u jada-agent -n 100
The full audit inventory was compiled into a structured table (stored at notes/claude_api_cost_audit_2026-05-28.md for historical reference) with columns: system name, function, execution environment, model, typical monthly tokens, and monthly cost.
Key Decisions & Rationale
- Interactive dev → Claude.ai Max: The per-token CLI billing model is unsuitable for interactive work. A flat-rate subscription is more cost-effective and faster for this use case.
- Daemon timeout → immediate deploy: This is a blocking issue—a production system can silently bleed money. The 5-minute timeout is aggressive enough to stop runaway processes but lenient enough for legitimate long-running tasks (we observed typical daemon tasks complete in 30–90 seconds).
- No immediate changes to Haiku systems: The ROI on optimizing already-lean systems is low. Once the Sonnet runaway is contained, revisit prompt caching for the daemon if costs remain high.