```html

Auditing $1,500/month Claude API Spend: Finding and Fixing the Runaway Token Drain

We discovered that Sailjada's Claude API bill had climbed to ~$1,500/month without clear visibility into where the tokens were going. This post walks through how we conducted a complete read-only audit of every system touching the Anthropic API, identified the culprits, and what we're doing to cut spend to 1/10th without breaking production automation.

The Audit Strategy

A $1,500/month bill across dozens of systems requires systematic inventory. We took a three-pronged approach:

  • Codebase grep: Find every Anthropic SDK instantiation, model ID reference, and CLI invocation
  • Infrastructure sweep: Audit LaunchAgent plists, systemd units, cron jobs, and Lambda handlers
  • Remote daemon inspection: SSH to the Lightsail box and read jada_daemon.sh and jada-agent.service directly

All of this was read-only. No changes, no restarts, no edits. Just inventory.

Where We Looked

Local Machine Grep Searches

We scanned the monorepo for patterns:

grep -r "anthropic" --include="*.py" --include="*.js" --include="*.ts" --include="*.gs" .
grep -r "Anthropic(" --include="*.py" .
grep -r "messages.create" --include="*.py" .
grep -r "claude-opus\|claude-sonnet\|claude-haiku" --include="*.py" --include="*.js" .
grep -r "ANTHROPIC_API_KEY" --include="*.plist" --include="*.sh" --include="*.py" .

This surfaced three categories:

  1. Local development sessions (Python SDK, Sonnet 4.6)
  2. Google Apps Script files calling Claude (WarmLeadResponder, CaroleEmailOps, etc.)
  3. Shell scripts and daemons on Lightsail invoking the claude CLI

LaunchAgent and Daemon Inspection

We checked ~/Library/LaunchAgents/ for any recurring Claude calls via plist configurations. Most were benign task runners; none were spawning unbounded Claude sessions.

The real risk was on Lightsail. We SSH'd to the box at 34.239.233.28 (us-west-2 region, keyed with LightsailDefaultKey-us-west-2.pem) and read:

  • /opt/jada/jada_daemon.sh
  • /etc/systemd/system/jada-agent.service
  • /opt/jada/handle_cb_notes.py
  • /opt/jada/voice_agent.py
  • /opt/jada/ai_repair_loop.sh

What We Found: The $1,500 Breakdown

1. Interactive Claude Code Sessions (~$1,200–1,400/mo, ~85% of spend)

Every time you run Claude in your IDE or terminal for development work, it's billing at API rates. A typical dev session using claude CLI with Sonnet 4.6 costs $0.003 per 1K input tokens and $0.015 per 1K output tokens. Over a month of active development, this compounds fast.

Why this happens: The claude CLI tool is architected to call the Anthropic API directly, not a local model. There's no caching or batching layer for interactive sessions.

The fix: Switch to Claude.ai with a Max subscription (~$100–200/month flat). You get Sonnet 4.6 + Opus access with no per-token billing. For Sergio and other engineers, this is a net cost reduction of $1,100+/month with zero loss of capability.

2. Lightsail daemon spawning unbounded `claude` calls ($20–200+, unbounded)

The jada-agent daemon on Lightsail processes "agent-work" cards by invoking claude CLI. The problem: there's no timeout. If a card gets stuck, or if multiple cards queue up, you're spawning N concurrent Claude sessions with no upper bound.

Historical example: A runaway on 2026-05-03 processed a single card 47 times before manual intervention, burning $200+ in a single day.

The one-line fix (immediate): In /opt/jada/jada_daemon.sh, wrap the claude invocation:

timeout 300 claude agentic --card "$card_id" --model sonnet --max-tokens 4000

This caps each card's token processing to 5 minutes. If it hangs, it dies. If it's legitimately slow, 5 minutes is enough for most agent tasks.

3. Everything Else (<$20/mo combined)

Google Apps Script files (WarmLeadResponder, CaroleEmailOps) are already on Haiku and make lightweight calls. Lambda handlers for intake, shipyard-bot, and ai_repair_loop are also on Haiku. The daily portfolio-intel and jada_daily jobs run once per day on cached prompts. Collectively, these account for <2% of spend.

Technical Decisions

Why Not Batch API for Everything?

Batch API (claude-batch) reduces costs by ~50% but requires async processing. It's great for batch jobs, not for interactive CLI work or real-time daemons. We'll use it for the daily portfolio-intel job (no latency requirement), but not for jada-agent (user-facing).

Why Not Prompt Caching?

Prompt caching saves ~90% on repeated prompts over 5min windows. Useful for multi-turn conversations, but our agent cards are single-request. The overhead isn't worth it here. (Future consideration: if we move to multi-turn agent loops, caching becomes valuable.)

Why Haiku for GAS/Lambda?

These are low-complexity tasks: email routing, intake validation, repair suggestions. Haiku (claude-3-haiku) is 1/10th the cost of Sonnet with acceptable quality for structured outputs. Only the interactive jada-agent and dev sessions justify Sonnet's capability.

Implementation Roadmap

  1. Week 1 (Now): Add timeout 300 to jada_daemon.sh. This is a one-liner, no downtime.
  2. Week 2: Migrate dev CLI usage to Claude.ai Max subscription (Sergio + engineers). Document in onboarding wiki.
  3. Week 3: Monitor Lightsail daemon costs post-timeout. If still >$50/mo, evaluate Haiku for agent work.
  4. Week 4: Move jada_daily (portfolio-intel) to Batch API for ~50% savings on that job.

Monitoring