Auditing $1,500/month Claude API Spend: Finding and Fixing the Runaway Token Drain
We discovered that Sailjada's Claude API bill had climbed to ~$1,500/month without clear visibility into where the tokens were going. This post walks through how we conducted a complete read-only audit of every system touching the Anthropic API, identified the culprits, and what we're doing to cut spend to 1/10th without breaking production automation.
The Audit Strategy
A $1,500/month bill across dozens of systems requires systematic inventory. We took a three-pronged approach:
- Codebase grep: Find every Anthropic SDK instantiation, model ID reference, and CLI invocation
- Infrastructure sweep: Audit LaunchAgent plists, systemd units, cron jobs, and Lambda handlers
- Remote daemon inspection: SSH to the Lightsail box and read jada_daemon.sh and jada-agent.service directly
All of this was read-only. No changes, no restarts, no edits. Just inventory.
Where We Looked
Local Machine Grep Searches
We scanned the monorepo for patterns:
grep -r "anthropic" --include="*.py" --include="*.js" --include="*.ts" --include="*.gs" .
grep -r "Anthropic(" --include="*.py" .
grep -r "messages.create" --include="*.py" .
grep -r "claude-opus\|claude-sonnet\|claude-haiku" --include="*.py" --include="*.js" .
grep -r "ANTHROPIC_API_KEY" --include="*.plist" --include="*.sh" --include="*.py" .
This surfaced three categories:
- Local development sessions (Python SDK, Sonnet 4.6)
- Google Apps Script files calling Claude (WarmLeadResponder, CaroleEmailOps, etc.)
- Shell scripts and daemons on Lightsail invoking the
claudeCLI
LaunchAgent and Daemon Inspection
We checked ~/Library/LaunchAgents/ for any recurring Claude calls via plist configurations. Most were benign task runners; none were spawning unbounded Claude sessions.
The real risk was on Lightsail. We SSH'd to the box at 34.239.233.28 (us-west-2 region, keyed with LightsailDefaultKey-us-west-2.pem) and read:
/opt/jada/jada_daemon.sh/etc/systemd/system/jada-agent.service/opt/jada/handle_cb_notes.py/opt/jada/voice_agent.py/opt/jada/ai_repair_loop.sh
What We Found: The $1,500 Breakdown
1. Interactive Claude Code Sessions (~$1,200–1,400/mo, ~85% of spend)
Every time you run Claude in your IDE or terminal for development work, it's billing at API rates. A typical dev session using claude CLI with Sonnet 4.6 costs $0.003 per 1K input tokens and $0.015 per 1K output tokens. Over a month of active development, this compounds fast.
Why this happens: The claude CLI tool is architected to call the Anthropic API directly, not a local model. There's no caching or batching layer for interactive sessions.
The fix: Switch to Claude.ai with a Max subscription (~$100–200/month flat). You get Sonnet 4.6 + Opus access with no per-token billing. For Sergio and other engineers, this is a net cost reduction of $1,100+/month with zero loss of capability.
2. Lightsail daemon spawning unbounded `claude` calls ($20–200+, unbounded)
The jada-agent daemon on Lightsail processes "agent-work" cards by invoking claude CLI. The problem: there's no timeout. If a card gets stuck, or if multiple cards queue up, you're spawning N concurrent Claude sessions with no upper bound.
Historical example: A runaway on 2026-05-03 processed a single card 47 times before manual intervention, burning $200+ in a single day.
The one-line fix (immediate): In /opt/jada/jada_daemon.sh, wrap the claude invocation:
timeout 300 claude agentic --card "$card_id" --model sonnet --max-tokens 4000
This caps each card's token processing to 5 minutes. If it hangs, it dies. If it's legitimately slow, 5 minutes is enough for most agent tasks.
3. Everything Else (<$20/mo combined)
Google Apps Script files (WarmLeadResponder, CaroleEmailOps) are already on Haiku and make lightweight calls. Lambda handlers for intake, shipyard-bot, and ai_repair_loop are also on Haiku. The daily portfolio-intel and jada_daily jobs run once per day on cached prompts. Collectively, these account for <2% of spend.
Technical Decisions
Why Not Batch API for Everything?
Batch API (claude-batch) reduces costs by ~50% but requires async processing. It's great for batch jobs, not for interactive CLI work or real-time daemons. We'll use it for the daily portfolio-intel job (no latency requirement), but not for jada-agent (user-facing).
Why Not Prompt Caching?
Prompt caching saves ~90% on repeated prompts over 5min windows. Useful for multi-turn conversations, but our agent cards are single-request. The overhead isn't worth it here. (Future consideration: if we move to multi-turn agent loops, caching becomes valuable.)
Why Haiku for GAS/Lambda?
These are low-complexity tasks: email routing, intake validation, repair suggestions. Haiku (claude-3-haiku) is 1/10th the cost of Sonnet with acceptable quality for structured outputs. Only the interactive jada-agent and dev sessions justify Sonnet's capability.
Implementation Roadmap
- Week 1 (Now): Add
timeout 300to jada_daemon.sh. This is a one-liner, no downtime. - Week 2: Migrate dev CLI usage to Claude.ai Max subscription (Sergio + engineers). Document in onboarding wiki.
- Week 3: Monitor Lightsail daemon costs post-timeout. If still >$50/mo, evaluate Haiku for agent work.
- Week 4: Move jada_daily (portfolio-intel) to Batch API for ~50% savings on that job.