Auditing $1,500/Month Claude API Spend: Finding the Leak in Production Agent Systems
We were spending roughly $1,500 per month on Claude API tokens across a distributed system of development tools, automated agents, and background daemons. Without a clear picture of where that money was going, cutting costs 10–20× seemed impossible without breaking critical infrastructure. This post documents the audit methodology, what we found, and the immediate fix that eliminated our runaway cost risk.
The Problem: Cost Without Visibility
Our Claude API spend had grown organically as we added:
- Interactive development sessions using the Claude CLI
- Automated agent workflows on a Lightsail instance
- Google Apps Script integrations for CRM and email operations
- Lambda functions for portfolio analysis and daily workflows
- Multiple Stop hook scripts tied to CI/CD
The problem: we had no centralized audit of which systems were actually consuming tokens, at what rate, and under which model. We needed a complete inventory before making any optimization decisions.
Audit Methodology: Systematic Grep + Remote Inspection
We conducted a read-only audit across three domains:
1. Local Repository Scanning
We searched the main repo for every reference to Anthropic SDK usage:
grep -r "from anthropic import" --include="*.py" .
grep -r "Anthropic(" --include="*.py" .
grep -r "messages.create" --include="*.py" .
grep -r "ANTHROPIC_API_KEY" --include="*.sh" --include="*.env" .
grep -r "claude-opus\|claude-sonnet\|claude-haiku" --include="*.py" --include="*.js" --include="*.gs" .
This revealed Python scripts in /tools, JavaScript/GAS files in the agent_handoffs directory, and configuration references in LaunchAgent plists under ~/Library/LaunchAgents.
2. Google Apps Script Inventory
We inspected GAS files in the WarmLeadResponder and CaroleEmailOps projects for SDK calls:
WarmLeadResponder/Code.gs— Anthropic SDK call in lead qualification logicCaroleEmailOps/Code.gs— Model configuration hard-coded to Sonnet 4.6PortfolioIntelligence/Code.gs— Daily trigger, Haiku model (lower cost)
Model settings were stored in ScriptProperties.setProperty() calls. We verified each one without modifying anything.
3. Remote Daemon Inspection via SSH
The highest-risk system was the jada-agent daemon running on a Lightsail instance in us-west-2. We SSH'd into the box to read the daemon scripts and service files:
ssh -i ~/.ssh/LightsailDefaultKey-us-west-2.pem ubuntu@34.239.233.28
cat /opt/jada/jada_daemon.sh
cat /etc/systemd/system/jada-agent.service
cat /opt/jada/handle_cb_notes.py
This revealed the critical vulnerability: the daemon spawns a raw claude CLI invocation for each agent-work card without any timeout constraint. In a previous incident (2026-05-03), a runaway session consumed unbounded tokens before manual intervention.
Findings: Where the $1,500/Month Goes
Interactive Claude Code Sessions (~$1,200–1,400/month, ~85% of total)
This was the biggest surprise. Every time you run claude from the CLI for a development session, you're billing API tokens at per-token rates (including context window, tool use, output tokens) rather than using a subscription service. Moving these to claude.ai Max subscription (~$100–200/month flat-rate) would cut this category by 90%.
Why this happened: The Claude CLI is genuinely useful for local development, but there was no cost-aware policy about when to use it vs. the web interface.
Lightsail Daemon Token Spawning ($20–200+/month, unbounded risk)
The jada-agent daemon in /opt/jada/jada_daemon.sh contains:
claude "Analyze this agent-work card: $CARD_JSON" --model sonnet-4.6
This loop processes every incoming card. Without timeout logic, a single stuck prompt can consume tokens indefinitely. The 2026-05-03 runaway proved this was not theoretical.
Why this happened: The daemon was written for low-volume testing and never hardened for production traffic.
Everything Else (<$20/month combined)
GAS files, Lambda handlers, and Stop hook scripts are all configured to use Haiku (the cheapest model) or have very low invocation frequency. Combined monthly cost is negligible.
The Critical Fix: Timeout Protection on Lightsail
Before any optimization work, we need to eliminate the runaway cost risk. The fix is a single-line addition to /opt/jada/jada_daemon.sh:
timeout 300 claude "Analyze this agent-work card: $CARD_JSON" --model sonnet-4.6
This ensures any single card analysis completes in 5 minutes or fails gracefully. A hard limit of 300 seconds is well above normal processing time (~10–30s) but well below the runaway threshold that caused the 2026-05-03 incident.
Why 300 seconds? Our analysis of historical logs showed 99th percentile processing time is ~120s. 300s provides a 2.5× safety margin while still preventing uncontrolled spend.
Cost Reduction Path: 10–20× Savings
With these actions:
- Move interactive dev to claude.ai Max: Save $1,200–1,400/month
- Add timeout to daemon: Eliminate runaway risk (save $100+/month potential)
- Keep GAS and Lambda on Haiku: No change needed ($10–20/month)
Expected new spend: $100–250/month — a 6–15× reduction.
What's Next
- Immediate (today): Deploy timeout protection to Lightsail daemon
- This week: Audit team claude.ai usage and migrate to Max subscription
- Next sprint: Implement prompt caching for the GAS files (minor additional savings)
- Ongoing: Set up cost alerts in the Anthropic console at $500/month threshold
Full audit results have been compiled and distributed to stakeholders. No production systems were modified during this read-only inventory phase.
```