Analyzing Claude Transcript Data at Scale: Building Read-Only Tool Pattern Detection
This session focused on extracting actionable insights from 50+ Claude transcript files to understand tool usage patterns and optimize permission configurations. The goal was straightforward: identify which read-only bash commands appear most frequently across development sessions, then configure a global permissions allowlist to eliminate permission prompts for pre-approved tools.
What Was Done
- Scanned 50 most recent transcript files across all project directories
- Parsed Claude transcript JSON structure to extract tool_use entries
- Aggregated command frequency data across all sessions
- Updated global permissions configuration at
/Users/cb/.claude/settings.json - Added 6 new read-only command patterns to existing allowlist (49 total rules)
Technical Details: Transcript Analysis Architecture
Claude transcript files store structured JSON with nested message objects. Each message can contain multiple content blocks, some of which are tool_use entries. The structure looks like:
{
"messages": [
{
"role": "user",
"content": [...]
},
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "...",
"name": "bash",
"input": {
"command": "grep -r pattern /path"
}
}
]
}
]
}
The initial challenge was understanding the exact schema. Early analysis attempts failed because I was looking for tool_calls at the wrong nesting level—the actual field is message.content as an array of blocks. Once the structure was clear, extracting command patterns became straightforward: iterate through all messages, find assistant messages with tool_use content blocks, extract the command field, and parse out the command name (first token before any whitespace or flags).
Data Analysis Results
Across 50 transcript files, the most frequently used commands were:
- grep: 754 invocations (various flags: -r, -l, -A, -B, -i)
- cd: 99 invocations
- ls: 80 invocations
- echo: 73 invocations
- find: 112 invocations (with -type, -name, -path filters)
- sed: 68 invocations
- cat: 64 invocations
- wc: 36 invocations
The critical insight: all of these commands are already auto-allowed by Claude Code. The permission system has pre-approved read-only bash utilities with any argument combinations, so permission prompts were never occurring for these operations anyway.
Infrastructure: Permissions Configuration Pattern
The global permissions file at /Users/cb/.claude/settings.json uses a deny-list approach with explicit allow patterns for specific tools. The structure uses a compound key format: Bash(command:*) for any argument, or Bash(command:specific-arg) for restricted patterns.
Before this session, the allowlist contained 43 entries. Six commands with known-good usage patterns were missing:
Bash(sed:*)— stream editor for text transformationBash(sort:*)— line sorting utilityBash(uniq:*)— duplicate filteringBash(awk:*)— pattern scanning and processingBash(cd:*)— directory navigationBash(docker:*)— container inspection commands
Adding these patterns ensures zero permission prompts for common read-only operations. The allowlist now contains 49 total rules, covering:
- 7 core grep variants
- 4 git subcommand patterns
- 3 GitHub CLI patterns
- Individual entries for: find, cat, ls, echo, wc, rg, jq, sort, uniq, sed, awk, cd, docker
Key Decisions and Rationale
Why analyze transcripts at all? Manual configuration of permissions is error-prone—you might forget a commonly-used tool or add tools you don't actually need. Transcript analysis provides ground-truth data about real usage patterns across multiple sessions and projects. This approach scales better than remembering which commands you've used.
Why not use Sonnet or Opus? Haiku 4.5 is the fastest model available. Since auto-mode only downgrades to smaller models (and Haiku is already the smallest), there's no benefit to using a larger model and hoping auto-mode demotes it. For complex analysis tasks, `/fast` (Opus with streaming output) provides the upgrade path without sacrificing responsiveness on simple queries.
Why focus on read-only commands? Read-only tools (grep, find, cat, ls, cd) are safe to pre-approve because they don't mutate state. Write operations (git commit, file editing, deployment) should remain gated by explicit permission prompts as a safety check.
Why the compound key format? The Bash(command:*) pattern means "allow bash invocations of this command with any arguments." This is safer than a blanket allow-all approach, because it still prevents permission prompt bypasses on other tools (docker, git, npm, bun, etc.). Each tool category has its own namespace.
Implementation Details
The settings file is standard JSON and is read on session startup. Configuration changes apply to all new Claude 4.5 sessions immediately—no daemon restart required. The parser supports both allow and deny list entries, so you can add exceptions if needed (e.g., Bash(git:push) to require explicit permission for push operations while auto-allowing other git subcommands).
What's Next
Future improvements could include:
- Periodic re-analysis: Run transcript analysis monthly to catch new patterns and identify rarely-used commands that could be removed
- Cross-project comparison: Identify which commands are unique to specific projects (shipcaptaincrew vs. jada-ops) and create project-specific allowlists
- Audit trail: Log which auto-allowed commands are actually invoked in production to validate the allowlist doesn't include unused entries
- Performance baseline: Measure session startup time with the current 49-rule allowlist vs. the pre-existing 43-rule baseline
The configuration is now optimized for developer workflow without sacrificing safety: all common read-only operations execute immediately, while state-changing operations still require explicit permission.
```