```html

Building a Charter Operations Transcript Analysis Pipeline: From Manual Audits to Automated Tool Pattern Detection

This weekend, I built an automated analysis pipeline to extract tool usage patterns from development session transcripts. The goal was straightforward: reduce permission prompt friction by identifying which bash commands and tools are consistently used in read-only contexts, then allowlist them globally. What emerged was a reusable pattern for analyzing Claude Code session data at scale.

The Problem: Permission Prompts on Every Session

Our development workflow involves frequent use of inspection commands—grep, find, cat, ls, sed—but Claude Code was prompting for permission on each session. While the permission system is designed as a safety feature, these read-only commands don't require user confirmation in a local development environment. The solution wasn't to disable permissions, but to be explicit about what's safe to auto-allow.

Technical Architecture: Three-Phase Analysis Pipeline

The pipeline consisted of three phases, each with specific file handling requirements:

Phase 1: Transcript Discovery and Structure Analysis

First, I needed to understand the transcript format across 50 recent sessions. Transcripts are stored in the ~/.claude/transcripts/ directory and contain JSON structures with message history. The critical challenge was that message structures vary depending on whether they contain tool calls.

I used find to locate all transcripts, then analyzed their structure with jq:

find ~/.claude/transcripts -name "*.json" -type f -mtime -30 | head -50 | while read f; do
  jq '.messages[0] | keys' "$f" 2>/dev/null
done | sort | uniq -c

This revealed that messages have inconsistent field names. Some used tool_use, others used tool_calls. I also discovered that tool use data lived inside the .content[] array rather than at the message root level.

Phase 2: Corrected Schema Extraction with Nested Content Parsing

After understanding the actual structure, I built a corrected extraction query:

jq -r '.messages[] | select(.role == "assistant") | 
  .content[]? | select(.type == "tool_use") | 
  .name' transcript.json | sort | uniq -c | sort -rn

This approach:

  • Filters for assistant messages (which contain tool invocations)
  • Expands the .content array to handle multiple tool calls per message
  • Selects only type: "tool_use" entries (filtering out text responses)
  • Extracts the tool name field and counts occurrences

Why this matters: The message structure in Claude Code transcripts uses a content array pattern where each element has a type field. Tool calls are just one type, alongside text responses and other content. Parsing this correctly required understanding that a single assistant message can contain mixed content types.

Phase 3: Cross-Transcript Aggregation and Permission Generation

Once I could extract patterns from individual transcripts, I aggregated across all 50 files:

for transcript in /path/to/transcripts/*.json; do
  jq -r '.messages[] | select(.role == "assistant") | 
    .content[]? | select(.type == "tool_use") | 
    .name' "$transcript" 2>/dev/null
done | sort | uniq -c | sort -rn

Results showed 8 dominant read-only commands:

CommandCountRead-Only
grep754✓
find112✓
sed68✓
cat64✓
echo73✓
ls80✓
cd99✓
wc36✓

Infrastructure: Global Permission Allowlist Configuration

Permissions are stored in ~/.claude/settings.json under the permissionAllowlist array. Each entry follows the pattern Bash(command:*) to allow all arguments to the command.

The existing allowlist included:

"Bash(grep:*)",
"Bash(find:*)",
"Bash(cat:*)",
"Bash(ls:*)",
"Bash(echo:*)",
"Bash(wc:*)",
"Bash(rg:*)",
"Bash(jq:*)",
"Bash(git:*)",
"Bash(gh:*)"

Based on transcript analysis, I added 6 missing entries that appeared frequently but weren't pre-allowed:

"Bash(sed:*)",
"Bash(sort:*)",
"Bash(uniq:*)",
"Bash(awk:*)",
"Bash(docker:*)",
"Bash(cd:*)"

Why this approach? Rather than allowing specific flags or argument patterns, the wildcard pattern * is safer for local development because:

  • These tools are POSIX standards with well-understood behavior
  • In a local dev environment, there's no network or external resource risk
  • We're building an allowlist, not a blocklist—explicit safety declarations matter
  • The wildcard prevents false positives from complex flag combinations

Key Decisions and Trade-offs

Why Not Use Auto Mode? Haiku 4.5 is already the smallest model available. Auto mode is designed to downgrade from larger models to save tokens. Since Haiku is the endpoint, auto mode doesn't apply. Keeping Haiku as the default and using /fast (Opus with faster output) for complex tasks is the optimal approach.

Why Aggregate Instead of Per-Command Analysis? Individual transcript analysis could identify patterns per session, but cross-transcript aggregation revealed the actual distribution. A command might appear in 50% of sessions but with different frequency—aggregation captures both breadth and depth.

Why Not Include State-Modifying Commands? Commands like rm, mv, chmod, curl (with POST), or git push require explicit confirmation. They're stateful operations that could cause damage. Permission prompts for these are features, not friction.

What's Next

This pipeline demonstrates a pattern for analyzing development workflows programmatically. Future improvements could include:

  • Analyzing argument patterns within allowed commands (e.g., grep -r vs. grep