Migrating a 10K-Token Blog Workflow to 5-Layer Model Workspace Protocol: Baseline Measurement and Pilot Design
What Was Done
We identified and measured a significant token-efficiency problem in the tech.queenofsandiego.com blog generation workflow, then designed a 5-layer Model Workspace Protocol migration to reduce per-session context overhead by ~97% (from ~10,925 tokens to a target ~370 tokens). This post documents the baseline measurement methodology, the architectural redesign, and the decision framework that makes this workflow the ideal pilot for generalizing the pattern across other mission-critical systems.
The Problem: Baseline Token Audit
The existing blog workflow operated from a monolithic context structure. We measured three distinct sources of always-on context load:
- Repos root CLAUDE.md (~942 tokens): Core developer instructions, tool definitions, and project scaffolding rules. This file is loaded by every session across all projects.
- Queen of San Diego site CLAUDE.md (~3,420 tokens): Deployment safety rules, pricing constraints, competitor analysis, and business logic. Every blog post session auto-loaded this entire file, though only ~200 tokens were blog-specific.
- tech_blog_generator.py (~6,563 tokens): The Python generator script itself, read into context for every session to provide tool definitions and examples.
Total baseline: ~10,925 tokens per session start. Approximately 90% waste.
The root cause: Layer 0 (project navigation), Layer 1 (routing), Layer 2 (stage-specific rules), and Layer 3 (business constraints) were all mixed into two CLAUDE.md files. Sessions needed to load layers they weren't currently using, creating false dependencies and preventing parallelization across multiple blog drafts.
Architecture: 5-Layer Model Workspace Protocol
We designed a new directory structure under repos/workspaces/tech-blog/ following this hierarchy:
repos/workspaces/tech-blog/
├── CLAUDE.md # Layer 0: ~80 tokens (nav only)
├── CONTEXT.md # Layer 1: router + task dispatch
├── reference/
│ └── voice.md # Style guide, examples, tone
├── 01-topic/
│ └── CONTEXT.md # Layer 2: topic ideation rules
├── 02-research/
│ └── CONTEXT.md # Layer 2: research methodology
├── 03-draft/
│ └── CONTEXT.md # Layer 2: writing + review gates
├── 04-publish/
│ └── CONTEXT.md # Layer 2: deployment + SEO rules
└── [per-run artifacts]/
├── topic_brief.md
├── research_notes.md
├── draft_v1.md
└── final_output.md
Layer assignments:
- Layer 0 (CLAUDE.md): Minimal map. Names the four stages, describes the CONTEXT.md router pattern, and explains when to load each stage's CONTEXT.md. ~80 tokens.
- Layer 1 (CONTEXT.md): Task router and dispatcher. Reads user input, determines which stage(s) are relevant, and loads only those stage CONTEXT.md files. Handles handoff between stages.
- Layer 2 (four stage CONTEXT.md files): Stage-specific rules, examples, guardrails, and success criteria. 01-topic handles ideation and trend analysis. 02-research handles source evaluation and fact-checking. 03-draft handles prose style, voice consistency, and internal review. 04-publish handles deployment, metadata, and monitoring.
- Layer 3 (reference/voice.md): Durable style guide, tone examples, and audience profile. Loaded only when drafting or reviewing.
- Layer 4 (per-run artifacts): Session outputs. Each blog post session produces its own artifact directory, enabling parallel work and preventing context pollution across concurrent drafts.
Why This Structure Works
Selective Loading: A session starting at the "topic" stage loads Layer 0 (~80 tokens) + Layer 1 (~300 tokens) + 01-topic CONTEXT.md (~200 tokens) = ~580 tokens. The same session does not load deploy-safety rules, pricing constraints, or competitor analysis. Those remain in Layer 3, loaded only at the 04-publish stage.
Parallelization: Two drafts can run concurrently without context collision. Each gets its own artifact directory. The Layer 1 router prevents race conditions on the stage CONTEXT.md files themselves (which are read-only during a session).
Reusability: The 5-layer pattern isn't blog-specific. Once validated here, the same structure applies to the deposit/booking workflow, the QoS email templates, the SCC admin dashboard, and the ticket-runner itself.
Low Risk: The blog workflow is non-mission-critical. Failures during the pilot don't impact revenue, customer access, or core systems. It's the ideal testbed.
Measurement Plan
We will measure three scenarios:
- Baseline (old): Start a blog post session with the monolithic CLAUDE.md. Record token load via Claude API usage logs.
- Pilot (new): Start an identical scenario using the 5-layer structure. Measure the same way.
- Real-world (new): Complete 3–5 actual blog post cycles (topic → research → draft → publish) using the new structure. Record wall-clock time, number of context switches, and cumulative tokens per post.
Success criteria: <500 tokens per session start (real-world average), zero breakage, no loss of quality or speed compared to the baseline.
Infrastructure and Tooling
The new structure lives in the existing Git repo at repos/workspaces/tech-blog/. Deployment is file-only—no infrastructure changes needed. The tech_blog_generator.py tool remains in repos/tools/tech_blog_generator.py but will be refactored to:
- Read the Layer 1 router (CONTEXT.md) to determine stage.
- Inject stage-specific CONTEXT.md into the prompt dynamically, rather than always including the full generator script.
- Route outputs to per-run artifact directories using timestamped naming.
No API changes, no new Lambda functions, no CloudFront invalidations, no Route53 updates.
Key Decisions
- Why nested CONTEXT.md over single generator config? CONTEXT.md integrates stage rules with examples and reasoning, making edits human-friendly. A JSON config would require dual maintenance (config + examples in CLAUDE.md).
- Why four stages? They map to distinct mental models and skill sets (ideation, research, writing, ops). Clear boundaries make it easier to gate approvals and handoff.
- Why measure wall-clock and tokens separately? Token count ≠ perceived latency. A 370-token session might feel faster or slower depending on which LLM tier is being used. We need both metrics to make billing and UX claims.
- Why not migrate everything at once? The