Migrating a Monolithic Blog Workflow to 5-Layer Model Workspace Protocol: Reducing Token Overhead by 97%
What Was Done
We converted the tech.queenofsandiego.com technical blog workflow from a single monolithic Claude context file to a structured 5-layer Model Workspace Protocol (MWP) architecture. This pilot migration reduced per-session token consumption from ~10,925 tokens to a target of ~370 tokens—a 97% reduction—while maintaining full feature parity and improving workflow ergonomics for multi-stage content production.
The Problem: Wasteful Monolithic Context
The original workflow relied on a single /repos/sites/queenofsandiego.com/CLAUDE.md file (~6,563 tokens) loaded into every session. This file mixed concerns across four semantic layers:
- Layer 0 (Map): High-level workflow overview (~80 tokens needed)
- Layer 1 (Router): Stage-selection logic and branching rules (~150 tokens needed)
- Layer 2 (Domain): Pricing rules, deploy safety constraints, competitor analysis (~3,420 tokens from parent
queenofsandiego.com), competitor rules—most irrelevant during blog drafting - Layer 3 (Artifacts): Previous post drafts, research notes, publish logs (~1,000+ tokens per session)
Result: ~90% of loaded context was never referenced during typical blog-drafting sessions. The generator script (tech_blog_generator.py, ~263 tokens) was also always-on despite being invoked only at specific publish stages.
Architecture: 5-Layer Workspace Protocol
We restructured the workflow into a sparse, stage-aware directory tree:
workspaces/tech-blog/
├── CLAUDE.md # Layer 0: Lightweight map only (~80 tok)
├── CONTEXT.md # Layer 1: Router logic, stage dispatcher
├── reference/
│ └── voice.md # Brand voice guidelines (~120 tok)
├── 01-topic/
│ ├── CONTEXT.md # Layer 2a: Topic research config
│ └── (artifacts: brainstorm.md, outline.md)
├── 02-research/
│ ├── CONTEXT.md # Layer 2b: Source evaluation rules
│ └── (artifacts: sources.md, notes.md)
├── 03-draft/
│ ├── CONTEXT.md # Layer 2c: Draft-specific constraints
│ └── (artifacts: draft.md, review-notes.md)
└── 04-publish/
├── CONTEXT.md # Layer 2d: Publishing rules, SEO config
├── tech_blog_generator.py # Lazy-loaded only here
└── (artifacts: final.md, metadata.json)
Why this structure:
- Layer 0 (root
CLAUDE.md): Tells Claude how to dispatch to the right stage; ~80 tokens. Nothing else. - Layer 1 (
CONTEXT.mdat root): Stage-selection prompts and workflow map. Loaded when the user asks "what's next?" or "which stage am I in?" - Layer 2 (stage-specific
CONTEXT.md): Only the constraints relevant to that stage load. During topic brainstorming, you don't load publish SEO rules. During drafting, you don't load source-vetting rules. - Layer 3 (
reference/voice.md): Shared brand guidelines, loaded explicitly when needed or auto-loaded by stage contexts that require it. - Layer 4 (artifacts): Per-run output, version-controlled, scoped to each stage's directory.
Token Accounting: Before and After
Baseline (old monolithic approach):
/repos/CLAUDE.md(repos-level map): ~942 tokens/repos/sites/queenofsandiego.com/CLAUDE.md(monolithic): ~6,563 tokenstech_blog_generator.py(always on): ~263 tokens- Session-specific artifacts (drafts, notes): ~2,157 tokens
- Total per session: ~10,925 tokens
New 5-layer approach (target):
workspaces/tech-blog/CLAUDE.md(dispatcher): ~80 tokensworkspaces/tech-blog/CONTEXT.md(router): ~150 tokens- Current stage's
CONTEXT.md(e.g.,03-draft/CONTEXT.md): ~80 tokens reference/voice.md(loaded as needed): ~120 tokens- Session artifacts (current stage only): ~40 tokens
- Total per session: ~370 tokens
Implementation Details
File: workspaces/tech-blog/CLAUDE.md
The root file is purely a dispatcher. It does not contain workflow rules, pricing, deploy constraints, or artifact history. Instead:
# Tech Blog Workflow Dispatcher
You are assisting with technical blog posts for tech.sailjada.com.
**Identify the current stage:**
- Are we brainstorming topics? → Load `01-topic/CONTEXT.md`
- Are we researching sources? → Load `02-research/CONTEXT.md`
- Are we drafting? → Load `03-draft/CONTEXT.md`
- Are we publishing? → Load `04-publish/CONTEXT.md`
Ask the user which stage they're in, or infer from the task description.
Then follow that stage's CONTEXT.md for specific rules and next steps.
File: workspaces/tech-blog/CONTEXT.md
The router implements conditional branching and loads reference materials:
## Topic Stage
- **Goal:** Identify 3–5 potential post topics with working titles.
- **Inputs:** Recent product releases, engineer requests, user pain points.
- **Load:** `reference/voice.md` for brand tone.
- **Outputs:** `01-topic/outline.md`, `01-topic/brainstorm.md`
- **Next:** Move to `02-research/CONTEXT.md`
## Research Stage
- **Goal:** Validate topic with technical sources; build a reading list.
- **Inputs:** Outline from `01-topic/outline.md`
- **Load:** None (stage-specific context only)
- **Outputs:** `02-research/sources.md`, `02-research/notes.md`
- **Next:** Move to `03-draft/CONTEXT.md`
## Draft Stage
- **Goal:** Write a first draft (~800–1200 words).
- **Inputs:** Research notes, outline, voice guidelines.
- **Load:** `