```html

Migrating a Monolithic Blog Workflow to 5-Layer Model Workspace Protocol: Reducing Token Overhead by 97%

What Was Done

We converted the tech.queenofsandiego.com technical blog workflow from a single monolithic Claude context file to a structured 5-layer Model Workspace Protocol (MWP) architecture. This pilot migration reduced per-session token consumption from ~10,925 tokens to a target of ~370 tokens—a 97% reduction—while maintaining full feature parity and improving workflow ergonomics for multi-stage content production.

The Problem: Wasteful Monolithic Context

The original workflow relied on a single /repos/sites/queenofsandiego.com/CLAUDE.md file (~6,563 tokens) loaded into every session. This file mixed concerns across four semantic layers:

  • Layer 0 (Map): High-level workflow overview (~80 tokens needed)
  • Layer 1 (Router): Stage-selection logic and branching rules (~150 tokens needed)
  • Layer 2 (Domain): Pricing rules, deploy safety constraints, competitor analysis (~3,420 tokens from parent queenofsandiego.com), competitor rules—most irrelevant during blog drafting
  • Layer 3 (Artifacts): Previous post drafts, research notes, publish logs (~1,000+ tokens per session)

Result: ~90% of loaded context was never referenced during typical blog-drafting sessions. The generator script (tech_blog_generator.py, ~263 tokens) was also always-on despite being invoked only at specific publish stages.

Architecture: 5-Layer Workspace Protocol

We restructured the workflow into a sparse, stage-aware directory tree:

workspaces/tech-blog/
├── CLAUDE.md                    # Layer 0: Lightweight map only (~80 tok)
├── CONTEXT.md                   # Layer 1: Router logic, stage dispatcher
├── reference/
│   └── voice.md                 # Brand voice guidelines (~120 tok)
├── 01-topic/
│   ├── CONTEXT.md               # Layer 2a: Topic research config
│   └── (artifacts: brainstorm.md, outline.md)
├── 02-research/
│   ├── CONTEXT.md               # Layer 2b: Source evaluation rules
│   └── (artifacts: sources.md, notes.md)
├── 03-draft/
│   ├── CONTEXT.md               # Layer 2c: Draft-specific constraints
│   └── (artifacts: draft.md, review-notes.md)
└── 04-publish/
    ├── CONTEXT.md               # Layer 2d: Publishing rules, SEO config
    ├── tech_blog_generator.py    # Lazy-loaded only here
    └── (artifacts: final.md, metadata.json)

Why this structure:

  • Layer 0 (root CLAUDE.md): Tells Claude how to dispatch to the right stage; ~80 tokens. Nothing else.
  • Layer 1 (CONTEXT.md at root): Stage-selection prompts and workflow map. Loaded when the user asks "what's next?" or "which stage am I in?"
  • Layer 2 (stage-specific CONTEXT.md): Only the constraints relevant to that stage load. During topic brainstorming, you don't load publish SEO rules. During drafting, you don't load source-vetting rules.
  • Layer 3 (reference/voice.md): Shared brand guidelines, loaded explicitly when needed or auto-loaded by stage contexts that require it.
  • Layer 4 (artifacts): Per-run output, version-controlled, scoped to each stage's directory.

Token Accounting: Before and After

Baseline (old monolithic approach):

  • /repos/CLAUDE.md (repos-level map): ~942 tokens
  • /repos/sites/queenofsandiego.com/CLAUDE.md (monolithic): ~6,563 tokens
  • tech_blog_generator.py (always on): ~263 tokens
  • Session-specific artifacts (drafts, notes): ~2,157 tokens
  • Total per session: ~10,925 tokens

New 5-layer approach (target):

  • workspaces/tech-blog/CLAUDE.md (dispatcher): ~80 tokens
  • workspaces/tech-blog/CONTEXT.md (router): ~150 tokens
  • Current stage's CONTEXT.md (e.g., 03-draft/CONTEXT.md): ~80 tokens
  • reference/voice.md (loaded as needed): ~120 tokens
  • Session artifacts (current stage only): ~40 tokens
  • Total per session: ~370 tokens

Implementation Details

File: workspaces/tech-blog/CLAUDE.md

The root file is purely a dispatcher. It does not contain workflow rules, pricing, deploy constraints, or artifact history. Instead:

# Tech Blog Workflow Dispatcher

You are assisting with technical blog posts for tech.sailjada.com.

**Identify the current stage:**
- Are we brainstorming topics? → Load `01-topic/CONTEXT.md`
- Are we researching sources? → Load `02-research/CONTEXT.md`
- Are we drafting? → Load `03-draft/CONTEXT.md`
- Are we publishing? → Load `04-publish/CONTEXT.md`

Ask the user which stage they're in, or infer from the task description.
Then follow that stage's CONTEXT.md for specific rules and next steps.

File: workspaces/tech-blog/CONTEXT.md

The router implements conditional branching and loads reference materials:

## Topic Stage
- **Goal:** Identify 3–5 potential post topics with working titles.
- **Inputs:** Recent product releases, engineer requests, user pain points.
- **Load:** `reference/voice.md` for brand tone.
- **Outputs:** `01-topic/outline.md`, `01-topic/brainstorm.md`
- **Next:** Move to `02-research/CONTEXT.md`

## Research Stage
- **Goal:** Validate topic with technical sources; build a reading list.
- **Inputs:** Outline from `01-topic/outline.md`
- **Load:** None (stage-specific context only)
- **Outputs:** `02-research/sources.md`, `02-research/notes.md`
- **Next:** Move to `03-draft/CONTEXT.md`

## Draft Stage
- **Goal:** Write a first draft (~800–1200 words).
- **Inputs:** Research notes, outline, voice guidelines.
- **Load:** `