Migrating tech-blog to 5-Layer Model Workspace Protocol: A Token-Efficient Workflow Redesign
What Was Done
We redesigned the tech.queenofsandiego.com blog publishing workflow from a monolithic single-CLAUDE.md structure to a distributed 5-layer Model Workspace Protocol (MWP) architecture. The migration reduces per-session token overhead by ~97% (from ~10,925 tokens to a target ~370 tokens) while maintaining full feature parity and improving workflow clarity for multi-stage content creation.
Baseline Analysis: The Problem
The original structure loaded three components into every session:
- Layer 0 (Map):
repos/CLAUDE.md(~942 tokens) — organization-wide navigation, irrelevant to a single blog post - Layer 3 (Rules): Embedded in
queenofsandiego.com/CLAUDE.md(~3,420 tokens) — deployment safety guardrails, pricing policy, competitor analysis, business context—most unnecessary for ideation or drafting - Generator:
tech_blog_generator.py(~6,563 tokens) — the actual orchestration script, loaded regardless of session phase
Combined, this forced ~10,925 tokens of always-on context. Worse, the Layer 3 rules were mixed with Layer 0 maps, making it impossible to load safety checks without drowning in unrelated business policy.
Technical Architecture: 5-Layer Model Workspace
We restructured as follows:
Layer 0: Workspace Root
New path: workspaces/tech-blog/CLAUDE.md (~80 tokens)
Purpose: Identity + navigation only.
Contains: Workspace name, brief description, pointer to Layer 1 router.
Never loads business rules or tool code.
Layer 1: Phase Router
New path: workspaces/tech-blog/CONTEXT.md (~120 tokens)
Purpose: Dispatch logic and phase workflow.
Contains: Session-phase detection, which Layer 2 CONTEXT to load,
which artifacts directory to use.
Layer 1.5: Reusable Reference
New path: workspaces/tech-blog/reference/voice.md (~150 tokens)
Purpose: Editorial voice, style guide, brand tone.
Why separate: Loaded only during research + draft phases, not topic selection.
Layer 2: Stage-Specific Contexts
Four separate CONTEXT.md files, each auto-loaded only when entering that stage:
workspaces/tech-blog/01-topic/CONTEXT.md(~100 tokens) — ideation, topic validation, editorial calendar rulesworkspaces/tech-blog/02-research/CONTEXT.md(~100 tokens) — research methodology, source evaluation, outline structureworkspaces/tech-blog/03-draft/CONTEXT.md(~100 tokens) — writing style, code example standards, technical depth guidanceworkspaces/tech-blog/04-publish/CONTEXT.md(~150 tokens) — SEO directives, deploy safety (separated from business rules), CI/CD integration
Layer 3: Guardrails (Separated)
Deploy safety and technical constraints moved to workspaces/tech-blog/reference/deploy-safety.md (~80 tokens), loaded only in the publish stage. Pricing and business rules remain in the root queenofsandiego.com space—not auto-loaded into blog sessions.
Layer 4: Per-Run Artifacts
Each session creates subdirectories under the active stage:
workspaces/tech-blog/03-draft/2026-06-03-deployment-patterns/
├── prompt.txt (session instructions)
├── outline.md (generated outline)
├── draft.md (in-progress article)
└── research.md (sourced facts + citations)
Workspace Directory Structure
workspaces/tech-blog/
├── CLAUDE.md (Layer 0: ~80 tokens)
├── CONTEXT.md (Layer 1 router: ~120 tokens)
├── reference/
│ ├── voice.md (Editorial tone: ~150 tokens)
│ └── deploy-safety.md (Publish guardrails: ~80 tokens)
├── 01-topic/
│ ├── CONTEXT.md (Ideation phase: ~100 tokens)
│ └── 2026-06-03-icm-migration/ (example artifact dir)
│ ├── prompt.txt
│ └── candidates.md
├── 02-research/
│ ├── CONTEXT.md (Research phase: ~100 tokens)
│ └── 2026-06-03-deployment-patterns/
│ ├── sources.md
│ └── outline.md
├── 03-draft/
│ ├── CONTEXT.md (Writing phase: ~100 tokens)
│ └── 2026-06-03-deployment-patterns/
│ ├── draft.md
│ ├── code-examples/
│ └── feedback.md
└── 04-publish/
├── CONTEXT.md (Publish phase: ~150 tokens)
└── 2026-06-03-deployment-patterns/
├── final.html
├── deploy.log
└── lighthouse-report.json
Key Decisions and Rationale
Why Separate Layer 3 (Rules) from Layer 0 (Maps)
The original monolith forced you to load all business rules even during topic brainstorming. Splitting them means:
- Ideation phase: Load only editorial voice + calendar rules (~250 tokens)
- Publish phase: Load deploy safety + SEO rules (~230 tokens)
- Never: Load unrelated pricing or competitor analysis into a blog session
Why Four Discrete Stage Directories, Not a Linear Progression
The workflow is nonlinear: you may revisit research mid-draft, or start multiple topics in parallel. Each stage directory is self-contained, so you can:
- Keep multiple drafts in progress simultaneously
- Resume a paused article weeks later without context thrashing
- Clone a successful 01-topic into 02-research without moving files
Why Artifact Directories Use ISO Date Prefixes
Each work session gets a unique subdirectory (2026-06-03-deployment-patterns) so the system can:
- Maintain audit trail of all drafts and rejected outlines
- Reference prior research without merge conflicts
- Roll back to yesterday's draft if today's edits don't land
Token Budget Impact
Original per-session