Implementing a 5-Layer Model Workspace Protocol for Tech Blog Post Generation
What Was Done
We migrated the tech.queenofsandiego.com blog workflow from a monolithic context structure to a 5-layer Model Workspace Protocol architecture. This pilot project reduced initial context load from ~10,925 tokens to a target of ~370 tokens—a 29× improvement—by decomposing a single giant CLAUDE.md into purpose-built files organized across a hierarchical directory structure.
The migration involved:
- Creating a dedicated
workspaces/tech-blog/directory tree with layer-specificCONTEXT.mdfiles - Extracting reusable reference material into
reference/voice.mdand other modular files - Building four stage-gated subdirectories:
01-topic,02-research,03-draft,04-publish - Instrumenting the
tech_blog_generator.pyscript to conditionally load only relevant layers - Capturing baseline measurements and creating a reproducible migration playbook
Technical Details: The Layer Architecture
Layer 0 — Root Navigation (minimal, ~80 tokens)
The root /workspaces/tech-blog/CLAUDE.md serves as a dispatcher. It does not contain pricing rules, deployment safety guardrails, or competitor research. Instead, it:
- Identifies the current stage (which subdirectory's
CONTEXT.md` we're in) - Points to the appropriate Layer-1 router file
- Lists available stages and their entry points
- Provides workspace conventions (file naming, artifact location, output formats)
Layer 1 — Conditional Router (~120 tokens)
The /workspaces/tech-blog/CONTEXT.md examines environment variables and file paths to determine which stage's Layer-2 context should be loaded. This router includes:
- Detection logic:
if STAGE == "01-topic" then load ./01-topic/CONTEXT.md - Cross-stage dependency notes (e.g., "drafts require research artifacts from 02-research/")
- Artifact schema for each stage (input files, expected outputs, state location)
Layer 2 — Stage-Specific Context (~200 tokens per stage)
Each stage directory contains its own CONTEXT.md`:
01-topic/CONTEXT.md: Topic brainstorming, keyword research prompts, competitor landscape02-research/CONTEXT.md: Deep-dive instructions, source evaluation, citation standards, note-taking format03-draft/CONTEXT.md: Writing voice (loaded fromreference/voice.md), structure templates, code example formatting04-publish/CONTEXT.md: Review checklist, SEO metadata requirements, deployment steps
Layer 3 & 4 — Per-Run Artifacts & Deployment Rules
Formerly baked into the monolithic CLAUDE.md, these now live separately:
/reference/voice.md: Brand voice, tone, audience target (write once, reuse across all posts)/reference/deploy-safety.md: CloudFront cache invalidation patterns, Route53 failover triggers, S3 versioning rules- Per-post state:
/drafts/[slug]/directory with research notes, outline, draft versions, and metadata
Infrastructure & Tooling Changes
File Structure
workspaces/tech-blog/
├── CLAUDE.md # Layer 0: Root navigation (~80 tok)
├── CONTEXT.md # Layer 1: Router (~120 tok)
├── reference/
│ ├── voice.md # Brand voice guidelines
│ └── deploy-safety.md # Infrastructure guardrails
├── 01-topic/
│ └── CONTEXT.md # Layer 2: Topic brainstorming
├── 02-research/
│ └── CONTEXT.md # Layer 2: Research methodology
├── 03-draft/
│ └── CONTEXT.md # Layer 2: Writing & structure
├── 04-publish/
│ └── CONTEXT.md # Layer 2: Review & deployment
└── drafts/
└── [slug]/
├── research-notes.md
├── outline.md
├── draft-v1.md
└── metadata.json
Script Integration
The tech_blog_generator.py` was updated to accept a --stage parameter and only load the corresponding CONTEXT.md file:
python tech_blog_generator.py \
--workspace /repos/workspaces/tech-blog \
--stage 03-draft \
--slug my-new-post
This conditional loading is key to achieving the token savings. The old monolithic approach loaded 6,563 tokens of generator code every session; now it lazy-loads only the stage-relevant subset.
Key Decisions & Trade-offs
Why 4 Linear Stages?
The workflow splits into four phases because each has distinct context needs:
- Topic brainstorming requires competitor research and keyword data but not deployment details
- Research needs citation standards and source evaluation criteria
- Drafting needs voice guidelines and formatting rules but not research methodology
- Publishing needs only the final checklist and deployment automation
A single monolithic context burdened every session with all four, even when only one was relevant.
Why Reference/ Separate from CONTEXT.md?
Voice, deployment safety, and pricing rules change infrequently and are reusable across all posts. Keeping them in reference/ allows:
- Version control without touching stage-specific CONTEXT.md files
- Sharing across multiple projects (e.g., if we add a video-tutorial workflow, it can reuse
reference/voice.md) - Cleaner auditing (all brand/deployment decisions in one place)
Why No Shared "Commons" Layer?
We avoided a Layer 2.5 "shared" context that would load for all stages. Experience showed that shared contexts often become dumping grounds for "might be useful" rules. By forcing explicit imports in each stage's CONTEXT.md, we keep scope tight and make token burn visible.
Measurement & Validation
Baseline (Before)
- Root CLAUDE.md: 942 tokens (repos metadata + site config)
- QOS site CLAUDE.md: 3,420 tokens (business rules, pricing, deployment, competitors)
tech_blog_generator.py: