```html

Implementing a 5-Layer Model Workspace Protocol for Tech Blog Post Generation

What Was Done

We migrated the tech.queenofsandiego.com blog workflow from a monolithic context structure to a 5-layer Model Workspace Protocol architecture. This pilot project reduced initial context load from ~10,925 tokens to a target of ~370 tokens—a 29× improvement—by decomposing a single giant CLAUDE.md into purpose-built files organized across a hierarchical directory structure.

The migration involved:

  • Creating a dedicated workspaces/tech-blog/ directory tree with layer-specific CONTEXT.md files
  • Extracting reusable reference material into reference/voice.md and other modular files
  • Building four stage-gated subdirectories: 01-topic, 02-research, 03-draft, 04-publish
  • Instrumenting the tech_blog_generator.py script to conditionally load only relevant layers
  • Capturing baseline measurements and creating a reproducible migration playbook

Technical Details: The Layer Architecture

Layer 0 — Root Navigation (minimal, ~80 tokens)

The root /workspaces/tech-blog/CLAUDE.md serves as a dispatcher. It does not contain pricing rules, deployment safety guardrails, or competitor research. Instead, it:

  • Identifies the current stage (which subdirectory's CONTEXT.md` we're in)
  • Points to the appropriate Layer-1 router file
  • Lists available stages and their entry points
  • Provides workspace conventions (file naming, artifact location, output formats)

Layer 1 — Conditional Router (~120 tokens)

The /workspaces/tech-blog/CONTEXT.md examines environment variables and file paths to determine which stage's Layer-2 context should be loaded. This router includes:

  • Detection logic: if STAGE == "01-topic" then load ./01-topic/CONTEXT.md
  • Cross-stage dependency notes (e.g., "drafts require research artifacts from 02-research/")
  • Artifact schema for each stage (input files, expected outputs, state location)

Layer 2 — Stage-Specific Context (~200 tokens per stage)

Each stage directory contains its own CONTEXT.md`:

  • 01-topic/CONTEXT.md: Topic brainstorming, keyword research prompts, competitor landscape
  • 02-research/CONTEXT.md: Deep-dive instructions, source evaluation, citation standards, note-taking format
  • 03-draft/CONTEXT.md: Writing voice (loaded from reference/voice.md), structure templates, code example formatting
  • 04-publish/CONTEXT.md: Review checklist, SEO metadata requirements, deployment steps

Layer 3 & 4 — Per-Run Artifacts & Deployment Rules

Formerly baked into the monolithic CLAUDE.md, these now live separately:

  • /reference/voice.md: Brand voice, tone, audience target (write once, reuse across all posts)
  • /reference/deploy-safety.md: CloudFront cache invalidation patterns, Route53 failover triggers, S3 versioning rules
  • Per-post state: /drafts/[slug]/ directory with research notes, outline, draft versions, and metadata

Infrastructure & Tooling Changes

File Structure

workspaces/tech-blog/
├── CLAUDE.md                    # Layer 0: Root navigation (~80 tok)
├── CONTEXT.md                   # Layer 1: Router (~120 tok)
├── reference/
│   ├── voice.md                 # Brand voice guidelines
│   └── deploy-safety.md         # Infrastructure guardrails
├── 01-topic/
│   └── CONTEXT.md               # Layer 2: Topic brainstorming
├── 02-research/
│   └── CONTEXT.md               # Layer 2: Research methodology
├── 03-draft/
│   └── CONTEXT.md               # Layer 2: Writing & structure
├── 04-publish/
│   └── CONTEXT.md               # Layer 2: Review & deployment
└── drafts/
    └── [slug]/
        ├── research-notes.md
        ├── outline.md
        ├── draft-v1.md
        └── metadata.json

Script Integration

The tech_blog_generator.py` was updated to accept a --stage parameter and only load the corresponding CONTEXT.md file:

python tech_blog_generator.py \
  --workspace /repos/workspaces/tech-blog \
  --stage 03-draft \
  --slug my-new-post

This conditional loading is key to achieving the token savings. The old monolithic approach loaded 6,563 tokens of generator code every session; now it lazy-loads only the stage-relevant subset.

Key Decisions & Trade-offs

Why 4 Linear Stages?

The workflow splits into four phases because each has distinct context needs:

  • Topic brainstorming requires competitor research and keyword data but not deployment details
  • Research needs citation standards and source evaluation criteria
  • Drafting needs voice guidelines and formatting rules but not research methodology
  • Publishing needs only the final checklist and deployment automation

A single monolithic context burdened every session with all four, even when only one was relevant.

Why Reference/ Separate from CONTEXT.md?

Voice, deployment safety, and pricing rules change infrequently and are reusable across all posts. Keeping them in reference/ allows:

  • Version control without touching stage-specific CONTEXT.md files
  • Sharing across multiple projects (e.g., if we add a video-tutorial workflow, it can reuse reference/voice.md)
  • Cleaner auditing (all brand/deployment decisions in one place)

Why No Shared "Commons" Layer?

We avoided a Layer 2.5 "shared" context that would load for all stages. Experience showed that shared contexts often become dumping grounds for "might be useful" rules. By forcing explicit imports in each stage's CONTEXT.md, we keep scope tight and make token burn visible.

Measurement & Validation

Baseline (Before)

  • Root CLAUDE.md: 942 tokens (repos metadata + site config)
  • QOS site CLAUDE.md: 3,420 tokens (business rules, pricing, deployment, competitors)
  • tech_blog_generator.py: