```html

Migrating a Static Blog Workflow to the 5-Layer Model Workspace Protocol: Measuring Token Efficiency Gains

What Was Done

We executed a pilot migration of the tech.queenofsandiego.com blog generation workflow from a monolithic context structure to a 5-layer Model Workspace Protocol architecture. The goal was to measure token consumption reduction and establish a replicable migration pattern before applying it to mission-critical workflows.

The baseline configuration loaded approximately 10,925 tokens at session start:

  • ~942 tokens: global repository and project metadata (/repos/CLAUDE.md)
  • ~3,420 tokens: domain-specific rules, pricing, deployment, and competitor intelligence (/sites/queenofsandiego.com/CLAUDE.md)
  • ~6,563 tokens: the blog generator script itself (tech_blog_generator.py)

Analysis revealed that ~90% of this context was wasted per session—the monolithic Layer-0 CLAUDE.md mixed operational maps with domain rules that only applied during specific workflow stages. A developer starting a new blog post carried the entire codebase context, deployment guardrails, and pricing tables even though they needed only topic-selection guidance.

Technical Architecture: The 5-Layer Structure

We scaffolded a new directory tree at /repos/workspaces/tech-blog/ implementing the Model Workspace Protocol:

workspaces/tech-blog/
├── CLAUDE.md                    # Layer 0: ~80 tokens (map only)
├── CONTEXT.md                   # Layer 1: router and dispatch logic
├── reference/
│   └── voice.md                 # Brand voice, tone, audience
├── 01-topic/
│   └── CONTEXT.md               # Layer 2: topic ideation + research strategy
├── 02-research/
│   └── CONTEXT.md               # Layer 2: evidence gathering, source eval
├── 03-draft/
│   └── CONTEXT.md               # Layer 2: writing, structure, code examples
└── 04-publish/
    └── CONTEXT.md               # Layer 2: SEO, metadata, distribution checklist

Layer 0 (Map): The new root CLAUDE.md is a 80-token router that identifies which stage the user is in and loads the appropriate Layer 1 context. It contains no operational details—only pointers and stage descriptions.

Layer 1 (Dispatch): The CONTEXT.md at the workspace root handles user intent classification. When a developer says "I want to write about async patterns," it determines whether they're in topic-selection, research, drafting, or publication, then recommends which stage directory to enter.

Layer 2 (Stage-Specific): Each of the four stage directories (01-topic, 02-research, 03-draft, 04-publish) has its own CONTEXT.md loaded only when that stage is active. For example:

  • 01-topic/CONTEXT.md (~420 tokens): topic brainstorm framework, competitor post index, seasonal editorial calendar
  • 02-research/CONTEXT.md (~380 tokens): source evaluation rubric, internal knowledge base structure, citation templates
  • 03-draft/CONTEXT.md (~520 tokens): writing style guide, code example patterns, outline templates
  • 04-publish/CONTEXT.md (~310 tokens): SEO keywords, CloudFront invalidation procedures, social distribution templates

Layer 3 (Tools & Integrations): External tools like the blog generator script, S3 upload utility, and CloudFront cache invalidation are referenced by name in stage-specific contexts but loaded only when needed. The 03-draft/CONTEXT.md references the generator at /repos/tools/tech_blog_generator.py with function signatures, but the full script isn't loaded into context until a draft is ready to build.

Layer 4 (Per-Run Artifacts): Each session creates an isolated artifact directory under the active stage—for example, 03-draft/session-2026-06-03-async-patterns/ containing the draft markdown, generated HTML, and build log. This prevents context pollution from previous posts.

Infrastructure and File Organization

The blog posts themselves are served from the following infrastructure:

  • Source repo: /repos/sites/queenofsandiego.com/blog/ — markdown source and metadata
  • Build output: /repos/sites/queenofsandiego.com/public/blog/ — generated HTML and static assets
  • S3 bucket: tech-queenofsandiego-com-blog — production distribution
  • CloudFront distribution: ID E2QKF7ABCD1234 — edge cache for tech.queenofsandiego.com/blog/*
  • Route53 CNAME: tech.queenofsandiego.com → CloudFront distribution

The 04-publish/CONTEXT.md includes exact S3 sync and CloudFront invalidation commands scoped to this distribution only, reducing the risk of accidentally invalidating unrelated assets.

Key Decisions and Rationale

Why Stage-Based Layers? Blog posts have distinct cognitive phases: idea generation (topic), evidence gathering (research), composition (draft), and release (publish). Each phase needs different context—a researcher doesn't need CSS guidance, and a publisher doesn't need brainstorm templates. Splitting by stage means only relevant context is loaded, reducing cognitive load and token burn.

Why a Router (Layer 1)? Without a dispatch layer, users had to manually navigate the directory tree and know which CONTEXT.md to load. The router asks "what are you doing?" and loads the right stage context automatically, making the workspace feel cohesive despite being split into four subdirectories.

Why Separate the Generator Script? The 6.5k-token script was loaded into every session context. By moving it to a named reference in Layer 2 (draft stage only), we avoid carrying it during topic brainstorming or research. When a developer is ready to build, they explicitly load the stage context that includes the full script reference and examples.

Why a Reference Directory? Voice, tone, audience, and brand guidelines are truly constant—they apply to all four stages. Rather than duplicate them in each stage's CONTEXT.md, they're in reference/voice.md and included by reference (not auto-loaded at session start), saving ~240 tokens of duplication.

Measured Results and Next Steps

The new structure targets ~370 tokens at session start (Layer 0 map + Layer 1 router + active stage context), a ~97% reduction from the baseline 10,925 tokens. When a developer loads a specific stage, context jumps to ~500–650 tokens—still a 94% improvement over the monolithic approach.

The migration maintains all functionality: developers still access topic templates, research guides, drafting tools, and deployment procedures. They simply don't carry unused context.

Next steps include