Migrating a Monolithic Blog Workflow to 5-Layer Model Workspace Protocol: Token Efficiency at Scale
What Was Done
We converted the tech.queenofsandiego.com blog generator from a single-file context architecture to a layered Model Workspace Protocol structure. The pilot migrated a 10,925-token baseline (90% waste) down to a target of ~370 tokens per session—a 30× reduction in always-on context overhead.
This wasn't a feature change; it was a structural refactor to eliminate token leakage. The old system auto-loaded an 80KB CLAUDE.md file containing business rules, deployment procedures, competitor pricing, and voice guidelines all in one Layer 0 context, regardless of whether the user was researching a topic, drafting an outline, or publishing live.
The Problem: Monolithic Context Waste
The baseline measurement revealed the cost structure:
- ~942 tokens: Cross-site
/repos/CLAUDE.md(organizational layer) - ~3,420 tokens:
queenofsandiego.com/CLAUDE.md(business layer—pricing, safety rules, deploy guardrails) - ~6,563 tokens:
tech_blog_generator.py(workflow orchestrator) - Total: 10,925 tokens loaded on every message, even for topic brainstorming
The architectural sin: Layer 0 (map/navigation) was fused with Layer 3 (execution rules). A user starting a blog post in the research phase got access to deploy-safety procedures they couldn't use. A user publishing loaded business logic irrelevant to their current task.
Technical Architecture: 5-Layer Separation
We restructured the workspace into discrete layers with file tree:
workspaces/tech-blog/
├── CLAUDE.md [Layer 0: ~80 tok — router only]
├── CONTEXT.md [Layer 1: ~200 tok — phase selector]
├── reference/
│ └── voice.md [Layer 2: ~150 tok — brand voice]
├── 01-topic/
│ ├── CONTEXT.md [Layer 2: topic ideation]
│ └── artifacts/ [per-run outputs]
├── 02-research/
│ ├── CONTEXT.md [Layer 2: research methodology]
│ └── artifacts/
├── 03-draft/
│ ├── CONTEXT.md [Layer 2: outline + draft rules]
│ └── artifacts/
└── 04-publish/
├── CONTEXT.md [Layer 2: formatting, deploy checks]
└── artifacts/
Layer 0 CLAUDE.md (~80 tokens): Pure navigation. Nothing but file tree, layer definitions, and a decision tree for which CONTEXT.md to load. No business logic.
Layer 1 CONTEXT.md (~200 tokens): Phase router. Asks "which stage?" and loads the appropriate stage CONTEXT.md. Stays resident; orchestrates jumps between stages.
Layer 2 CONTEXT.md files (per-stage, ~100–150 tokens each): Task-specific rules only. 01-topic/CONTEXT.md has idea-generation prompts and brand positioning; 04-publish/CONTEXT.md has deploy checklist and CloudFront cache invalidation syntax. Never both.
Layer 3 reference/voice.md (~150 tokens): Loaded on demand by individual stage contexts. Brand voice, tone, writing examples—referenced but not always resident.
Layer 4 artifacts/: Outputs from each stage. Outlines, drafts, final HTML. Loaded only when iterating within a stage.
Implementation Details
Root CLAUDE.md excerpt:
# tech-blog workspace — Layer 0 router
You are writing technical blog posts for tech.sailjada.com.
## Current Stage
Use /stage CONTEXT to set context. Valid: 01-topic, 02-research, 03-draft, 04-publish.
## Architecture
- Layer 0: This file (navigation only)
- Layer 1: ./CONTEXT.md (phase orchestrator)
- Layer 2: ./{stage}/CONTEXT.md (task rules)
- Layer 3: ./reference/voice.md (brand rules, on-demand)
- Layer 4: ./{stage}/artifacts/ (run outputs)
## Navigation
If unsure, ask: "What stage are you in?"
Load only the relevant stage CONTEXT.md.
Stage context example (03-draft/CONTEXT.md):
# Draft Stage — Blog Post Composition
You are drafting a blog post outline and first draft in HTML format.
## Input
- Topic from 01-topic/artifacts/
- Research summary from 02-research/artifacts/
## Output Rules
- HTML – structure only
- Code blocks in with language hints
- No , , wrappers
- Max 1200 words
- Technical language for engineers; explain *why*, not just what
## File Paths to Reference
- ../reference/voice.md for tone and examples
- ../02-research/artifacts/ for research outputs
Infrastructure: Deployment & Static Site
The tech blog lives at tech.sailjada.com, served via:
- S3 bucket:
tech-sailjada-com-static (us-west-2)
- CloudFront distribution:
E2X9EXAMPLE1ABC (root domain + wildcard)
- Route53 hosted zone:
sailjada.com (Z0ABC1EXAMPLE2XY)
- ALIAS record:
tech.sailjada.com → CloudFront distribution DNS
Published posts are output as static HTML files, checksummed for integrity, and synced to S3 with:
aws s3 sync ./artifacts/published/ s3://tech-sailjada-com-static/posts/ \
--delete \
--cache-control "public, max-age=3600"
Cache invalidation on publish (necessary to bypass CloudFront):
aws cloudfront create-invalidation \
--distribution-id E2X9EXAMPLE1ABC \
--paths "/posts/*"
Key Decisions: Why This Structure
- No monolithic CLAUDE.md: Each stage loads only what it needs. A research phase doesn't load deploy-safety rules that fire based on keywords in the prompt, introducing hallucination risk.
- Layer 1 as permanent fixture: Phase router stays resident so the user (or agent) can jump between stages without reloading the entire workspace. The overhead is ~200 tokens—the cost of a navigation layer—vs. 3,420 for full business context.
- reference/ for shared rules: Brand voice, tone, and examples live in one file. Each stage
CONTEXT.md says "load ../reference/voice.md" when relevant. This way,