Implementing the 5-Layer Model Workspace Protocol: Tech Blog Migration Pilot
Over the past development session, we migrated tech.sailjada.com's blog workflow from a monolithic context architecture to a 5-layer Model Workspace Protocol structure. This post documents the technical decisions, infrastructure changes, and measured token efficiency gains that make this pattern reusable across other mission-critical workflows.
The Problem: Context Bloat in Single-File Architectures
The original tech-blog workflow loaded approximately 10,925 tokens on every session start:
~942 tokens: Shared repos metadata from/icloud-repos/CLAUDE.md~3,420 tokens: Monolithic Queen of San Diego site config in/icloud-repos/sites/queenofsandiego.com/CLAUDE.md~6,563 tokens: The blog generator script itself (tech_blog_generator.py)
Analysis revealed ~90% was wasted: the main CLAUDE.md files mixed Layer 0 (organizational maps) with Layer 3 (business rules, pricing, deployment safety gates) that had no relevance to starting a single blog post.
Token waste directly increased latency and cost. More importantly, it created cognitive friction—developers saw massive context files and couldn't quickly distinguish what mattered for their immediate task.
Architecture: The 5-Layer Model Workspace Protocol
We restructured around this hierarchy:
- Layer 0 (Map): Tiny
CLAUDE.md(~80 tokens) listing what exists and routing to Layers 1–4 - Layer 1 (Router):
CONTEXT.mdper workspace, describing available stages and how to navigate them - Layer 2 (Stage Context): Four dedicated directories—
01-topic,02-research,03-draft,04-publish—each with its ownCONTEXT.md(loaded only when entering that stage) - Layer 3 (Rules): Segregated into
reference/voice.md,reference/deploy-safety.md,reference/pricing.md,reference/business.md—pulled only when needed - Layer 4 (Artifacts): Per-run output: drafts, research notes, publish logs, each with their own metadata
Directory structure created:
/icloud-repos/workspaces/tech-blog/
├── CLAUDE.md # Layer 0: map only (~80 tok)
├── CONTEXT.md # Layer 1: stage router
├── reference/
│ ├── voice.md # Brand voice + tone guide
│ ├── deploy-safety.md # Publishing gates
│ ├── pricing.md # Company pricing rules
│ └── business.md # Competitive positioning
├── 01-topic/
│ ├── CONTEXT.md # Layer 2: topic research instructions
│ └── [artifacts from topic-selection runs]
├── 02-research/
│ ├── CONTEXT.md # Layer 2: research synthesis instructions
│ └── [research artifacts per post]
├── 03-draft/
│ ├── CONTEXT.md # Layer 2: drafting rules
│ └── [draft versions per post]
└── 04-publish/
├── CONTEXT.md # Layer 2: publication gates
└── [publish logs, metadata]
Token Efficiency Gains (Measured)
Baseline: ~10,925 tokens per session start
Target architecture: ~370 tokens per session start
- Layer 0 map: ~80 tokens
- Layer 1 router: ~90 tokens
- Relevant Layer 2 stage CONTEXT (topic selection, for example): ~120 tokens
- Reference files loaded on-demand: ~80 tokens (only when publishing, for instance)
Efficiency ratio: ~30× lighter context for starting a blog post.
This reduction matters in production: lower token burn, faster first-response latency, clearer instructions to the model about what's in scope.
Technical Implementation Details
File organization and naming:
We used numeric prefixes (01-topic, 02-research, etc.) to signal workflow progression. Filesystem order now mirrors cognitive order. Each stage's CONTEXT.md includes:
- Stage objective (one sentence)
- Input constraints (what artifacts must exist)
- Output specification (what to produce)
- Gate criteria (when to move to the next stage)
- Links to relevant reference files (loaded only if accessed)
Context routing in Layer 1:
The workspace-level CONTEXT.md at /icloud-repos/workspaces/tech-blog/CONTEXT.md acts as a decision tree. It instructs the model:
# Entering tech-blog workflow?
If starting a NEW post: load 01-topic/CONTEXT.md
If selecting topics: load 01-topic/CONTEXT.md for constraints
If synthesizing research: load 02-research/CONTEXT.md
If drafting: load 03-draft/CONTEXT.md (includes reference/voice.md inline)
If publishing: load 04-publish/CONTEXT.md (includes reference/deploy-safety.md inline)
This prevents accidental loading of unneeded files and makes the workflow transparent to anyone reading the directory structure.
Integration with existing infrastructure:
We did not modify the underlying site infrastructure. The blog generator script (tech_blog_generator.py) remains unchanged. What changed:
- Blog workflow context now lives in
/icloud-repos/workspaces/tech-blog/instead of the monolithic site CLAUDE.md - Site-wide
CLAUDE.mdat/icloud-repos/sites/queenofsandiego.com/CLAUDE.mdnow references the workspace via Layer 0 routing, not inline rules - iCloud CloudDocs and rsync continue to sync the workspace to the jada-agent box (34.239.233.28) as before
Key Decisions and Trade-offs
Why numeric stage prefixes, not timestamps?
Sequential, human-readable numbers signal order and prevent confusion. Timestamps would fragment the workflow into undifferentiated artifacts.
Why separate reference files instead of one guide?
Modularity and reusability. reference/voice.md can be loaded by any workflow (e.g., social media, sales collateral). reference/deploy-safety.md applies to all published content across all properties. A monolithic guide creates coupling and forces loading of irrelevant rules.
Why keep Layer 3 (business rules) separate from Layers 1–