```html

Implementing a 5-Layer Model Workspace Protocol for the Tech Blog: Tokenomics, Architecture, and Migration Strategy

Over the past session, we conducted a comprehensive audit of the tech.queenofsandiego.com blog workflow and designed a 5-layer context isolation protocol to reduce per-session token consumption by ~97% (from ~10,925 tokens to a target ~370 tokens). This post walks through the problem, the architecture, and the migration mechanics.

The Problem: Bloated Always-On Context

The existing blog generator pipeline loaded a monolithic CLAUDE.md file (~6,563 tokens) into every session, along with inherited parent context from the Queen of San Diego site (~3,420 tokens) and repository-wide governance rules (~942 tokens). Of that ~10,925 token baseline:

  • ~90% was waste — the giant CLAUDE.md mixed Layer 0 (navigation map) with Layer 3 (pricing rules, deployment safety constraints, competitor analysis). A topic-research session needed none of the deploy-safety rules; a publish session needed none of the competitor research.
  • No stage awareness — the generator didn't distinguish between topic-discovery, research, draft, and publish phases, so each loaded identical context.
  • Coupled governance — shared business rules (pricing, brand voice, deployment safety) were hard-coded into the blog CLAUDE.md instead of referenced from a single source of truth.

For a high-volume blog (target: 2–3 posts/week), this overhead alone could consume 15k–20k tokens/week on context setup alone, before writing a single word.

The Solution: 5-Layer Model Workspace Protocol

We designed a hierarchical context structure that collapses all shared governance into Layer 1 routers and distributes only stage-specific context to each phase:

icloud-repos/workspaces/tech-blog/
├── CLAUDE.md                      # Layer 0: navigation map (~80 tokens)
├── CONTEXT.md                     # Layer 1: router + shared references (~120 tokens)
├── reference/
│   └── voice.md                   # Shared brand voice & style (~200 tokens, loaded once)
├── 01-topic/
│   ├── CONTEXT.md                 # Layer 2: topic-discovery phase context
│   └── artifacts/                 # Per-run: brainstorm, outline, competitor list
├── 02-research/
│   ├── CONTEXT.md                 # Layer 2: research phase context
│   └── artifacts/                 # Per-run: notes, sources, technical deep-dives
├── 03-draft/
│   ├── CONTEXT.md                 # Layer 2: draft phase context
│   └── artifacts/                 # Per-run: outline, sections, embedded code samples
└── 04-publish/
    ├── CONTEXT.md                 # Layer 2: publish phase context
    └── artifacts/                 # Per-run: final MD, frontmatter, images, deployment log

Each layer serves a distinct purpose:

  • Layer 0 (CLAUDE.md) — ~80 tokens. Minimal navigation and "how to use this workspace" guidance. Points to Layer 1.
  • Layer 1 (CONTEXT.md at root) — ~120 tokens. Router that imports shared references and describes the 4-stage workflow. Links to governance (QoS business rules, deploy-safety), voice guide, and stage-specific CONTEXT.md files.
  • Layer 2 (CONTEXT.md per stage) — ~150–200 tokens each. Only the specific prompts, examples, and constraints needed for that phase. E.g., 01-topic/CONTEXT.md includes competitor research templates but not deployment safety rules.
  • Layer 3 (reference/ docs) — Shared, stable guidance (brand voice, code examples, glossaries). Loaded once per session, reused across stages.
  • Artifacts (per-run working files) — Session output: brainstorms, research notes, draft sections, deployment logs. Never auto-loaded; used as input to the next stage.

Token Budget Breakdown (Measured)

Before:

  • repos/CLAUDE.md: ~942 tokens
  • qos/CLAUDE.md (auto-inherited): ~3,420 tokens
  • tech_blog_generator.py (rules + hardcoded context): ~6,563 tokens
  • Per-session total: ~10,925 tokens

After (projected):

  • Layer 0 (CLAUDE.md): ~80 tokens
  • Layer 1 (CONTEXT.md): ~120 tokens
  • Layer 2 (stage CONTEXT.md, e.g., 01-topic): ~180 tokens
  • reference/voice.md (cached across session): ~200 tokens
  • Per-session total: ~580 tokens (cold start); ~370 tokens (warm, with cache)

Savings: ~95% for cold starts, ~97% for warm sessions.

File Structure & Implementation Details

All files live in /Users/cb/icloud-repos/workspaces/tech-blog/ (symlinked from ~/icloud-repos/workspaces/tech-blog/ for local work).

Root CLAUDE.md (~80 tokens):

# tech-blog Workspace

Use this workspace to author posts for tech.sailjada.com. Follow the 4-stage workflow:
1. **01-topic/** — brainstorm topic, scope, audience
2. **02-research/** — gather sources, technical details, code examples
3. **03-draft/** — write first draft, structure sections
4. **04-publish/** — finalize, add metadata, deploy

Load `CONTEXT.md` in the root, then the CONTEXT.md for your current stage.
See `reference/voice.md` for brand voice & style.

Root CONTEXT.md (~120 tokens):

# Tech Blog Workflow Router

## Imports
- **Governance:** See `/Users/cb/icloud-repos/sites/sailjada.com/reference/business.md` for pricing, brand, deploy-safety.
- **Voice:** See `reference/voice.md` for tone, audience, style.

## 4-Stage Workflow
Each stage has its own CONTEXT.md:
- `01-topic/CONTEXT.md` — Topic discovery, competitor analysis, audience mapping
- `02-research/CONTEXT.md` — Source gathering, technical validation, code samples
- `03-draft/CONTEXT.md` — Outline, first draft, section structure
- `04-publish/CONTEXT.md` — Final edits, metadata, deployment procedure

Load the root CONTEXT.md + your stage's CONTEXT.md. Artifacts from stage N become input to stage N+1.

reference/voice.md (~200 tokens):

# Tech Blog Voice & Style

**Audience:** Full-stack engineers, DevOps, platform teams. Assume familiarity with AWS, Docker, CLI tools.

**Tone:** Technical but conversational. Explain the WHY behind decisions, not just the WHAT