```html

Migrating a 10K-Token Blog Workflow to 5-Layer Model Workspace Protocol: Baseline Measurement and Pilot Design

What Was Done

We identified and measured a significant token-efficiency problem in the tech.queenofsandiego.com blog generation workflow, then designed a 5-layer Model Workspace Protocol migration to reduce per-session context overhead by ~97% (from ~10,925 tokens to a target ~370 tokens). This post documents the baseline measurement methodology, the architectural redesign, and the decision framework that makes this workflow the ideal pilot for generalizing the pattern across other mission-critical systems.

The Problem: Baseline Token Audit

The existing blog workflow operated from a monolithic context structure. We measured three distinct sources of always-on context load:

  • Repos root CLAUDE.md (~942 tokens): Core developer instructions, tool definitions, and project scaffolding rules. This file is loaded by every session across all projects.
  • Queen of San Diego site CLAUDE.md (~3,420 tokens): Deployment safety rules, pricing constraints, competitor analysis, and business logic. Every blog post session auto-loaded this entire file, though only ~200 tokens were blog-specific.
  • tech_blog_generator.py (~6,563 tokens): The Python generator script itself, read into context for every session to provide tool definitions and examples.

Total baseline: ~10,925 tokens per session start. Approximately 90% waste.

The root cause: Layer 0 (project navigation), Layer 1 (routing), Layer 2 (stage-specific rules), and Layer 3 (business constraints) were all mixed into two CLAUDE.md files. Sessions needed to load layers they weren't currently using, creating false dependencies and preventing parallelization across multiple blog drafts.

Architecture: 5-Layer Model Workspace Protocol

We designed a new directory structure under repos/workspaces/tech-blog/ following this hierarchy:

repos/workspaces/tech-blog/
├── CLAUDE.md                    # Layer 0: ~80 tokens (nav only)
├── CONTEXT.md                   # Layer 1: router + task dispatch
├── reference/
│   └── voice.md                 # Style guide, examples, tone
├── 01-topic/
│   └── CONTEXT.md               # Layer 2: topic ideation rules
├── 02-research/
│   └── CONTEXT.md               # Layer 2: research methodology
├── 03-draft/
│   └── CONTEXT.md               # Layer 2: writing + review gates
├── 04-publish/
│   └── CONTEXT.md               # Layer 2: deployment + SEO rules
└── [per-run artifacts]/
    ├── topic_brief.md
    ├── research_notes.md
    ├── draft_v1.md
    └── final_output.md

Layer assignments:

  • Layer 0 (CLAUDE.md): Minimal map. Names the four stages, describes the CONTEXT.md router pattern, and explains when to load each stage's CONTEXT.md. ~80 tokens.
  • Layer 1 (CONTEXT.md): Task router and dispatcher. Reads user input, determines which stage(s) are relevant, and loads only those stage CONTEXT.md files. Handles handoff between stages.
  • Layer 2 (four stage CONTEXT.md files): Stage-specific rules, examples, guardrails, and success criteria. 01-topic handles ideation and trend analysis. 02-research handles source evaluation and fact-checking. 03-draft handles prose style, voice consistency, and internal review. 04-publish handles deployment, metadata, and monitoring.
  • Layer 3 (reference/voice.md): Durable style guide, tone examples, and audience profile. Loaded only when drafting or reviewing.
  • Layer 4 (per-run artifacts): Session outputs. Each blog post session produces its own artifact directory, enabling parallel work and preventing context pollution across concurrent drafts.

Why This Structure Works

Selective Loading: A session starting at the "topic" stage loads Layer 0 (~80 tokens) + Layer 1 (~300 tokens) + 01-topic CONTEXT.md (~200 tokens) = ~580 tokens. The same session does not load deploy-safety rules, pricing constraints, or competitor analysis. Those remain in Layer 3, loaded only at the 04-publish stage.

Parallelization: Two drafts can run concurrently without context collision. Each gets its own artifact directory. The Layer 1 router prevents race conditions on the stage CONTEXT.md files themselves (which are read-only during a session).

Reusability: The 5-layer pattern isn't blog-specific. Once validated here, the same structure applies to the deposit/booking workflow, the QoS email templates, the SCC admin dashboard, and the ticket-runner itself.

Low Risk: The blog workflow is non-mission-critical. Failures during the pilot don't impact revenue, customer access, or core systems. It's the ideal testbed.

Measurement Plan

We will measure three scenarios:

  • Baseline (old): Start a blog post session with the monolithic CLAUDE.md. Record token load via Claude API usage logs.
  • Pilot (new): Start an identical scenario using the 5-layer structure. Measure the same way.
  • Real-world (new): Complete 3–5 actual blog post cycles (topic → research → draft → publish) using the new structure. Record wall-clock time, number of context switches, and cumulative tokens per post.

Success criteria: <500 tokens per session start (real-world average), zero breakage, no loss of quality or speed compared to the baseline.

Infrastructure and Tooling

The new structure lives in the existing Git repo at repos/workspaces/tech-blog/. Deployment is file-only—no infrastructure changes needed. The tech_blog_generator.py tool remains in repos/tools/tech_blog_generator.py but will be refactored to:

  • Read the Layer 1 router (CONTEXT.md) to determine stage.
  • Inject stage-specific CONTEXT.md into the prompt dynamically, rather than always including the full generator script.
  • Route outputs to per-run artifact directories using timestamped naming.

No API changes, no new Lambda functions, no CloudFront invalidations, no Route53 updates.

Key Decisions

  • Why nested CONTEXT.md over single generator config? CONTEXT.md integrates stage rules with examples and reasoning, making edits human-friendly. A JSON config would require dual maintenance (config + examples in CLAUDE.md).
  • Why four stages? They map to distinct mental models and skill sets (ideation, research, writing, ops). Clear boundaries make it easier to gate approvals and handoff.
  • Why measure wall-clock and tokens separately? Token count ≠ perceived latency. A 370-token session might feel faster or slower depending on which LLM tier is being used. We need both metrics to make billing and UX claims.
  • Why not migrate everything at once? The