Estate Verification Architecture: Building Trust in Distributed File Systems
One of the core challenges in managing a large distributed estate of files, deployments, and artifacts is answering a simple but critical question: does what we claim exist actually exist? This post covers the verification pipeline we built to cross-check day-logs against the real filesystem, and the patterns we're using to keep artifact claims synchronized with reality.
The Problem: Claim vs. Reality Divergence
We maintain detailed logs of daily work sessions—what was completed, what artifacts were produced, where they were saved. In a complex estate spanning multiple storage roots, S3 buckets, and managed directories, there's always a gap between what we recorded we did and what actually exists on disk or in cloud storage. This gap grows silently until you need that artifact and discover it doesn't exist.
Our day-logs live in /Users/cb/icloud-jada-ops/logs/ with entries like:
## Session: 10:30-12:15 — Marketing Copywriting
**Output**: report in `brand/marketing-copy-audit-jul03.md`
## Session: 14:45-16:00 — Site Deploy (tech.sailjada.com)
**Output**: CloudFront cache invalidation for `*.html` files; deployed to S3 `sailjada-web-prod`
The question: are these claims true?
Solution: Cross-Check Pipeline
We built a three-tool verification pattern:
- estate_map: Enumerate the complete filesystem topology at known roots (e.g.,
/Users/cb/icloud-jada-ops/,/Users/cb/dablio/), providing baseline structure - estate_search: Full-text filename search across all registered roots; used to locate artifacts by name (e.g., search for
marketing-copy-audit-jul03,MOATS-ANALYSIS-JUL03) - estate_read: Read and parse specific files or directories to inspect content structure and metadata
The verification flow for each claimed artifact is straightforward:
FOR EACH session in day_log:
FOR EACH claimed artifact in session:
IF artifact_type == file:
search_results = estate_search(artifact_name)
IF search_results.empty():
MARK "MISSING" with context
ELSE:
verify_path = search_results[0].path
content = estate_read(verify_path)
MARK "VERIFIED" with path and metadata
IF artifact_type == deployment:
manifest = estate_read(/state/deployment-manifest.json)
IF manifest.s3_bucket == claimed_bucket AND manifest.timestamp.date == session_date:
MARK "VERIFIED" with deployment details
ELSE:
MARK "INCONSISTENT"
Verification Results: 2026-07-03
We ran cross-checks on the complete 2026-07-03 day-log. Results:
- 5 sessions verified
- 4/5 artifacts confirmed present
- 1 missing artifact detected
Missing: Marketing copywriting report claimed in brand/marketing-copy-audit-jul03.md. The file doesn't exist in /Users/cb/icloud-jada-ops/brand/ (which contains only README.md, sailjada-homepage.html, testimonials.md).
Verified artifacts:
- MOATS analysis report:
/Users/cb/icloud-jada-ops/compliance/MOATS-ANALYSIS-JUL03.md✓ - Site deployment to
sailjada-web-prodS3 bucket with CloudFront distributionE1A2B3C4D5E6F7, cache invalidation pattern/*.html, deployed 2026-07-03T15:47:22Z ✓ - Estate audit report:
/Users/cb/icloud-jada-ops/AUDIT-2026-07-03-full.md✓ - Jett birthday coordination folder:
/Users/cb/icloud-jada-ops/2026-07-10-jett/✓
Infrastructure Details: How Deployments Are Tracked
For deployments (like the site update), we maintain a deployment manifest at /Users/cb/icloud-jada-ops/state/deployment-manifest.json that grounds truth in metadata:
{
"domain": "tech.sailjada.com",
"s3_bucket": "sailjada-web-prod",
"cloudfront_dist_id": "E1A2B3C4D5E6F7",
"last_deployed": "2026-07-03T15:47:22Z",
"cache_invalidation_pattern": "/*.html",
"git_commit": "abc123def456"
}
This structure allows us to verify that:
- The S3 bucket name matches the claim
- The CloudFront distribution ID is correctly registered (critical for cache invalidation automation)
- The invalidation pattern (e.g.,
/*.htmlfor full-site HTML refresh) is enforced - The deployment is time-stamped and linked to a git commit for audit trails
Key Design Decisions
Why search-by-filename instead of hardcoded paths? The estate has multiple roots and evolved storage layout. Requiring precise paths creates brittleness; searching allows flexibility. We verify results by reading content, not just trusting filename matches.
Why a manifest file for deployments? It's the single source of truth for infrastructure configuration. No need to cross-check S3 directly or query AWS API—the manifest is auditable, versionable, and lives in the estate filesystem. It's read-only from our side; external deployment automation updates it on push.
Why report missing artifacts as an action item, not an error? A missing file could mean: file was never generated (session incomplete), saved elsewhere (user error), or legitimately not needed (plan changed). We flag it for investigation rather than failing hard.
What's Next
We're integrating this verification pipeline into nightly automation to catch divergence early. Next steps include adding write-time artifact registration (when a session completes, require explicit artifact claim in structured format) and linking to session transcripts for root-cause analysis when claims don't match reality.
The full cross-check report was saved to /Users/cb/.claude/reports/daylog-crosscheck-20260704.md and is ready for handoff to the ops queue.