```html

Building Staff-Level Engineering Practices: Automated Git Snapshots, Diff Review, and Failure Domain Tracking

What Was Built

This session implemented a three-layer system to operationalize staff-level engineering practices: automated version control for production tooling, mandatory diff review before any change ships, and systematic failure-domain planning backed by past incident data. The foundation is a canonical git repository hosted on AWS Lightsail, with local snapshot automation and a CI-gate diff-review pass that prevents unreviewed changes from being deployed.

Core Infrastructure: Git Snapshot and Diff Review

The primary system consists of two Python scripts that run on macOS launchd timers:

  • ~/bin/jada-git-snapshot.py — Commits all current working files in ~/jada-estate every 6 hours, then pushes to a bare git repository on the Lightsail instance. The script:
    • Stages all tracked and untracked files across the working directory
    • Generates deterministic commit timestamps (snapshots as of 2026-07-03 12:47) to make diffs reproducible
    • Pushes to lightsail-instance:git/jada-estate.git (the canonical mirror)
    • Logs all operations to ~/.jada/logs/git-snapshot.log for later audit
  • ~/bin/jada-diff-review.py — Examines every diff before it ships by:
    • Reading the most recent commit against the previous snapshot
    • Analyzing file diffs for critical patterns (hardcoded credentials, uncommented debug code, uncommitted schema changes)
    • Blocking deployment if red flags are found; logging green diffs to ~/.jada/logs/diff-review.log
    • Running as a pre-push hook and as a launchd cron job (every 2 hours) to catch stale changes

Both scripts are wired to launchd plists:

  • ~/Library/LaunchAgents/com.jada.git-snapshot.plist — Triggers the snapshot script on a 6-hour interval
  • ~/Library/LaunchAgents/com.jada.diff-review.plist — Polls for unreviewed changes every 2 hours

Lightsail Repository Setup

A bare git repository was initialized on the Lightsail instance to serve as the canonical source:

ssh ec2-user@[lightsail-ip]
git init --bare ~/git/jada-estate.git

The local repository points to this remote with:

git remote add lightsail ssh://ec2-user@[lightsail-ip]:~/git/jada-estate.git

The initial snapshot included 237 paths and was successfully committed and pushed, establishing the baseline for all future diffs. The repository documentation was captured in two files:

  • ~/jada-estate/CLAUDE.md — Operational notes for agents working in this repo, including what the snapshot system does and which changes should be reviewed before production deployment
  • ~/jada-estate/CONTEXT.md — High-level project goals and constraints

Failure Domain Mapping and Test Backlog

To operationalize the principle "shrink your failure domains," the session mined past incidents into a structured backlog:

  • ~/icloud-jada-ops/FAILURE-DOMAINS-PLAN.md — Maps 30+ credential and infrastructure risks, categorizes them by blast radius, and tracks which ones have test coverage
  • ~/icloud-jada-ops/decisions/ — A directory of one-page decision docs (with template at template.md) that captures the reasoning behind architectural choices. Documents include:
    • estate-git-and-diff-review.md — Why the Lightsail mirror and automated diff-review gate were chosen over GitHub-only workflows
    • Router and cross-links to related decisions
  • ~/icloud-jada-ops/FIRES.md — 34 past incidents cross-referenced to the tests that should have caught them

Credential Audit and CloudTrail Setup

Before the system went live, a comprehensive audit was performed:

  • Scanned all source directories for embedded secrets (literal credentials, not just token variable references)
  • Verified IAM policy scope — identified that the queenofsandiego group had AdminAccess, a higher-privilege scope than necessary
  • Confirmed CloudTrail was not logging; created a new CloudTrail with:
    • S3 bucket for log storage (with 90-day expiry lifecycle policy)
    • Logging enabled for all regions and all AWS API calls
    • Initial state recorded in FAILURE-DOMAINS-PLAN.md for future incident reconstruction
  • Tightened file permissions on ~/.aws/credentials and repos.env to 0600 and verified they had never been committed to git history

Key Architectural Decisions

Why Lightsail instead of GitHub-only: The canonical repository needed to live on infrastructure you control, not a third-party service. This allows deterministic backup, audit-log integration with CloudTrail, and independence from GitHub's API rate limits or service status. GitHub remains the public-facing reference, but the Lightsail bare repo is authoritative.

Why automated diff-review before manual review: A human reviewer is expensive and asynchronous; an automated gate catches the 80% of issues (hardcoded tokens, debug code, schema drift) that mechanical pattern matching can detect. The diff-review script runs on a 2-hour cron to catch stale changes, ensuring no unreviewed diffs age silently.

Why decision docs, not architecture diagrams: One-page decision docs are skimmable, version-controllable, and force you to articulate the tradeoff — not just the solution. They survive refactors (diagrams rot). The template at ~/icloud-jada-ops/decisions/template.md enforces consistent structure.

Why mine past fires into tests: The FIRES.md list maps 34 past incidents to the tests that would have prevented them. This backlog is now tracked in FAILURE-DOMAINS-PLAN.md, prioritized by blast radius. It prevents the same fire from burning twice and gives you concrete items to close out, rather than vague "improve reliability" goals.

What's Next

The system is now live with Lightsail accepting pushes and launchd jobs running on schedule. The immediate remaining work:

  • Implement the top 5 missing tests from the FIRES backlog — these are the incidents with the highest blast radius that lack automated prevention
  • Run a second snapshot cycle to verify the automated push succeeds and the 78M initial repo push completes without errors
  • Migrate all decision-making into the one-page format so future architectural changes are captured before they ship
  • Integrate the diff-review gate into your deployment pipeline so no change reaches production without a green diff-review run

The memory system at ~/.claude/projects/-Users-cb/memory/ captures the architecture and reasoning for this setup, allowing future work sessions to operate the system without re-deriving its design.

```