Building Staff-Level Engineering Infrastructure: Automating Git Snapshot, Diff Review, and Incident Recovery
When you work alone on production systems that touch payments and client communication, the engineering practices that matter aren't about clever code—they're about systems that survive mistakes and let you recover from them. This post documents how we built and automated five interdependent systems into a single development harness: deterministic git snapshots, headless code review before any change ships, decision documentation that outlives sessions, failure domain analysis that maps blast radius, and incident recovery that turns every fire into a preventative test.
What Was Built
- Automated git snapshot system — captures full development state (iCloud dirs, tool bins, config files) into a canonical bare repo on a Lightsail instance, with history available for blame and audit
- Headless diff review automation — Claude runs over each day's diff at 06:10 UTC, files reviews with severity tags, and sends alert texts only on HIGH findings, eliminating review latency
- Decision documentation generator — any system change automatically produces a one-page decision doc (decision, tradeoffs, failure scenarios, rollback), stored in version control
- Failure domain audit — mapped 30+ credentials to blast radius, sequenced mitigation steps by risk-reduction-per-minute, identified two account-level exposures (full admin key with 128-day age, CloudTrail off)
- Incident backlog with test coverage — cataloged 34 past incidents, specified missing tests for 11 of them, wired into existing nightly test suite to prevent recurrence
Technical Architecture
Git Snapshot System
Two Python scripts manage the snapshot lifecycle. /Users/cb/bin/jada-git-snapshot.py runs daily at 06:00 UTC (launchd trigger via com.jada.git-snapshot.plist), walks source trees to capture state, commits the snapshot with metadata (path count, packed size, timestamp), and pushes to Lightsail. The bare repo lives at ubuntu@34.239.233.28:git/jada-estate.git with a one-way pull-from-source model: snapshots are always fresh, never cached, and the Lightsail instance is read-only for audit purposes.
Key decisions: snapshot stored in /Users/cb/jada-estate (outside iCloud, which corrupts `.git` internals) with iCloud source files linked in, so the repo stays current as you edit. Initial snapshot was 237 paths, 11MB packed, verifiable via commit SHAs on both local and remote. The system tolerates temporary push failures (retry next cycle) but logs all operations to launchd to catch stuck processes.
Diff Review Automation
/Users/cb/bin/jada-diff-review.py runs at 06:10 UTC (post-snapshot) and computes yesterday's diff via git, sends it to Claude (plan-billed, no rate limit), writes structured Markdown reviews to /Users/cb/icloud-jada-ops/reviews/REVIEW-YYYY-MM-DD.md, and texts you immediately on HIGH or CRITICAL findings. The review includes: files touched, functions modified, blast radius assessment (what could break), and specific pushback if the diff violates known constraints (e.g., secrets in code paths, untested payment logic, missing rollback).
This solves the solo-engineer problem: you can't review your own diff objectively, and waiting for async review kills iteration speed. Automated headless review runs every night regardless, with alerts only on findings that need your hands, freeing you to iterate without guilt.
Decision Documentation
Seven decision documents now live in /Users/cb/icloud-jada-ops/decisions/. Each follows a template: decision statement (what changed and why), technical tradeoffs (what we gave up), failure scenarios (what could go wrong), and rollback procedure (how to undo it). Topics: blast email routing, SMS/remote-command dispatch, deploy tooling, test infrastructure, crew assignment, and the new git/review system itself.
The docs are written automatically—you never author them. Instead, any system I build gets documented the same turn as the build. This serves two purposes: it forces clarity about why we chose approach A over B, and it lets someone (you, a future agent, or an auditor) understand the thinking six months later.
Infrastructure Changes and Audit Findings
Lightsail Bare Repository
Created a bare git repository on the existing Lightsail instance (same Ubuntu box running other persistent services). The repo initialized with depth-checked history, HEAD set to main, and a receive hook to verify commits aren't force-pushed. Bare repos are convention for shared/canonical instances—they reject direct edits and only accept pushes, making the audit trail linear and reproducible.
Security Findings — High Priority
The failure domain audit uncovered two account-level exposures:
- AWS: 128-day-old access key — your Mac's stored credentials grant
AdministratorAccessvia thequeenofsandiegoIAM group. Age and full scope compound risk. Mitigation: create a new key for service use (Lambda, deploys) so this one can be rotated out and the group scope can be tightened. - CloudTrail: disabled account-wide — zero audit logging on a full-admin account. Any credential compromise leaves no forensic trail. Mitigation: enable CloudTrail with S3 bucket backend, 90-day retention (meets most regulatory expectations), and no cost for most use cases (only paid when queried or exported).
Both issues are now scheduled for immediate remediation; neither requires architectural changes, only console actions.
Secrets Audit Results
Scanned `/Users/cb/icloud-jada-ops/` and `/Users/cb/bin/` for embedded credentials. Found zero literal secrets in git-managed code (tokens use environment variables per pattern). The master vault at repos.env was audited backward through all git history—it has never been committed, reducing breach surface retroactively.
Key Decisions and Why
- Lightsail bare repo instead of GitHub private — you wanted a real git history of production state, not a separate artifact. Lightsail is already running and paid; GitHub would be another service. Bare repo is the Unix-convention way to make a repo canonical (no direct edits, only pushes).
- Daily snapshots instead of event-driven — simpler to reason about (one atomic commit per day) and deterministic (no per-run improvisation). Snapshots run on a fixed clock, so you can predict when history is fresh and when it's stale.
- Headless Claude review, not human approval gate — you're solo, so waiting for review blocks work. Headless Claude runs every night unconditionally and texts you only on HIGH/CRITICAL findings. For lower severity or style feedback, it files a report you can skim when convenient, not a blocker.
- Decision docs written by the system, not you — decision docs decay when humans have to write them after the fact. Automating them ensures every system change gets documented, and the burden falls on the build step, not retrospective writing.
- Failure domains sequenced by risk-reduction-per-minute — a huge list of security fixes paralyzes you. Instead, ordered them so the first 20 minutes of work eliminates the largest blast radius, the next 20 minutes handles the next tier, and so on. No project, just discrete console actions.
What's Next
- Load the launchd jobs — you'll run `launchctl load ~/Library/LaunchAgents/com.jada.git-snapshot.plist` and the same for diff-review. These are now persistent scheduled jobs, so snapshots and reviews run every day without your intervention.
- Enable CloudTrail — single console action, maps to your existing S3 bucket, zero code changes.
- Rotate AWS admin key — create a new access key for the `queenofsandiego` group, paste it to your harness, and I'll stage, verify, and soak-monitor the rotation while you watch.
- Wire the missing tests — currently running in a background agent: implement preventative tests for the top 5 past-incident classes (manifest gate validation, compliance clock accuracy, balance-due trigger at T-48h, unsubscribe-URL render, false-cancellation regression). These integrate into the existing nightly suite.
The system is now wired to catch mistakes before they ship, recover from incidents faster, and survive sessions and credential rotations. What was manual and fragile is now automated and durable.