Putting Production Tooling Under Version Control: Snapshots, Automated Review, and Audit Logging
The Problem
Production-critical tooling—deployment scripts, payment processors, email senders, infrastructure utilities—was living in iCloud directories and ~/bin with no version control, no diff history, no audit trail, and no code review before changes shipped. When a script that moves money or sends client email gets a bad edit, there's no way to answer "what changed," "who changed it," "when," or "how do I revert." Staff-level engineering requires that load-bearing systems be reconstructible and auditable.
What We Built
An end-to-end system to snapshot production tooling into a canonical git repository, automatically review diffs before they land, and log all changes for audit:
- Automated snapshots: Python script that mirrors
~/icloud-jada-ops/tools/and~/bin/into a git repository on a Lightsail instance, running hourly via launchd. - Diff-before-merge gate: A second automation that reviews every diff before it commits, catching mistakes before they reach the canonical repo.
- CloudTrail + S3 audit log: All changes logged to an S3 bucket with 90-day retention and CloudTrail integration, so we know exactly what shipped when.
- Deterministic recovery: Any change can be reverted to a tagged release; the entire history is reconstructible.
Technical Architecture
The Snapshot Pipeline
Every hour, jada-git-snapshot.py runs via launchd job com.jada.git-snapshot.plist:
~/iCloud Drive (tools/, bin/)
↓
jada-git-snapshot.py (copies, excludes secrets, excludes large artifacts)
↓
Local worktree at /Users/cb/jada-estate/
↓
Git commit with timestamp
↓
SSH push to Lightsail bare repo
The snapshot script:
- Copies trees from
~/iCloud\ Drive/icloud-jada-ops/tools/and~/bin/into the worktree. - Excludes compiled binaries, caches, and files over 10MB (checked with
find ... -size) to keep the repo lean. - Scans for embedded secrets (literal AWS keys, Stripe tokens, etc.) using regex patterns and fails loudly if found; secret references (environment variable names) are safe and allowed.
- Compares the new tree against the last committed tree; if there are no changes, the script exits cleanly without creating a commit.
- Commits with a message:
snapshot YYYY-MM-DD HH:MM (N paths), showing exactly what changed. - Pushes to the Lightsail remote with error handling; if the push times out or fails, logs the error and exits (the next hourly run retries).
The launchd job runs at */1 * * * * (hourly) with StandardErrorPath and StandardOutPath` pointing to log files in ~/Library/Logs/jada-git/, so we have a running record of each snapshot.
The Diff-Review Gate
jada-diff-review.py runs every 15 minutes on the local worktree, invoked by com.jada.diff-review.plist:
jada-diff-review.py (runs in /Users/cb/jada-estate/)
↓
git diff origin/main
↓
Scan changes for:
- Modified scripts calling system() / subprocess without shell=False
- New secrets in staged/committed files
- Diffs larger than 500 lines (unusual for incremental snapshots)
↓
If clean: no action (snapshot will push normally)
If dirty: git reset --hard origin/main and log alert
This catches mistakes before they reach the canonical repo. It's deterministic: the same diff always produces the same verdict. It's also non-blocking—if the review fails, the commit is rewound and logged; the next snapshot will try again.
Infrastructure: The Lightsail Bare Repository
On the Lightsail instance (located at the private IP resolved via AWS API), we initialized a bare git repository at a fixed path. This serves as the canonical upstream:
git init --bare /path/to/jada-estate-mirror.git
cd /path/to/jada-estate-mirror.git
git symbolic-ref HEAD refs/heads/main
Why Lightsail over GitHub Public?
- These scripts contain API endpoint paths, S3 bucket names, email templates, and other details that shouldn't be on GitHub (even private).
- A private GitHub repo still requires managing OAuth tokens and deploy keys; Lightsail lets us use SSH keypairs we control directly.
- Lightsail sits in our private network and doesn't expose the repository via a public API.
SSH access is authenticated with a keypair stored on the local machine; the Lightsail security group only allows SSH from the home IP (checked once at setup; this is a known IP, not a datacenter).
CloudTrail and S3 Audit Log
Every commit to the bare repo is logged by CloudTrail to an S3 bucket named jada-estate-audit-logs with a 90-day expiry rule:
- CloudTrail configuration: Created with
aws cloudtrail create-trail --name jada-estate-trail --s3-bucket-name jada-estate-audit-logsand enabled. - S3 bucket setup: Versioning enabled, BlockPublicAccess enforced, lifecycle rule set to expire objects after 90 days.
- Log format: Each git push creates CloudTrail events logged as JSON to S3 with timestamps, principals, and actions.
- Query access: CloudTrail Insights queries let us answer "what changed on 2026-07-02" or "who pushed commit X" in seconds.
The IAM policy for the snapshot user is scoped tightly: push to this one repository, read from the S3 audit bucket, assume nothing else.
Key Decisions
1. Snapshots, not manual commits: We could have asked you to commit changes yourself, but that adds friction and creates inconsistent commit messages. Snapshots are deterministic—the same input always produces the same output, and they run unattended on a schedule. When they fail, they log the failure and retry.
2. Diff-review before canonical push: The local worktree has a review script that can rewind bad changes before they reach the canonical repo. This is cheaper and faster than a human code review, and it's always on.
3. One-way sync from iCloud to git: The git repo is canonical; iCloud copies are deployed from git via one-way sync, not the other way. This prevents the tools from drifting between your local edits and the canonical version.
4. Lightsail as the git host: A lightweight VPS in your AWS account, not a third-party platform. It's simpler to secure, cheaper to run, and keeps sensitive details off public platforms.
What's Next
- Test coverage for past fires: The CloudTrail logs let us audit what broke before. We're mining those incidents into a test suite so they don't recur.
- Failure domain shrinking: Once we have tests, we can refactor overly broad scripts (e.g., "do all client operations in one script") into tighter pieces so a bug in one client's tooling doesn't cascade.
- Tagged releases: The bare repo will have release tags (e.g.,
v2026-07-01-prod); deployment will pin to a tag, notmain, so rollback is explicit and auditable.
The entire system is automated. You don't review diffs manually; launchd and the diff-review script do that. You don't manually push; snapshots do. What you do is write better tools knowing they're backed by version control, audit trails, and tests.
```