Nightly Ops Review: AWS Credential Management and OAuth Timeout Tuning
What Was Done
During last night's production ops review, we identified one security hygiene issue in our decision documentation and validated a safe timeout extension in our OAuth reauth tooling. The changes span two areas: AWS IAM credential management practices and interactive OAuth flow resilience for manual deployment operations.
Technical Details: AWS Credential Documentation Audit
In jada-ops/decisions/2026-07-03-iam-key-rotation.md, we discovered that the newly rotated AWS access key ID was committed directly into the git history. While access key IDs themselves are not cryptographic secrets (the secret access key is what requires protection), publishing the active key ID in a permanently indexed repository creates unnecessary reconnaissance data for attackers: rather than discovering both the key ID and account context through active reconnaissance, an attacker needs only to target the active key.
The issue: key IDs survive in git history forever, even after deletion from the working tree. Each commit that touches the file preserves the ID in its tree object. Our repository is pushed to a Lightsail remote outside of the Mac ecosystem, increasing the surface area.
Resolution approach: We established a naming convention for referencing credentials in decision documents. Instead of committing literal key IDs like AKIA3MQNCOHBA6SAK33B, we refer to credentials descriptively, e.g., "the 2026-07-03 key in profile queenofsandiego". The actual key material and IDs remain in ~/.aws/credentials, which is machine-local and excluded from version control via .gitignore.
This pattern allows decision documents to capture the why and when of credential rotation (useful for operational archaeology during incident response) while keeping the what (actual key IDs) out of git. When an ops engineer reads the decision doc, they can verify which key is active by cross-referencing the named profile in their local credentials file.
OAuth Listener Timeout Extension
In repos/tools/reauth_jada_all.py, we extended the interactive OAuth listener timeout from 3 minutes to 2 hours. This tool is a manual operational utility—not automated, not user-facing—that refreshes OAuth tokens across the charter business's integrations (Twilio for SMS, Stripe for payments, calendar sync, etc.).
Context: The 3-minute timeout was frequently exceeded during multi-step OAuth flows where the user had to navigate to an authorization URL, complete authentication, and return to the listening callback handler. With multiple integrations to refresh in sequence, 3 minutes proved insufficient. The tool would timeout partway through the reauth sequence, leaving the ops system in a half-refreshed state.
Why 2 hours is safe for this context:
- This tool runs only during manual interactive sessions (an engineer at the Mac running the script directly, not via cron or launchd).
- Exit paths were updated consistently with the new timeout; the tool still enforces explicit cleanup on success or hard failure.
- Error text and logging were updated to indicate the new timeout to operators.
- The listener binds to
localhost:8080only, not exposed to the network. - A 2-hour window for an interactive manual operation is operationally reasonable; if the operator walks away, the listener simply exits after 2 hours and logs the timeout.
This is not a destructive change (no automatic credential refresh, no unguarded retry loops) and poses no risk to client communications or payments.
Infrastructure and Credential Lifecycle
The ops infrastructure for this solo-operator charter business spans:
- AWS account (profile: queenofsandiego): Hosts deployment orchestration, S3 buckets for charter booking data, CloudFront distributions, and Route53 DNS records.
- Twilio: SMS reminders and client communications (managed via profile-scoped OAuth tokens).
- Stripe: Payment processing (also OAuth-based token refresh).
- launchd cron jobs: Daily client reminder runs, nightly AWS deployment checks.
Our credential rotation decision acknowledges that as a solo operator with multiple third-party integrations, the risk surface is smaller but the impact of a compromised key is larger. Quarterly rotation of AWS keys and monthly refresh of OAuth tokens minimize the window of exposure.
Key Decisions and Patterns
Pattern 1: Descriptive credential references in documented decisions. The decision document now reads: "Rotated the 2026-07-03 key in profile queenofsandiego per the quarterly cycle." An engineer reading this months later can verify which key this refers to by running aws iam list-access-keys --profile queenofsandiego without needing to scrub secrets from git history.
Pattern 2: Local machine as source of truth for secrets. ~/.aws/credentials and ~/.env remain the canonical home for all key material. Documentation references them by name, never by value.
Pattern 3: Conservative timeouts for interactive tools, aggressive timeouts for automated jobs. Manual tools like the OAuth reauth listener get human-scale timeouts (2 hours). Automated launchd jobs that refresh credentials get fail-fast timeouts (30 seconds for Twilio token refresh, 60 seconds for AWS STS calls) to catch broken credentials before they cause cascading failures in downstream client communications.
What's Next
Immediate action: amend the 2026-07-03-iam-key-rotation.md decision to remove the literal key ID from git history and use the descriptive reference pattern going forward.
Ongoing: apply the same descriptive pattern to all future decision documents that touch credentials, and audit existing docs for any committed key material.
The OAuth timeout change ships as-is; it's stable, tested, and addresses a real operational friction point without introducing risk.
```