Diagnosing and Staging a Critical Deposit Collection Outage: Apps Script Access Control and Endpoint Validation
During a June 2026 operations session, we discovered that deposit collection functionality across 11 public event pages had silently failed. The "Reserve" widget—critical to inbound booking funnels—was returning 403 and 404 errors, meaning zero deposits were being collected while the outage remained undetected. This post documents the diagnostic approach, infrastructure changes, and staging work to restore revenue collection.
The Problem: Dark Funnel Failure
Event pages at sailjada.com and queenofsandiego.com depend on Google Apps Script endpoints to handle reservation deposits. Two distinct deployments power this:
- Primary endpoint (10 event pages):
https://script.google.com/macros/s/.../execreturning 403 Forbidden - Worship endpoint (1 page):
https://script.google.com/macros/s/.../execreturning 404 Not Found
The 403 indicated access control was revoked; the 404 suggested the deployment was deleted entirely. Without visibility into why, we couldn't know if this had been down for hours or days—and we couldn't quantify lost deposits.
Diagnosis: Hunting the Apps Script Project
The first challenge was locating which Google Apps Script project owned these deployments. We took a systematic approach:
- Search ops documentation for any reference to the Apps Script project name or deployment ID across
~/Library/Mobile Documents/com~apple~CloudDocs/jada-ops/ - Scan source repositories for clasp configurations (Google's Apps Script CLI stores project metadata in
.clasp.json) - Locate all .clasp.json files to extract project IDs and match against the live endpoint IDs
- Cross-reference appsscript.json and project notes to confirm which project was in use and its current state
This identified the primary project: 1dDpSK8JZda7XUpKIGlyyAX19KLL4JqFjYVtpcunB5ZE3-NMX_9v0lQJ5. The worship endpoint required separate investigation into a second project that had apparently been decommissioned.
Infrastructure Context: Endpoint Hosting and S3 Distribution
Before we could fix the Apps Script access, we needed to understand the full flow:
- Event page HTML is stored in S3 and served via CloudFront
- JavaScript embedded in those pages makes AJAX calls to the Apps Script
/execendpoints - S3 paths follow the pattern:
s3://jada-events-web/[event-slug]/index.html - CloudFront distribution caches these pages and handles HTTPS/CORS
The critical insight: if we needed to change the endpoint URL (e.g., migrate to a new Apps Script deployment), we'd need to:
- Update the inline JavaScript in all 11 event HTML files
- Reupload them to S3
- Invalidate the CloudFront cache (or wait for TTL expiry)
- Verify the new endpoints were live and returning 200
We prepared for this scenario, but first we tested whether the existing endpoints could be repaired in-place.
Access Control Fix: Google Apps Script Deployment Permissions
The root cause became clear: the primary Apps Script project's deployment had its access control set to restrict who could invoke the /exec endpoint. To fix this required a single action in the Google Apps Script console:
- Open project
1dDpSK8JZda7XUpKIGlyyAX19KLL4JqFjYVtpcunB5ZE3-NMX_9v0lQJ5atscript.google.com - Navigate to Deploy → Manage deployments
- Select the active deployment
- Change "Who has access" from the restricted state to "Anyone"
- Change "Execute as" to "Me" (to preserve the original owner's permissions for database access)
- Click Deploy
Critically, this redeployment into the same deployment slot preserves the /exec URL—meaning no HTML changes are required. The URL stays the same; only the access control policy changes.
Staging and Verification Strategy
Rather than immediately apply changes to production pages, we:
- Dry-ran endpoint rewrites across all 11 event page HTML files (in local clones) to confirm they would parse and update correctly
- Tested every live endpoint manually via curl to confirm HTTP status codes before and after changes
- Prepared S3 upload scripts with exact bucket paths and file checksums to ensure atomicity
- Documented the rollback procedure in case the new endpoints were down
This staging approach meant we could apply the fix with high confidence and minimal rollback risk.
Integrated Monitoring: Calendar Token and Crew Dispatch Integrity
While debugging the deposit outage, we discovered related systemic issues in supporting infrastructure:
- Calendar token refresh failures in the crew-dispatch workflow prevented automated crew assignment
- Gmail authentication tokens in DynamoDB tables were stale, blocking automated email dispatch
- Google Sheets access for the charter helper functions was timing out
Rather than band-aid fixes, we created a comprehensive reauth toolkit to refresh all OAuth tokens in one pass. This toolkit:
- Inventories all stored credentials in DynamoDB (tables:
crew-dispatch,email-list, others) - Tests token refresh for each credential
- Identifies which are truly dead vs. which have transient network issues
- Provides an interactive reauth option without hanging the process
This prevented a cascade failure where the deposit endpoints could be fixed but crew assignment and confirmation emails would still fail.
Key Decisions and Trade-offs
Why redeploy in-place rather than create a new deployment? URL stability. Changing the deployment means editing all 11 event pages and syncing to S3—much higher risk of coordinating failures. If the in-place redeploy succeeds, we get the fix with zero client changes.
Why test the full token chain before claiming victory? A 200 response from the Apps Script endpoint is only the first step. If the underlying Sheets queries or database writes fail, deposits still fail silently. The real test is end-to-end: a test deposit from a real browser.
Why document the worship endpoint separately? It's a different project and a different failure