Diagnosing and Staging a Critical Deposit System Outage: Apps Script Access Control and Deployment State Recovery
This post documents the technical investigation and remediation approach for a silent but total failure of the deposit/reservation system across 11 production event pages. The outage manifested as HTTP 403 (access denied) and 404 (not found) responses from Google Apps Script endpoints that power the "Reserve" widget on sailjada.com and queenofsandiego.com properties.
What Was Done
Over a two-hour investigation window, we:
- Identified the root cause: Apps Script deployment access control reverted to restricted, and one deployment was deleted entirely
- Mapped all 11 affected event pages and their endpoint dependencies
- Verified the source code and infrastructure were intact; only the deployment/access state was broken
- Staged a two-step remediation (redeploy + access control change) ready for immediate execution
- Documented the exact Google Console steps required to restore service (no code changes needed)
Technical Details: The Outage Anatomy
Symptom: Every public event page returning broken deposit widgets. The JavaScript fetch to the Apps Script endpoint was silently failing, leaving users with no way to complete reservations or deposits.
Root Cause Analysis:
- Two Google Apps Script projects serve deposit endpoints:
- Project
1dDpSK8JZda7XUpKIGlyyAX19KLL4JqFjYVtpcunB5ZE3-NMX_9v0lQJ5(primary, serves 10 pages): deployment IDAKfycbw...44Pme8wCA, returning 403 Forbidden - Project (worship site): deployment deleted, returning 404 Not Found
- Project
- The 403 indicates the deployment exists but access control is set to a restricted state (likely "Private" or "Only me")
- The 404 indicates the deployment was removed entirely, possibly during a failed redeploy or manual cleanup
Why This Breaks Revenue Immediately: The deposit flow is entirely client-side JavaScript calling these endpoints. No human can work around it; no background job can retry it. Every visitor hitting a Reserve button gets a network error. The failure is silent to the user (no alert), so they may simply leave without attempting to contact support.
Infrastructure: Apps Script Deployment Model
Google Apps Script deployments follow a specific access control model that's often misunderstood:
- Deployment ID: Immutable identifier for a specific published version. Remains the same if you redeploy into the same slot.
- /exec Endpoint: The HTTPS URL that clients call. Derived from the deployment ID. Format:
https://script.google.com/macros/s/[DEPLOYMENT_ID]/exec - Access Control: Set at deployment time, determines who can invoke the endpoint:
Anyone— any HTTP client, no auth required. Correct for public widgets.Anyone with a Google Account— requires the client to be signed into Google.Only me/Private— only the script owner. Breaks all public clients.
- Execute as: Which identity the script runs under (typically the owner). Separate from access control.
The staged fix requires two actions:
- For the primary project (10 pages): Navigate to Google Apps Script console, open the project, go to Deploy → Manage deployments, select the active deployment, change Who has access to Anyone, and click Deploy. No code changes needed; this only resets the access control.
- For the worship site: The deployment was deleted. Redeploy the script first (same code, new deployment ID), then apply the access control change above.
Impact on Event Pages
Affected pages are served from S3 buckets and distributed via CloudFront. Each embeds the endpoint URL directly in the JavaScript widget:
// Hardcoded in page templates across:
// - /Users/cb/Documents/repos/sites/queenofsandiego.com/events/*.html
// - /Users/cb/Documents/repos/sites/sailjada.com/events/*.html
const depositEndpoint = "https://script.google.com/macros/s/AKfycbw...44Pme8wCA/exec";
If a redeploy changes the deployment ID, the pages must be updated and resynced to S3. However, the staged approach assumes redeploying into the same slot, which preserves the endpoint URL. This minimizes the change surface and avoids a page redeploy cycle.
Key Decisions and Why
Why We Didn't Immediately Redeploy: Redeploying can change the deployment ID (and thus the /exec URL) if you deploy into a new slot. This would require updating and resyncing all 11 event pages. Instead, the staged approach first tries to fix the access control on the existing deployment. If the deployment is still valid but just locked down, this is a one-step fix with zero page changes.
Why We Verified the Source Code: Before touching the deployment, we confirmed that the source code in the Git repos was intact and matched what was last deployed. This ensures the issue is not bad code, but only a deployment/access state issue. This is critical because if the source were corrupted, redeploying would perpetuate the outage.
Why This Matters for Cost Savings: A silent, total booking funnel failure can cost hundreds or thousands in lost deposits over hours without alerting anyone. The fix is literally two clicks in the Google console. Automating a health check on these endpoints (polling the /exec URL for 200 responses) would catch similar issues within minutes and trigger a page alert.
What's Next
- Execute the fix: Two minutes in the Google Apps Script console (primary project + worship redeploy).
- Verify live: Curl both /exec endpoints and confirm 200 responses, not 403 or 404.
- Test the widget: Load a few event pages in a browser, click Reserve, and confirm the deposit form appears.
- Add monitoring: Set up a CloudWatch or Lambda health check that polls both endpoints every 5 minutes and alerts if either returns non-200. This can be a simple GET request with no payload.
- Document recovery playbook: Create a runbook in the jada-ops repo so any team member can re-apply access control without guessing.
Once this is live, the next high-priority item is sending the Giovanna charter offer (Base Cost $3,334, ready to ship). Both are revenue-direct and both require minimal effort relative to payoff.
```