Diagnosing and Patching Google OAuth Reauthentication in Production EC2 Environments
During this development session, we identified a critical path configuration issue in our Google OAuth reauthentication module that required immediate diagnosis and patching across our EC2 infrastructure. This post documents the technical approach, infrastructure patterns, and lessons learned from the investigation.
The Problem: Path Resolution in Production
Our application relies on a Python module called reauth_google.py that handles OAuth token refresh cycles for Google API integrations. During a routine infrastructure audit, we discovered that this module was using hardcoded relative paths that didn't resolve correctly in our production EC2 environment, leading to intermittent authentication failures and cryptic error logs.
The core issue stemmed from the module attempting to locate credential files and configuration secrets using paths that assumed a specific working directory at runtime—an assumption that didn't hold when the application was invoked from different contexts within our containerized deployment.
Technical Investigation and Diagnosis
Our investigation began with SSH connectivity verification to our production EC2 instances in the us-east-1 region. We needed to:
- Confirm SSH key-based authentication was functioning properly
- Locate the exact file path of
reauth_google.pyin the deployed environment - Understand the directory structure and how secrets were being accessed
- Review the current implementation to identify the path resolution bug
Rather than exploring credential directories directly (which would create unnecessary security exposure), we narrowed our SSH commands to connect and locate the specific module. This approach minimizes the attack surface while still gathering the diagnostic information needed.
Command structure used:
ssh -i /path/to/key.pem ec2-user@<instance-ip> 'find ~ -name "reauth_google.py" -type f'
This simple, targeted approach allowed us to locate the exact file path without traversing sensitive directories.
Root Cause Analysis
Once located, we inspected the module's implementation. The problematic pattern was evident:
# Before: Hardcoded relative path (INCORRECT)
secrets_path = "~/repos/.secrets/google_oauth.json"
config = json.load(open(secrets_path))
This approach fails because:
~expansion doesn't work with string literals in Python without explicitos.path.expanduser()- The relative path assumes the current working directory is predictable, which it isn't in containerized or systemd-managed environments
- No fallback mechanism exists if the path doesn't resolve
- Error messages don't clearly indicate which path was attempted
The Fix: Absolute Path Resolution with Environment Variables
We implemented a robust solution using environment variables and absolute path resolution:
# After: Environment-aware absolute path (CORRECT)
import os
from pathlib import Path
secrets_dir = os.environ.get('APP_SECRETS_DIR', '/opt/app/secrets')
secrets_path = Path(secrets_dir) / 'google_oauth.json'
try:
with open(secrets_path, 'r') as f:
config = json.load(f)
except FileNotFoundError:
raise ValueError(f"Google OAuth secrets not found at {secrets_path}")
This approach provides:
- Environment flexibility: The
APP_SECRETS_DIRenvironment variable can be set per environment (dev, staging, production) - Sensible defaults: Falls back to
/opt/app/secretsif the environment variable isn't set - Better error handling: Explicit
FileNotFoundErrorwith the actual path attempted - Path safety: Uses
pathlib.Pathfor cross-platform compatibility
Infrastructure Configuration Changes
To support this change, we needed to ensure the environment variable is properly set across our deployment pipeline:
For EC2 instances (via user data):
#!/bin/bash
export APP_SECRETS_DIR=/opt/app/secrets
echo "APP_SECRETS_DIR=${APP_SECRETS_DIR}" >> /etc/environment
For systemd service files:
[Service]
Environment="APP_SECRETS_DIR=/opt/app/secrets"
ExecStart=/usr/bin/python3 /opt/app/main.py
For Docker deployments:
FROM python:3.11-slim
ENV APP_SECRETS_DIR=/opt/app/secrets
COPY reauth_google.py /opt/app/
Why This Matters for Infrastructure Patterns
This issue highlights a critical principle in cloud-native application design: never embed environment-specific paths in code. Instead:
- Use environment variables for configuration that varies by deployment
- Document expected environment variables in your deployment templates
- Provide clear error messages when required paths don't exist
- Test path resolution in your actual deployment environments, not just locally
Our EC2 instances across multiple availability zones (us-east-1a, us-east-1b) now consistently set this variable through the same Infrastructure-as-Code templates, ensuring consistency from staging through production.
Testing and Validation
After applying the patch, we validated the changes by:
- SSH'ing into test instances and verifying the environment variable was set correctly
- Running the
reauth_google.pymodule in isolation to confirm path resolution - Checking application logs for successful OAuth token refresh cycles
- Monitoring CloudWatch metrics for any remaining authentication errors
What's Next
Following this incident, we're implementing:
- Automated testing that validates path resolution across different working directories
- Configuration validation at application startup that fails fast if required paths don't exist
- A deployment checklist that explicitly verifies environment variables are set correctly
This incident reinforced that even mature applications need periodic infrastructure audits to catch assumptions about the runtime environment that don't hold in production.
```