```html

Diagnosing and Patching Google OAuth Reauthentication in Production EC2 Environments

During this development session, we identified a critical path configuration issue in our Google OAuth reauthentication module that required immediate diagnosis and patching across our EC2 infrastructure. This post documents the technical approach, infrastructure patterns, and lessons learned from the investigation.

The Problem: Path Resolution in Production

Our application relies on a Python module called reauth_google.py that handles OAuth token refresh cycles for Google API integrations. During a routine infrastructure audit, we discovered that this module was using hardcoded relative paths that didn't resolve correctly in our production EC2 environment, leading to intermittent authentication failures and cryptic error logs.

The core issue stemmed from the module attempting to locate credential files and configuration secrets using paths that assumed a specific working directory at runtime—an assumption that didn't hold when the application was invoked from different contexts within our containerized deployment.

Technical Investigation and Diagnosis

Our investigation began with SSH connectivity verification to our production EC2 instances in the us-east-1 region. We needed to:

  • Confirm SSH key-based authentication was functioning properly
  • Locate the exact file path of reauth_google.py in the deployed environment
  • Understand the directory structure and how secrets were being accessed
  • Review the current implementation to identify the path resolution bug

Rather than exploring credential directories directly (which would create unnecessary security exposure), we narrowed our SSH commands to connect and locate the specific module. This approach minimizes the attack surface while still gathering the diagnostic information needed.

Command structure used:

ssh -i /path/to/key.pem ec2-user@<instance-ip> 'find ~ -name "reauth_google.py" -type f'

This simple, targeted approach allowed us to locate the exact file path without traversing sensitive directories.

Root Cause Analysis

Once located, we inspected the module's implementation. The problematic pattern was evident:

# Before: Hardcoded relative path (INCORRECT)
secrets_path = "~/repos/.secrets/google_oauth.json"
config = json.load(open(secrets_path))

This approach fails because:

  • ~ expansion doesn't work with string literals in Python without explicit os.path.expanduser()
  • The relative path assumes the current working directory is predictable, which it isn't in containerized or systemd-managed environments
  • No fallback mechanism exists if the path doesn't resolve
  • Error messages don't clearly indicate which path was attempted

The Fix: Absolute Path Resolution with Environment Variables

We implemented a robust solution using environment variables and absolute path resolution:

# After: Environment-aware absolute path (CORRECT)
import os
from pathlib import Path

secrets_dir = os.environ.get('APP_SECRETS_DIR', '/opt/app/secrets')
secrets_path = Path(secrets_dir) / 'google_oauth.json'

try:
    with open(secrets_path, 'r') as f:
        config = json.load(f)
except FileNotFoundError:
    raise ValueError(f"Google OAuth secrets not found at {secrets_path}")

This approach provides:

  • Environment flexibility: The APP_SECRETS_DIR environment variable can be set per environment (dev, staging, production)
  • Sensible defaults: Falls back to /opt/app/secrets if the environment variable isn't set
  • Better error handling: Explicit FileNotFoundError with the actual path attempted
  • Path safety: Uses pathlib.Path for cross-platform compatibility

Infrastructure Configuration Changes

To support this change, we needed to ensure the environment variable is properly set across our deployment pipeline:

For EC2 instances (via user data):

#!/bin/bash
export APP_SECRETS_DIR=/opt/app/secrets
echo "APP_SECRETS_DIR=${APP_SECRETS_DIR}" >> /etc/environment

For systemd service files:

[Service]
Environment="APP_SECRETS_DIR=/opt/app/secrets"
ExecStart=/usr/bin/python3 /opt/app/main.py

For Docker deployments:

FROM python:3.11-slim
ENV APP_SECRETS_DIR=/opt/app/secrets
COPY reauth_google.py /opt/app/

Why This Matters for Infrastructure Patterns

This issue highlights a critical principle in cloud-native application design: never embed environment-specific paths in code. Instead:

  • Use environment variables for configuration that varies by deployment
  • Document expected environment variables in your deployment templates
  • Provide clear error messages when required paths don't exist
  • Test path resolution in your actual deployment environments, not just locally

Our EC2 instances across multiple availability zones (us-east-1a, us-east-1b) now consistently set this variable through the same Infrastructure-as-Code templates, ensuring consistency from staging through production.

Testing and Validation

After applying the patch, we validated the changes by:

  1. SSH'ing into test instances and verifying the environment variable was set correctly
  2. Running the reauth_google.py module in isolation to confirm path resolution
  3. Checking application logs for successful OAuth token refresh cycles
  4. Monitoring CloudWatch metrics for any remaining authentication errors

What's Next

Following this incident, we're implementing:

  • Automated testing that validates path resolution across different working directories
  • Configuration validation at application startup that fails fast if required paths don't exist
  • A deployment checklist that explicitly verifies environment variables are set correctly

This incident reinforced that even mature applications need periodic infrastructure audits to catch assumptions about the runtime environment that don't hold in production.

```