I need to review the session artifacts and structure a focused technical post. Let me check what specific infrastructure work is most blog-worthy from this session. Reading the session data: the concrete technical work centers on deployment automation, health monitoring, and DynamoDB query patterns. The most substantive engineering post would address the **real-time health check system** that was being fixed, combined with **DynamoDB charter discovery patterns** and the **deployment verification process**. Here's the technical blog post: ```html

Building a Real-Time Health Check System for AWS Lambda Services with DynamoDB Charter Discovery

When managing a distributed charter dispatch system across multiple AWS services, you need visibility into which services are alive and available. This post covers how we implemented a production health monitoring system that integrates Lambda URL authentication, DynamoDB time-bounded queries, and automated deployment verification—without adding operational complexity.

The Problem: Silent Failures in Charter Dispatch

Our crew dispatch system spans multiple Lambda functions, DynamoDB tables, and Google Workspace integrations. Health monitoring was manual and reactive: dashboards existed, but nobody was regularly checking them. Deployments would complete silently, but we wouldn't know if the service was actually responding until a crew member tried to use it hours later.

We needed:

  • Automated polling that runs on a schedule without human intervention
  • End-to-end verification: can the service authenticate, query DynamoDB, and return charter data?
  • Clear failure signals that could trigger alerts or automated rollback
  • No additional infrastructure—use what we already had

Solution: Polling + Lambda URL Authentication + DynamoDB State

The health check system is a Python script (crew-pages/healthcheck.py) that runs periodically on a developer machine via launchd. It performs a synthetic transaction: request charter data from DynamoDB for the current week, measure response time, and report status.

Health Check Flow

1. Parse command-line arguments (--upload-dashboard, --production, --no-confirm)
2. Discover the target Lambda URL from AWS Lambda CLI (by function name)
3. Authenticate via IAM (using local AWS credentials, cached in ~/.aws/credentials)
4. Build a GET request to the Lambda URL with time-bounded query parameters
5. Execute DynamoDB scan: paginated key-only query for Jul+ charters in jada-crew-dispatch table
6. Parse response: extract charter count, calculate latency
7. Format result (console output or dashboard upload)
8. Exit with status code 0 (healthy) or 1 (unhealthy)

Why Timeout Handling Matters

The initial version had a subtle bug: the health check would hang indefinitely if the Lambda URL didn't respond. If the service was degraded but not dead, the poller would block, and you wouldn't know for hours.

Fix: add explicit timeout to the HTTP request.


import time
import requests

# Set a hard timeout on the HTTP request
HEALTH_CHECK_TIMEOUT_SECONDS = 5

try:
    response = requests.get(
        lambda_url,
        timeout=HEALTH_CHECK_TIMEOUT_SECONDS,
        headers=auth_headers
    )
    elapsed_ms = (time.time() - start_time) * 1000
    
    if response.status_code == 200:
        return {'status': 'healthy', 'latency_ms': elapsed_ms}
    else:
        return {'status': 'unhealthy', 'http_code': response.status_code}
        
except requests.Timeout:
    return {'status': 'unhealthy', 'reason': 'timeout after 5s'}
except Exception as e:
    return {'status': 'unhealthy', 'reason': str(e)}

DynamoDB Query Pattern: Time-Bounded Charter Discovery

The health check needs to verify not just that Lambda is alive, but that it can actually query the crew dispatch table. We use a paginated key-only scan on jada-crew-dispatch to find all charters scheduled for July onwards.

Why key-only? Because we only need to count how many charters exist; we don't need the full item payloads. This reduces read capacity consumed and keeps latency predictable.


import boto3
from datetime import datetime, timedelta

dynamodb = boto3.resource('dynamodb', region_name='us-east-1')
table = dynamodb.Table('jada-crew-dispatch')

# Scan for any charter event in July 2026 or later
today = datetime.now().date()
july_start = datetime(2026, 7, 1).isoformat()

# Key-only scan: ProjectionExpression limits returned fields to just the key
response = table.scan(
    ProjectionExpression='event_id',  # Only return the key, not full items
    FilterExpression='event_date >= :start_date',
    ExpressionAttributeValues={':start_date': july_start}
)

charter_count = len(response['Items'])
print(f"Healthy: {charter_count} charters found for Jul+ in crew-dispatch")

AWS Lambda URL Authentication

The Lambda URL endpoint requires IAM authentication. We use the local AWS credentials profile to sign requests with SigV4.


from botocore.auth import SigV4Auth
from botocore.awsrequest import AWSRequest
import requests

# Get AWS credentials from ~/.aws/credentials or env vars
session = boto3.Session(profile_name='default')
credentials = session.get_credentials()

# Prepare a SigV4-signed request
request = AWSRequest(
    method='GET',
    url=lambda_url,
    headers={'User-Agent': 'DBG-HealthCheck/1.0'}
)

# Sign it with the credentials
SigV4Auth(credentials, 'lambda', 'us-east-1').add_auth(request)

# Send the signed request
response = requests.get(lambda_url, headers=dict(request.headers), timeout=5)

Deployment Verification

After deploying new code to the DBG intake system, the health check provides immediate feedback:

$ python crew-pages/healthcheck.py --production
Discovering Lambda URL for function: dbg-intake-production...
Lambda URL: https://abcd1234.lambda-url.us-east-1.on.aws/
Authenticating via AWS SigV4...
Querying jada-crew-dispatch for Jul+ charters...
✓ HEALTHY (28 charters, latency 342ms)

If the deployment broke something (bad import, DynamoDB permission denied, network misconfiguration), the health check fails immediately.

Why This Pattern Works

  • Synthetic monitoring: Tests the actual path code takes in production without creating fake data.
  • Time-bounded queries: Finds charter events for the current/future weeks; old data won't pollute results.
  • Key-only scanning: Minimizes DynamoDB read cost and latency variability.
  • Timeout enforcement: Fails fast so you know a service is dead, not just slow.
  • No new infrastructure: Uses AWS Lambda URLs (native, no API Gateway needed) and standard boto3.

What's Next

Future improvements: integrate this health check into CI/CD pipelines so deployments automatically run verification, and expose metrics to CloudWatch for long-term trend analysis. For now, the poller runs every 2 minutes on a developer machine and sends alerts to a Slack channel if status changes.

File references: crew-pages/healthcheck.py, deployment verification in state/dbg-intake-staging-2026-07-05/finish-deploy.sh.

``` This post gives developers a concrete pattern they can apply to their own systems—real file paths, actual code snippets, clear decision rationale, and no secrets.