Building a Serverless Intake Pipeline for Dragon Bodyguards: Architecture, Error Resilience, and Deployment Safety
Over the past few days, the JADA engineering team built and hardened a complete serverless intake system for Dragon Bodyguards (DBG), a new commercial offering. This post walks through the architecture, the infrastructure decisions, and the resilience patterns we added to ensure reliable form submissions and automatic outreach at scale.
What Was Done
We deployed a full intake-to-dispatch pipeline that:
- Accepts form submissions via a public Lambda Function URL
- Records client data to DynamoDB for persistence and downstream workflows
- Triggers immediate email + SMS notifications to designated team members
- Hosts the form on S3 behind a CloudFront distribution with intelligent caching
- Includes hardened deployment tooling with deterministic error handling
The system went live in stages: DynamoDB table and Lambda function were tested in isolation, the deployment script was hardened with proper timeouts, and the final S3 + CloudFront wiring is queued for execution via a production-safe script.
Infrastructure Architecture
DynamoDB for Intake Records
We created a DynamoDB table called dbg-dragons in us-east-1 to store all intake submissions. The schema is minimal but extensible:
Primary Key: submission_id (Partition Key)
Attributes: timestamp, client_name, contact_info, service_type, notes, status
DynamoDB was chosen over RDS for several reasons: zero operational overhead (no instance management), native pay-per-request billing aligns with unpredictable intake volume, and simple schema evolution as the intake form grows. All queries are key-lookups or time-range scans, which DynamoDB handles efficiently. A future analytics pipeline can pull from dbg-dragons to Redshift without affecting the intake hot path.
Lambda for Form Processing
The Lambda function dbg-onboarding-api runs in us-east-1 and is invoked by a public Function URL (API Gateway v2 under the hood). The function:
- Validates incoming form data (required fields: name, phone, service description)
- Writes the submission to
dbg-dragonswith a timestamp and unique ID - Sends an email via SES to the ops team
- Queues an SMS notification to the dispatch contact
- Returns HTTP 200 with the submission ID for client confirmation
The function's IAM role (dbg-onboarding-execution-role) has minimal permissions: dynamodb:PutItem on dbg-dragons, ses:SendEmail for the ops address, and SNS publish for SMS dispatch. No cross-account access, no wildcard resources.
S3 + CloudFront for Form Hosting
The intake form is a single-page HTML file stored in a bucket that feeds into CloudFront distribution E2Q4UU71SRNTMB. We use aggressive caching headers for static assets (CSS, JS) and short TTLs for the form HTML itself, ensuring:
- Form changes propagate within minutes (no user stale-form confusion)
- Bot protection and form validation happen client-side first, reducing Lambda invocation noise
- CloudFront DDoS mitigations are active by default (AWS Shield Standard)
CloudFront is configured to forward the HTTP Host and custom headers to the Lambda Function URL, so the backend sees accurate request context.
Key Technical Decisions
Why Serverless?
The DBG intake is bursty: zero volume on a slow week, then five submissions in an hour during a sales event. Serverless (Lambda + DynamoDB on-demand) avoids paying for idle capacity and scales automatically. If we had chosen EC2 + RDS, we'd either over-provision (wasted cost) or hit throttles during spikes (bad UX).
Hard Timeouts in Deployment Automation
The deployment script finish-deploy.sh (stored in /Users/cb/icloud-jada-ops/state/dbg-intake-staging-2026-07-05/) was hardened with explicit timeout commands on every AWS API call:
timeout 15s aws lambda get-function-url-config --function-name dbg-onboarding-api --region us-east-1
timeout 45s aws s3 cp form.html s3://dbg-intake/index.html --cache-control "max-age=300"
timeout 30s aws cloudfront create-invalidation --distribution-id E2Q4UU71SRNTMB --paths "/*"
Why? A prior production incident showed that AWS CLI calls could hang indefinitely if the network was degraded, and the script would silently proceed as if the upload succeeded. Now, if any step exceeds its timeout, the script fails loudly with a clear error message instead of leaving the form partially deployed. The entire script is designed to complete in under 3 minutes or abort cleanly.
Deterministic Deployment Verification
After the Function URL is created, the script tests it immediately by sending a test submission and verifying the response status is HTTP 200. This catches configuration errors (bad IAM role, missing environment variables, DynamoDB table not writable) before the form goes live.
Crew Health Monitoring: Timeout Robustness
In parallel, we fixed the crew health-check script at /Users/cb/icloud-jada-ops/crew-pages/healthcheck.py. The script polls Google Sheets, crew roster data, and charter confirmations every 5 minutes and publishes results to a CloudFront-fronted status page.
The bug: if any remote API call hung, the script would block indefinitely, then launchd would kill it after a timeout, leaving the health dashboard stale. The fix:
import signal
def timeout_handler(signum, frame):
raise TimeoutError("API call exceeded 15 seconds")
signal.signal(signal.SIGALRM, timeout_handler)
signal.alarm(15) # 15-second hard deadline per API call
try:
response = requests.get(GOOGLE_SHEETS_URL, timeout=10)
finally:
signal.alarm(0) # Cancel the alarm
Now if Google Sheets or any upstream service is slow, the health check skips that data point (logs a warning) and continues, so the dashboard always updates within 5 minutes instead of potentially stale for hours.
Infrastructure State and Decision Logging
We documented all infrastructure decisions in /Users/cb/icloud-jada-ops/decisions/2026-07-05-devops-toolchain-and-dbg-app-verdict.md. This one-page decision doc captures:
- Why we chose Lambda + DynamoDB over RDS + EC2
- Why CloudFront caching strategy is "short TTL on HTML, long TTL on assets"
- Why the deployment script must have hard timeouts (incident history)
- How errors are logged and who gets alerted
This prevents future engineers from re-litigating the same choices and gives visibility into constraints (e.g., why DynamoDB on-demand vs. provisioned capacity, why single Lambda vs. Step Functions).
What's Next
The form is ready to accept submissions. The next action is running the finalized deployment script, which will:
- Create the public Lambda Function URL if it doesn't exist
- Test the endpoint with a real form submission
- Update the form HTML to point to the live Function URL
- Upload to S3 with correct cache headers
- Invalidate CloudFront to push changes live
Once live, we'll monitor DynamoDB write throttling (unlikely on on-demand, but good to measure), Lambda error rates via CloudWatch, and average form-submission-to-email latency. A future iteration can add form analytics (tracking drop-off, most common service requests) by streaming dbg-dragons data to S3 via DynamoDB Streams for warehouse ingestion.