Building a Deterministic Network Monitor for macOS: From Wi-Fi Failures to Automated Healing
The Problem
A production system hidden behind a residential Mac with two fallback network links — Marriott guest Wi-Fi and iPhone 15 USB tether — was experiencing frequent connectivity flakiness. The core issue: when Wi-Fi failed silently or got trapped behind a captive portal, the Mac wouldn't automatically failover to the tether, leaving services unreachable despite having a healthy backup link available.
Manual debugging each incident was unsustainable. The solution required a deterministic, always-on network monitor that could:
- Detect both link health in real time (every 30 seconds)
- Implement intelligent failover logic with cooldowns to prevent flapping
- Repair common failure modes automatically (Wi-Fi radio off, failed SSID join, portal lockout)
- Alert on genuine state transitions without noise
- Integrate with existing incident tracking and test infrastructure
Architecture: Five-Tier Escalation with Cooldowns
The monitor implements a deterministic state machine in /Users/cb/bin/jada-net-monitor.py. Each pass evaluates network health in escalating order, with per-tier cooldown windows to prevent flapping:
Tier 1: Wi-Fi radio off → turn it on (cooldown: 2 min)
Tier 2: Wi-Fi on but no association → rejoin preferred networks (cooldown: 3 min)
Tier 3: Wi-Fi routed but internet dead OR captive portal + tether healthy
→ force failover by turning Wi-Fi off (cooldown: 15 min for re-probe)
Tier 4: Everything dead → bounce Wi-Fi radio (cooldown: 10 min)
Tier 5: Total outage → retry every pass, queue alert for recovery moment
Each tier runs only if previous tiers passed. Cooldowns are tracked in temporary state files and checked before taking action. This prevents the classic launchd problem of a script flapping state every 30 seconds if network is genuinely unstable during a transition.
The monitor runs as a LaunchAgent at ~/Library/LaunchAgents/com.jada.net-monitor.plist with a 30-second interval timer. It captures both standard output and stderr to per-run log files in a dedicated directory for forensics.
Link Health Detection
The monitor probes both links independently:
- Wi-Fi: Check if associated, then verify internet reachability via ping to a low-latency public IP and HTTPS GET to a known-good endpoint (Slack). DNS resolution is verified first.
- USB Tether: Detect the interface directly (usually
bridge0on macOS when iPhone is tethered), verify it has a valid IP, then probe internet the same way.
The monitor detects captive portal blockage by comparing HTTP and HTTPS reachability: if HTTP succeeds but HTTPS fails, the network is likely behind a redirect. In this state, if the tether is healthy, the monitor forces failover by turning Wi-Fi off — the only reliable way to make macOS abandon a portal-blocked route.
A key discovery during development: macOS now hides SSIDs from networksetup output for privacy, claiming "not associated" even when connected. The monitor works around this by checking for a valid IP on the Wi-Fi interface instead of parsing SSID strings.
State and Alerting
The monitor tracks state in a simple file-based registry: which link is active, whether failover is engaged, last transition timestamp, and a rolling 24-hour event log. A separate status CLI at /Users/cb/bin/jada-net queries this state and displays it in human-readable form:
$ jada-net
Internet: UP (Wi-Fi)
Failover: OFF
Last event: 2026-07-06 06:30:15 (30m ago) — recovery from outage, was down 8s
Alerts are sent via AWS SES to a configured inbox on genuine state transitions only: failover engagement, recovery, captive portal requiring intervention, backup tether disconnection, or total outage onset. No noise for routine passes or temporary flips.
The monitor uses the queenofsandiego AWS profile explicitly, after discovering that the default AWS credentials from a prior aws login session had expired. This uncovered a systemic risk: other launchd jobs in the ecosystem (diff-review, nightly test runner) that use boto3 without a named profile were silently losing their alert paths. This was logged as FIRES incident I-43, backlog item MT-13.
Integration with Standing Systems
The monitor plugs into the existing operational infrastructure:
- Nightly tests:
/Users/cb/icloud-repos/sites/queenofsandiego.com/tests/test_net_monitor.pycontains four unit tests covering monitor state transitions, cooldown logic, and alert queueing. These run as part of the standard nightly suite and page on-call if the monitor itself crashes. - Decision doc:
/Users/cb/icloud-jada-ops/decisions/net-monitor.mddocuments the architecture, failure modes, and tuning parameters for future changes. - FIRES incident tracking: I-43 is the open incident for this deployment; MT-13 is the backlog item for auditing other scripts' AWS credential paths.
- CONTEXT.md: Router configuration and network layout are documented for quick reference by on-call staff.
Testing and Validation
The monitor ran a soak test over three 30-second intervals on initial deployment. During testing, a real ~30-second route blip occurred naturally — the monitor detected the outage, queued the alert, and upon recovery sent the full incident summary with outage duration. No manual intervention was required; the system behaved exactly as designed.
The nightly test suite passes cleanly. The monitor was loaded into launchd and confirmed firing on schedule. State files and logs confirmed zero lock contention, no spurious action retries, and proper cooldown enforcement.
Operational Commands
Day-to-day operation requires one command: jada-net to check current status and recent events. If the output shows failover: ON TETHER, Wi-Fi is off intentionally — typically because a captive portal is in the way. Clicking through the portal in a browser will allow the monitor to rejoin on the next probing cycle (15 minutes by default, or manually by turning Wi-Fi back on).
To pause monitoring: launchctl unload ~/Library/LaunchAgents/com.jada.net-monitor.plist
To resume: launchctl load ~/Library/LaunchAgents/com.jada.net-monitor.plist
Key Decisions and Trade-offs
Why escalating tiers instead of binary failover? A strict primary/backup model would failover on any Wi-Fi hiccup, burning through tether data. Tiered healing repairs most issues (radio off, failed join) before resorting to failover, dramatically reducing false triggers.
Why force Wi-Fi off for portal lockout? macOS refuses to abandon a captive-portal-blocked route via the standard network APIs. Turning off the radio is the only reliable escape — and it's safer than it sounds, because the 15-minute re-probe cycle ensures we rejoin as soon as the portal genuinely clears.
Why SES instead of Slack for alerts? Email is simpler to integrate with launchd (no webhooks, no token rotation risk), and ensures alerts reach you even if Slack itself is unreachable. This prevents the alert system from depending on the very link it's monitoring.
What's Next
Immediate: audit other launchd jobs (diff-review, nightly test runner) for AWS credential assumptions (MT-13).
Future: extend monitoring to track Marriott's captive portal refresh rate and preemptively re-auth before it expires. Add metrics export for historical analysis of failure patterns.
```