array_start.sh → array_started.sh (bulk rename — actual file has always been array_started.sh; docs consistently used the wrong name across 9 files) failover.sh → fallback.sh (bulk rename — script is Fallback/fallback.sh; docs used the old name throughout) FAILOVER_HOST2_RUNS_FOR_HOST1_TIER1 → FALLBACK_HOST2_COVERS_HOST1_TIER1 (README.md — wrong variable name, actual vars use FALLBACK_ prefix and COVERS_ not RUNS_FOR_) docker_watchdog.sh attribution: "started by array_started.sh" → "called every minute by watchdog_orchestrator.sh" (docker_watchdog is not in ARRAY_START_SCRIPTS) server_reboot.sh shutdown sequence: step 7 "rc.docker stop" → calls array_stopping.sh (user_scripts, fallback, rsync, mover, docker_container_stop) Rsync manual ARRAY_START_SCRIPTS list: removed stability_watchdog and docker_watchdog (now in watchdog_orchestrator); failover → fallback
12 KiB
━━━━━ FALLBACK ━━━━━
Mutual automatic failover between two independent unRAID servers. When one goes down the other starts its containers, cuts over DNS, and keeps users online. When it comes back everything hands back in the correct sequence — covering DDNS stops, containers stop, rsync writeback runs, containers start on the primary, primary DDNS starts last — so users hit the returning server only after it's actually ready.
Built from scratch. Refined through a year of production testing. The DDNS sequencing and handback order were the hardest parts to get right. Both directions are exercised regularly with
fallback_test.sh.
━━━ THE PROBLEM THAT BUILT THIS ━━━
Running a self-hosted stack means being the operator. These are the specific problems that drove this build:
A Single Point of Failure for an Entire Household
HOST1 runs Emby, NPM, Authelia, NextCloud, VaultWarden, and every service the household
uses daily. When HOST1 goes down — even briefly — all of those services go down with it.
The fix: a second server with mirrored critical data that covers the first automatically.
From a user's perspective, a brief interruption and then everything is back.
DNS Cutting Over Before the Server Was Ready
Early attempts started containers on the covering server then updated DNS. Problem: DNS
propagated in under a minute. Users hit the new IP before Emby had finished starting,
before Authelia had loaded its sessions, before NPM had loaded its proxy configurations.
The fix: warm standby for the auth stack. NPM, LLDAP, and Authelia run actively on both
servers at all times. When DNS cuts over, auth is already running and ready.
Split Brain DNS During Handback
When HOST1 returned, the obvious sequence was: start HOST1 containers, then switch DNS
back. Problem: between "start containers" and "DNS switches" both servers' DDNS containers
were running, both updating the same domain with different IPs. Users got routed randomly
between servers — intermittent auth failures, no clear error state anywhere.
The fix: stop DDNS on the covering server first, before anything else moves. There is
never a window where two DDNS containers update the same record.
Rsync Running Into Active Container I/O
Syncing data back while containers were still running — to minimise downtime — produced
slower transfers, potential file inconsistency, and database dirty state risk.
The fix: stop containers before syncing. The outage window is only the rsync duration —
typically minutes. Clean static source at full bandwidth, predictable state every time.
No Way to Validate the System Before Needing It
A failover system that has never been tested is not a failover system — it is a hope.
The fix: fallback_test.sh — a controlled simulation using an iptables DROP rule to make
the remote appear unreachable, triggering the full sequence without taking anything offline.
A safety trap removes the rule on any exit — crash, error, ctrl-c, or clean completion.
━━━ THE TWO-SERVER SETUP ━━━
HOST1 — unRAID-Gmer4Lfe
Hardware: Threadripper 1950X, 128GB RAM, ZFS cache pools
Location: Primary site
DDNS: Gmer4Lfe.com
Role: Primary — full service stack, arr_sync union for all media shares
HOST2 — unRAID-Jayred365
Hardware: Intel i5 10th gen, 64GB RAM
Location: Remote — 50 miles away
DDNS: Gmer4Lfe.us
Role: Secondary — own stack + covers HOST1 + mirrors critical data
Hardware does not need to match. Everything is accessed through /mnt/user/ — unRAID's
fused share layer. HOST1 has a Threadripper with ZFS pools. HOST2 has completely different
hardware. Fallback containers on HOST2 mount /mnt/user/Movies and see mirrored data
because the share names match. The hardware underneath is irrelevant.
What must match between servers:
Share names /mnt/user/Movies must exist on both servers (mirrored data)
Container names "Emby" on HOST2 must be the container HOST2 starts for HOST1
Network names Docker custom networks must match for NPM routing to work
Union Library — arr_sync.sh Is the Source of Truth
Both servers run arr instances simultaneously. arr_sync.sh keeps all databases in bidirectional union before every rsync cycle — either server can download anything and it propagates to the other:
HOST1 runs: Radarr (Movies), Sonarr (Tv_Shows), Lidarr (Music)
HOST2 runs: Radarr (Anime_Movies), Sonarr (Anime_Shows)
All arr databases merged via arr_sync. rsync spreads files additively.
Either server downloads to any share — the other catches up on the next cycle.
When HOST2 runs HOST1's arr containers during a Tier 4 fallback, those arrs already know the full library — arr_sync kept them in union while both servers were online.
Auth Stack — Warm on Both Servers
NPM, LLDAP, and Authelia run actively on both servers at all times. HOST2 needs them running to serve his own users daily — this is not a fallback-only configuration. HOST1 is source of truth: all changes mirror to HOST2 every 15 minutes via critical sync.
When DNS cuts over, auth is already running on the covering server. The 30-60 second dead zone where auth is coming up after DNS has already switched does not exist.
━━━ HOW IT WORKS ━━━
Both servers run fallback.sh independently as a continuous background process. Each server
makes all decisions from two pings every FALLBACK_CHECK_INTERVAL seconds:
ping REMOTE_TAILSCALE_IP # is the other server reachable?
ping EXTERNAL_IP # do I have internet? (default: 8.8.8.8)
No SSH signaling between servers. No shared state file. No election algorithm. Each server acts entirely from its own network perspective.
States:
| State | Remote | Internet | Action |
|---|---|---|---|
| NORMAL | up | up | Silent — own containers, own DDNS on |
| FALLBACK | down | up | Start tier containers, cut DDNS over |
| NO_INTERNET | — | down | Stop own DDNS immediately, wait |
| DARK | down | down | Same as NO_INTERNET — cannot determine cause |
━━━ DDNS — THE CRITICAL PART ━━━
This took a year to get right. Do not change the sequencing.
fallback.sh is the sole authority over when any DDNS container starts. Network state
returning is not permission to start DDNS. Only completion of the full handback sequence
grants that permission.
ONE DOMAIN → ONE DDNS ACTIVE → AT ALL TIMES
Gmer4Lfe.com → HOST1's DDNS normally → HOST2's DDNS during HOST1 outage
Gmer4Lfe.us → HOST2's DDNS normally → HOST1's DDNS during HOST2 outage
Own DDNS: ON when this server has internet. OFF when internet is lost.
Remote DDNS: ON as Tier 1 fallback action. OFF as FIRST handback action.
Auto-start: NEVER — DDNS never starts automatically on internet return.
Why auto-start is forbidden: If HOST1 lost internet and its DDNS auto-started when internet returned, there is a window where both servers are updating the same domain with different IPs. Users get routed randomly — some to the primary with fresh data, some to the covering server. Authentication sessions don't transfer between servers. This is split brain and produces the most confusing symptoms: intermittent auth failures with no clear error state anywhere.
The brief gap where neither DDNS is updating the record is intentional. DNS TTL caches the last value. During the rsync + container start window, cached DNS still routes users to the covering server where containers are still running. By the time the cache expires, the primary's DDNS has started and the record points at the right server.
━━━ TIERED FALLBACK ━━━
Starting the full stack for a 5-minute power blip wastes resources — most brief outages resolve before Tier 2 would even activate. Tiers start only what is needed for the actual outage duration.
| Tier | Delay | Coverage | Why This Timing |
|---|---|---|---|
| Tier 1 | Immediate | Vital services + Live TV | People are watching — cannot wait 4 hours |
| Tier 2 | 4hr (HOST*_TIER2_DELAY) | NextCloud, Immich, Jellyseerr | 4hr covers most ISP and power events |
| Tier 3 | 12hr (HOST*_TIER3_DELAY) | Dashboard, AdGuard, Git, Collabora | Secondary — useful but not daily-critical |
| Tier 4 | 24hr (HOST*_TIER4_DELAY) | Arrs + downloaders | Significant I/O — only worth starting at 24hr |
Tier 1 always includes the remote domain's DDNS container as the first entry — DNS coverage happens before any other container starts.
For the actual container lists and tier delay values, see Manual-Fallback.md.
━━━ HANDBACK SEQUENCE ━━━
When the remote server returns after a FALLBACK event. Every step has a reason. Do not reorder.
-
Strike confirmation — FALLBACK_HANDBACK_STRIKES consecutive remote-up checks before handback begins. Prevents false triggers from brief network recovery.
-
Pre-flight checks — version parity, remote array mounted, remote Docker daemon up. Any failure aborts and retries next cycle.
-
Staged reverse handback: Tier 4 → 3 → 2 — Emby and vital services stay on the covering server serving users throughout this phase. Each tier: stop local containers → rsync writeback (if outage exceeded tier threshold) → start on remote.
-
DDNS handoff — stop remote DDNS immediately before Tier 1 goes down. This prevents split brain during the Tier 1 rsync window.
-
Tier 1 handback — stop local vital services, rsync writeback, start on remote.
-
Start remote DDNS last — DNS cuts back to the primary only after all containers are confirmed running.
-
Return to NORMAL — state file reset, own DDNS restored if it was stopped.
━━━ MUTUAL FALLBACK — BOTH DIRECTIONS ━━━
The same fallback.sh handles both directions without any code changes. detect_hosts()
determines which server is local and which is remote at runtime, then selects the correct
container arrays and tier delays from config via MY_ID.
HOST2 covers HOST1: FALLBACK_HOST2_COVERS_HOST1_TIER* (in host2.conf)
HOST1 covers HOST2: FALLBACK_HOST1_COVERS_HOST2_TIER* (in host1.conf)
Both servers run identical scripts. MY_ID selects the correct arrays. No hostname comparisons anywhere in the script code.
━━━ INDEPENDENCE — ALWAYS ONE RSYNC STOP AWAY ━━━
HOST2 is designed to be fully independent if needed. If HOST2 ever wants to separate from HOST1: stop HOST1 pushing data. Any changes HOST2 makes to his own data stick permanently. His server becomes fully independent immediately — no script changes, no migration, no data movement required. The fallback and rsync scripts are configuration-driven.
━━━ SCRIPTS IN THIS FOLDER ━━━
| Script | Role | When It Runs |
|---|---|---|
fallback.sh |
Continuous state machine — monitors remote, manages fallback and handback | Continuously (started by array_started.sh) |
fallback_test.sh |
7-phase test harness — validates the entire fallback lifecycle via iptables simulation | On demand — maintenance window only |
━━━ HOW THE SCRIPTS RELATE ━━━
array_started.sh
│
└── starts fallback.sh (continuous loop)
│
├── Every FALLBACK_CHECK_INTERVAL seconds:
│ ping remote, ping internet
│ → determine state → act on containers + DDNS
│
└── Test path:
│
fallback_test.sh
│
├── Phase 1: pre-flight — both servers ready
├── Phase 2: iptables DROP rule → remote appears down
├── Phase 3: wait for fallback.sh to detect → FALLBACK state
├── Phase 4: verify Tier 1 containers started locally
├── Phase 5: remove DROP rule → remote reachable again
├── Phase 6: wait for fallback.sh to complete handback → NORMAL
└── Phase 7: verify Tier 1 containers stopped locally
fallback_test.sh contains no fallback logic. It exercises the real fallback.sh through
connectivity manipulation. Any change to fallback.sh is automatically reflected in the
test result.