feat: slskd reconnect guard in downloaders_reset, mass v2 sync

- downloaders_reset: connection check block before slskd API sections;
  triggers PUT /api/v0/server reconnect if disconnected, polls 60s,
  gates Stuck Searches and Dead Transfer Records on SLSKD_CONNECTED
- Sync all modified/new/deleted files from v2 refactor across Docker_Essentials,
  Media, Monitors, Partnership, Rsync, Tools, Transcodes, unRAID_Essentials,
  common.sh, master confs, and new Manual/README docs
This commit is contained in:
Gmer4Lfe
2026-05-19 20:00:10 -04:00
parent 5cb16d4b18
commit e13f2fa14f
81 changed files with 12164 additions and 10656 deletions
+275
View File
@@ -0,0 +1,275 @@
# ━━━━━ FALLBACK ━━━━━
Mutual automatic failover between two independent unRAID servers. When one goes down the
other starts its containers, cuts over DNS, and keeps users online. When it comes back
everything hands back in the correct sequence — covering DDNS stops, containers stop, rsync
writeback runs, containers start on the primary, primary DDNS starts last — so users hit the
returning server only after it's actually ready.
> **Built from scratch. Refined through a year of production testing.** The DDNS sequencing
> and handback order were the hardest parts to get right. Both directions are exercised
> regularly with `fallback_test.sh`.
---
## ━━━ THE PROBLEM THAT BUILT THIS ━━━
Running a self-hosted stack means being the operator. These are the specific problems
that drove this build:
**A Single Point of Failure for an Entire Household**
HOST1 runs Emby, NPM, Authelia, NextCloud, VaultWarden, and every service the household
uses daily. When HOST1 goes down — even briefly — all of those services go down with it.
The fix: a second server with mirrored critical data that covers the first automatically.
From a user's perspective, a brief interruption and then everything is back.
**DNS Cutting Over Before the Server Was Ready**
Early attempts started containers on the covering server then updated DNS. Problem: DNS
propagated in under a minute. Users hit the new IP before Emby had finished starting,
before Authelia had loaded its sessions, before NPM had loaded its proxy configurations.
The fix: warm standby for the auth stack. NPM, LLDAP, and Authelia run actively on both
servers at all times. When DNS cuts over, auth is already running and ready.
**Split Brain DNS During Handback**
When HOST1 returned, the obvious sequence was: start HOST1 containers, then switch DNS
back. Problem: between "start containers" and "DNS switches" both servers' DDNS containers
were running, both updating the same domain with different IPs. Users got routed randomly
between servers — intermittent auth failures, no clear error state anywhere.
The fix: stop DDNS on the covering server first, before anything else moves. There is
never a window where two DDNS containers update the same record.
**Rsync Running Into Active Container I/O**
Syncing data back while containers were still running — to minimise downtime — produced
slower transfers, potential file inconsistency, and database dirty state risk.
The fix: stop containers before syncing. The outage window is only the rsync duration —
typically minutes. Clean static source at full bandwidth, predictable state every time.
**No Way to Validate the System Before Needing It**
A failover system that has never been tested is not a failover system — it is a hope.
The fix: `fallback_test.sh` — a controlled simulation using an iptables DROP rule to make
the remote appear unreachable, triggering the full sequence without taking anything offline.
A safety trap removes the rule on any exit — crash, error, ctrl-c, or clean completion.
---
## ━━━ THE TWO-SERVER SETUP ━━━
```
HOST1 — unRAID-Gmer4Lfe
Hardware: Threadripper 1950X, 128GB RAM, ZFS cache pools
Location: Primary site
DDNS: Gmer4Lfe.com
Role: Primary — full service stack + source of truth for Movies/Shows/Music
HOST2 — unRAID-Jayred365
Hardware: Intel i5 10th gen, 64GB RAM
Location: Remote — 50 miles away
DDNS: Gmer4Lfe.us
Role: Secondary — own stack + covers HOST1 + mirrors critical data
```
**Hardware does not need to match.** Everything is accessed through `/mnt/user/` — unRAID's
fused share layer. HOST1 has a Threadripper with ZFS pools. HOST2 has completely different
hardware. Fallback containers on HOST2 mount `/mnt/user/Movies` and see mirrored data
because the share names match. The hardware underneath is irrelevant.
**What must match between servers:**
```
Share names /mnt/user/Movies must exist on both servers (mirrored data)
Container names "Emby" on HOST2 must be the container HOST2 starts for HOST1
Network names Docker custom networks must match for NPM routing to work
```
### Split Source of Truth — No Conflicts
Both servers run arr instances simultaneously with zero conflict — they manage completely
different shares:
```
HOST1 owns: Movies (Radarr), Tv_Shows (Sonarr), Music (Lidarr)
HOST2 owns: Anime_Movies (his Radarr), Anime_Shows (his Sonarr)
Each server mirrors the other's shares continuously via rsync.
```
The rule: never run two arr instances against the same share simultaneously. Different arrs
managing different shares is fine. When HOST2 runs HOST1's arrs during a Tier 4 fallback,
HOST1's Tdarr is not running (HOST1 is down) — no conflict.
### Auth Stack — Warm on Both Servers
NPM, LLDAP, and Authelia run actively on both servers at all times. HOST2 needs them running
to serve his own users daily — this is not a fallback-only configuration. HOST1 is source of
truth: all changes mirror to HOST2 every 15 minutes via critical sync.
When DNS cuts over, auth is already running on the covering server. The 30-60 second dead
zone where auth is coming up after DNS has already switched does not exist.
---
## ━━━ HOW IT WORKS ━━━
Both servers run `fallback.sh` independently as a continuous background process. Each server
makes all decisions from two pings every `FALLBACK_CHECK_INTERVAL` seconds:
```bash
ping REMOTE_TAILSCALE_IP # is the other server reachable?
ping EXTERNAL_IP # do I have internet? (default: 8.8.8.8)
```
No SSH signaling between servers. No shared state file. No election algorithm. Each server
acts entirely from its own network perspective.
**States:**
| State | Remote | Internet | Action |
|-------|--------|----------|--------|
| NORMAL | up | up | Silent — own containers, own DDNS on |
| FALLBACK | down | up | Start tier containers, cut DDNS over |
| NO_INTERNET | — | down | Stop own DDNS immediately, wait |
| DARK | down | down | Same as NO_INTERNET — cannot determine cause |
---
## ━━━ DDNS — THE CRITICAL PART ━━━
> **This took a year to get right. Do not change the sequencing.**
`fallback.sh` is the sole authority over when any DDNS container starts. Network state
returning is not permission to start DDNS. Only completion of the full handback sequence
grants that permission.
```
ONE DOMAIN → ONE DDNS ACTIVE → AT ALL TIMES
Gmer4Lfe.com → HOST1's DDNS normally → HOST2's DDNS during HOST1 outage
Gmer4Lfe.us → HOST2's DDNS normally → HOST1's DDNS during HOST2 outage
Own DDNS: ON when this server has internet. OFF when internet is lost.
Remote DDNS: ON as Tier 1 fallback action. OFF as FIRST handback action.
Auto-start: NEVER — DDNS never starts automatically on internet return.
```
**Why auto-start is forbidden:** If HOST1 lost internet and its DDNS auto-started when
internet returned, there is a window where both servers are updating the same domain with
different IPs. Users get routed randomly — some to the primary with fresh data, some to
the covering server. Authentication sessions don't transfer between servers. This is split
brain and produces the most confusing symptoms: intermittent auth failures with no clear
error state anywhere.
The brief gap where neither DDNS is updating the record is intentional. DNS TTL caches the
last value. During the rsync + container start window, cached DNS still routes users to the
covering server where containers are still running. By the time the cache expires, the
primary's DDNS has started and the record points at the right server.
---
## ━━━ TIERED FALLBACK ━━━
Starting the full stack for a 5-minute power blip wastes resources — most brief outages
resolve before Tier 2 would even activate. Tiers start only what is needed for the actual
outage duration.
| Tier | Delay | Coverage | Why This Timing |
|------|-------|----------|-----------------|
| Tier 1 | Immediate | Vital services + Live TV | People are watching — cannot wait 4 hours |
| Tier 2 | 4hr (HOST*_TIER2_DELAY) | NextCloud, Immich, Jellyseerr | 4hr covers most ISP and power events |
| Tier 3 | 12hr (HOST*_TIER3_DELAY) | Dashboard, AdGuard, Git, Collabora | Secondary — useful but not daily-critical |
| Tier 4 | 24hr (HOST*_TIER4_DELAY) | Arrs + downloaders | Significant I/O — only worth starting at 24hr |
Tier 1 always includes the remote domain's DDNS container as the first entry — DNS
coverage happens before any other container starts.
For the actual container lists and tier delay values, see `Manual-Fallback.md`.
---
## ━━━ HANDBACK SEQUENCE ━━━
When the remote server returns after a FALLBACK event. Every step has a reason.
Do not reorder.
1. **Strike confirmation** — FALLBACK_HANDBACK_STRIKES consecutive remote-up checks before
handback begins. Prevents false triggers from brief network recovery.
2. **Pre-flight checks** — version parity, remote array mounted, remote Docker daemon up.
Any failure aborts and retries next cycle.
3. **Staged reverse handback: Tier 4 → 3 → 2** — Emby and vital services stay on the
covering server serving users throughout this phase. Each tier: stop local containers
→ rsync writeback (if outage exceeded tier threshold) → start on remote.
4. **DDNS handoff** — stop remote DDNS immediately before Tier 1 goes down. This prevents
split brain during the Tier 1 rsync window.
5. **Tier 1 handback** — stop local vital services, rsync writeback, start on remote.
6. **Start remote DDNS last** — DNS cuts back to the primary only after all containers
are confirmed running.
7. **Return to NORMAL** — state file reset, own DDNS restored if it was stopped.
---
## ━━━ MUTUAL FALLBACK — BOTH DIRECTIONS ━━━
The same `fallback.sh` handles both directions without any code changes. `detect_hosts()`
determines which server is local and which is remote at runtime, then selects the correct
container arrays and tier delays from config via MY_ID.
```
HOST2 covers HOST1: FALLBACK_HOST2_COVERS_HOST1_TIER* (in master_host2.conf)
HOST1 covers HOST2: FALLBACK_HOST1_COVERS_HOST2_TIER* (in master_host1.conf)
```
Both servers run identical scripts. MY_ID selects the correct arrays. No hostname
comparisons anywhere in the script code.
---
## ━━━ INDEPENDENCE — ALWAYS ONE RSYNC STOP AWAY ━━━
HOST2 is designed to be fully independent if needed. If HOST2 ever wants to separate from
HOST1: stop HOST1 pushing data. Any changes HOST2 makes to his own data stick permanently.
His server becomes fully independent immediately — no script changes, no migration, no data
movement required. The fallback and rsync scripts are configuration-driven.
---
## ━━━ SCRIPTS IN THIS FOLDER ━━━
| Script | Role | When It Runs |
|--------|------|--------------|
| `fallback.sh` | Continuous state machine — monitors remote, manages fallback and handback | Continuously (started by `array_start.sh`) |
| `fallback_test.sh` | 7-phase test harness — validates the entire fallback lifecycle via iptables simulation | On demand — maintenance window only |
---
## ━━━ HOW THE SCRIPTS RELATE ━━━
```
array_start.sh
└── starts fallback.sh (continuous loop)
├── Every FALLBACK_CHECK_INTERVAL seconds:
│ ping remote, ping internet
│ → determine state → act on containers + DDNS
└── Test path:
fallback_test.sh
├── Phase 1: pre-flight — both servers ready
├── Phase 2: iptables DROP rule → remote appears down
├── Phase 3: wait for fallback.sh to detect → FALLBACK state
├── Phase 4: verify Tier 1 containers started locally
├── Phase 5: remove DROP rule → remote reachable again
├── Phase 6: wait for fallback.sh to complete handback → NORMAL
└── Phase 7: verify Tier 1 containers stopped locally
```
`fallback_test.sh` contains no fallback logic. It exercises the real `fallback.sh` through
connectivity manipulation. Any change to `fallback.sh` is automatically reflected in the
test result.