docs: fix watchdog README and Manual — 15min schedule, dead vars, renamed vars, abort defaults, removed Required containers from stability tier

This commit is contained in:
Gmer4Lfe
2026-06-01 19:33:49 -04:00
parent d3454a3dab
commit 4b28529736
2 changed files with 35 additions and 38 deletions
+22 -25
View File
@@ -26,7 +26,7 @@ For the orchestrator that calls these scripts see `Orchestrators/watchdog_orches
All watchdog scripts use a two-tier output model: `echo` lines are always visible; All watchdog scripts use a two-tier output model: `echo` lines are always visible;
`log` lines only appear when `--log` is passed. `log` lines only appear when `--log` is passed.
All watchdog scripts are **single-pass scripts** called once per minute by the orchestrator. All watchdog scripts are **single-pass scripts** called once every 15 minutes by the orchestrator.
Without `--log`, only state transitions, warnings, errors, and the conclusion line are Without `--log`, only state transitions, warnings, errors, and the conclusion line are
visible. Per-check detail is suppressed on clean cycles. visible. Per-check detail is suppressed on clean cycles.
@@ -528,18 +528,17 @@ RAM below `SYS_WATCHDOG_MEM_GB` AND OOM kills this cycle ≥ `SYS_WATCHDOG_OOM_L
Both conditions required — RAM alone uses the standard strike system. Both conditions required — RAM alone uses the standard strike system.
OOM confirms the system is dying faster than watchdogs can heal. OOM confirms the system is dying faster than watchdogs can heal.
**Tier 3 — STANDARD (`SYS_WATCHDOG_STRIKES` consecutive failures → reboot)** **Tier 3 — STANDARD (`SYS_WATCHDOG_STRIKE_LIMIT` consecutive failures → reboot)**
| Check | Threshold | | Check | Threshold |
|-------|-----------| |-------|-----------|
| Free RAM | `SYS_WATCHDOG_MEM_GB` (reboot trigger — earlier tiers handled by resource_watchdog) | | Free RAM | `SYS_WATCHDOG_MEM_GB` (reboot trigger — earlier tiers handled by resource_watchdog) |
| Load average | `SYS_WATCHDOG_LOAD_MULTIPLIER` × cpu_count | | Load average | `SYS_WATCHDOG_LOAD_MULTIPLIER` × cpu_count |
| CPU temperature | `SYS_WATCHDOG_CPU_TEMP` | | CPU temperature | `SYS_WATCHDOG_CPU_TEMP_MAX` |
| Zombie processes | `SYS_WATCHDOG_ZOMBIES` | | Zombie processes | `SYS_WATCHDOG_ZOMBIE_LIMIT` |
| /var/log usage | `SYS_WATCHDOG_VAR_LOG_PCT` | | /var/log usage | `SYS_WATCHDOG_LOG_PCT` |
| /tmp usage | `SYS_WATCHDOG_TMP_PCT` | | /tmp usage | `SYS_WATCHDOG_TMP_PCT` |
| Array disk errors | mdstat error delta > 0 | | Array disk errors | mdstat error delta > 0 |
| NIC state | interface operstate != "up" | | NIC state | interface operstate != "up" |
| Required containers | containers in `HOST*_WATCHDOG_REQUIRED_CONTAINERS` (host*.conf) |
### RAM Tiers ### RAM Tiers
@@ -563,8 +562,8 @@ Prevent reboot — running them would risk data loss:
```bash ```bash
SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # ZFS pool degraded/faulted SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # ZFS pool degraded/faulted
SYS_WATCHDOG_ABORT_ON_PARITY=true # Parity check/rebuild running SYS_WATCHDOG_ABORT_ON_PARITY=false # aborting parity mid-check is worse than crashing
SYS_WATCHDOG_ABORT_ON_MOVER=true # Mover running SYS_WATCHDOG_ABORT_ON_MOVER=false # aborting move mid-run is worse than crashing
``` ```
Tier 1 CRITICAL bypasses all abort conditions — an imminent crash outweighs data Tier 1 CRITICAL bypasses all abort conditions — an imminent crash outweighs data
@@ -574,10 +573,10 @@ safety concerns.
```bash ```bash
SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 # window in hours SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 # window in hours
SYS_WATCHDOG_MAX_REBOOTS=3 # max reboots within the window SYS_WATCHDOG_REBOOT_LIMIT=3 # max reboots within the window
``` ```
If the server reboots `SYS_WATCHDOG_MAX_REBOOTS` times within the window, the watchdog If the server reboots `SYS_WATCHDOG_REBOOT_LIMIT` times within the window, the watchdog
switches from rebooting to notifying only. Prevents a boot loop where the watchdog switches from rebooting to notifying only. Prevents a boot loop where the watchdog
reboots → something crashes again immediately → reboot again. reboots → something crashes again immediately → reboot again.
@@ -629,7 +628,6 @@ HOST1_QBIT_USERNAME="admin"
HOST1_QBIT_PASSWORD="your-password" HOST1_QBIT_PASSWORD="your-password"
# ── Docker Watchdog ──────────────────────────────────────────────────────────── # ── Docker Watchdog ────────────────────────────────────────────────────────────
DOCKER_WATCHDOG_INTERVAL=900 # seconds between cycles (15 minutes)
WATCHDOG_STARTUP_GRACE=600 # seconds before restarts begin after boot WATCHDOG_STARTUP_GRACE=600 # seconds before restarts begin after boot
CONTAINER_DELAY=15 # seconds between dependency + dependent restart CONTAINER_DELAY=15 # seconds between dependency + dependent restart
@@ -684,10 +682,9 @@ STORAGE_WATCHDOG_STATE_FILE="/tmp/storage_watchdog_state.db"
# Reboot trigger only — warn/shutdown/recover RAM tiers handled by resource_watchdog.sh # Reboot trigger only — warn/shutdown/recover RAM tiers handled by resource_watchdog.sh
# RW_RAM_HARD_GB > SYS_WATCHDOG_MEM_GB always (resource_watchdog acts before watchdog reboots) # RW_RAM_HARD_GB > SYS_WATCHDOG_MEM_GB always (resource_watchdog acts before watchdog reboots)
SYS_WATCHDOG_STRIKE_LIMIT=2 SYS_WATCHDOG_STRIKE_LIMIT=2
SYSTEM_WATCHDOG_INTERVAL=300
SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 SYS_WATCHDOG_REBOOT_WINDOW_HRS=12
SYS_WATCHDOG_MAX_REBOOTS=3 SYS_WATCHDOG_REBOOT_LIMIT=3
SYS_WATCHDOG_OOM_LIMIT=3 SYS_WATCHDOG_OOM_LIMIT=3
SYS_WATCHDOG_MEM_GB=4 # strike system → reboot (last resort — below resource_watchdog hard stop) SYS_WATCHDOG_MEM_GB=4 # strike system → reboot (last resort — below resource_watchdog hard stop)
@@ -695,20 +692,20 @@ SYS_WATCHDOG_MEM_GB=4 # strike system → reboot (last resort — b
SYS_WATCHDOG_ROOTFS_CRITICAL_PCT=99 SYS_WATCHDOG_ROOTFS_CRITICAL_PCT=99
SYS_WATCHDOG_FD_CRITICAL_PCT=95 SYS_WATCHDOG_FD_CRITICAL_PCT=95
SYS_WATCHDOG_LOAD_MULTIPLIER=4 SYS_WATCHDOG_LOAD_MULTIPLIER=4
SYS_WATCHDOG_CPU_TEMP=85 SYS_WATCHDOG_CPU_TEMP_MAX=95
SYS_WATCHDOG_ZOMBIES=20 SYS_WATCHDOG_ZOMBIE_LIMIT=50
SYS_WATCHDOG_VAR_LOG_PCT=80 SYS_WATCHDOG_LOG_PCT=95
SYS_WATCHDOG_TMP_PCT=85 SYS_WATCHDOG_TMP_PCT=90
SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true
SYS_WATCHDOG_ABORT_ON_PARITY=true SYS_WATCHDOG_ABORT_ON_PARITY=false
SYS_WATCHDOG_ABORT_ON_MOVER=true SYS_WATCHDOG_ABORT_ON_MOVER=false
# State files: # State files (all in State_Files/ — survive reboots):
SYS_WATCHDOG_STATE_FILE="/tmp/system_watchdog_state.db" SYS_WATCHDOG_STATE_FILE="State_Files/system_watchdog_state.db"
SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db" DOCKER_WATCHDOG_FAILED_FILE="State_Files/docker_watchdog_failed.db"
SYS_WATCHDOG_REBOOT_LOG="/boot/config/system_watchdog_reboots.db" SYS_WATCHDOG_REBOOT_LOG="State_Files/system_watchdog_reboots.db"
SYS_WATCHDOG_OOM_FILE="/tmp/system_watchdog_oom.db" SYS_WATCHDOG_OOM_FILE="State_Files/system_watchdog_oom.db"
# ── Network Watchdog ─────────────────────────────────────────────────────────── # ── Network Watchdog ───────────────────────────────────────────────────────────
NETWORK_WATCHDOG_ENABLED=true NETWORK_WATCHDOG_ENABLED=true
+13 -13
View File
@@ -1,10 +1,10 @@
# ━━━━━ WATCHDOGS ━━━━━ # ━━━━━ WATCHDOGS ━━━━━
**Four tiers that run every minute through `watchdog_orchestrator.sh`, **Four tiers that run every 15 minutes through `watchdog_orchestrator.sh`,
each with a clear lane:** reduce system pressure → heal containers → check system components → each with a clear lane:** reduce system pressure → heal containers → check system components →
reboot if nothing else worked. The orchestrator calls them in order, once per cron cycle. reboot if nothing else worked. The orchestrator calls them in order, once per cron cycle.
System component checks (storage, WebGUI) run inside a thin `system_watchdog.sh` orchestrator System component checks (storage, WebGUI, network) run inside a thin `system_watchdog.sh`
that can also be run standalone. orchestrator that can also be run standalone.
--- ---
@@ -79,7 +79,7 @@ System component checks run after containers are healed. Stability watchdog runs
reboot is always the last option, not the first. reboot is always the last option, not the first.
**None of these scripts run standalone loops.** Each is a single-pass script called **None of these scripts run standalone loops.** Each is a single-pass script called
once per minute by `Orchestrators/watchdog_orchestrator.sh`. The orchestrator handles once every 15 minutes by `Orchestrators/watchdog_orchestrator.sh`. The orchestrator handles
startup grace, overlap protection, heartbeat, and sequencing. `system_watchdog.sh` can startup grace, overlap protection, heartbeat, and sequencing. `system_watchdog.sh` can
also be run standalone to check all system component watchdogs at once. also be run standalone to check all system component watchdogs at once.
@@ -89,7 +89,7 @@ also be run standalone to check all system component watchdogs at once.
``` ```
Orchestrators/ Orchestrators/
watchdog_orchestrator.sh ──────────────────► resource_watchdog.sh (1st — every minute) watchdog_orchestrator.sh ──────────────────► resource_watchdog.sh (1st — every 15 min)
──────────────────► docker_watchdog.sh (2nd) ──────────────────► docker_watchdog.sh (2nd)
──────────────────► system_watchdog.sh (3rd — thin orchestrator) ──────────────────► system_watchdog.sh (3rd — thin orchestrator)
──────────────────► stability_watchdog.sh (4th — last resort) ──────────────────► stability_watchdog.sh (4th — last resort)
@@ -117,13 +117,13 @@ Docker_Essentials/
| Script | Role | Called By | | Script | Role | Called By |
|--------|------|-----------| |--------|------|-----------|
| `resource_watchdog.sh` | Three-level pressure reduction — throttle, pause, stop | `watchdog_orchestrator.sh` — 1st every minute | | `resource_watchdog.sh` | Three-level pressure reduction — throttle, pause, stop | `watchdog_orchestrator.sh` — 1st every 15 min |
| `docker_watchdog.sh` | Two-tier container healing — memory, CPU, HTTP, required | `watchdog_orchestrator.sh` — 2nd every minute | | `docker_watchdog.sh` | Two-tier container healing — memory, CPU, HTTP, required | `watchdog_orchestrator.sh` — 2nd every 15 min |
| `system_watchdog.sh` | Thin orchestrator — runs System/ component watchdogs in sequence | `watchdog_orchestrator.sh` — 3rd every minute | | `system_watchdog.sh` | Thin orchestrator — runs System/ component watchdogs in sequence | `watchdog_orchestrator.sh` — 3rd every 15 min |
| `stability_watchdog.sh` | Last-resort server watchdog — reboots when healing has failed | `watchdog_orchestrator.sh` — 4th every minute | | `stability_watchdog.sh` | Last-resort server watchdog — reboots when healing has failed | `watchdog_orchestrator.sh` — 4th every 15 min |
| `System/storage_watchdog.sh` | Pool growth rate + runaway log detection and remediation | `system_watchdog.sh` — every minute | | `System/storage_watchdog.sh` | Pool growth rate + runaway log detection and remediation | `system_watchdog.sh` — every 15 min |
| `System/webgui_watchdog.sh` | WebGUI availability — nginx → php-fpm → emhttp escalation | `system_watchdog.sh` — every minute | | `System/webgui_watchdog.sh` | WebGUI availability — nginx → php-fpm → emhttp escalation | `system_watchdog.sh` — every 15 min |
| `System/network_watchdog.sh` | Internet reachability, DDNS sync, Tailscale, NPM proxy | `system_watchdog.sh` — every minute | | `System/network_watchdog.sh` | Internet reachability, DDNS sync, Tailscale, NPM proxy | `system_watchdog.sh` — every 15 min |
> `watchdog_orchestrator.sh` is in `Orchestrators/`. `watchdog_skip_list_manager.sh` > `watchdog_orchestrator.sh` is in `Orchestrators/`. `watchdog_skip_list_manager.sh`
> is in `Tools/`. Neither is a watchdog — they sit at the edges of this system. > is in `Tools/`. Neither is a watchdog — they sit at the edges of this system.
@@ -133,7 +133,7 @@ Docker_Essentials/
## ━━━ HOW THE SCRIPTS RELATE ━━━ ## ━━━ HOW THE SCRIPTS RELATE ━━━
``` ```
Every minute — watchdog_orchestrator.sh fires: Every 15 minutes — watchdog_orchestrator.sh fires:
Step 1 — resource_watchdog.sh Step 1 — resource_watchdog.sh
│ RAM/load OK → pass through (no action) │ RAM/load OK → pass through (no action)