Files
Varaverk/Watchdogs/README-Watchdogs.md
T
Gmer4Lfe bf3e7cc2c4 Storage-mode awareness pass + doc update for System_Essentials through Partnership
All state/data file paths in scripts and PHP now resolve via STATE_DIR / DATA_DIR /
PERSISTENT_CONF_CACHE instead of hardcoded /boot/config/ or /tmp/ paths, so the
ecosystem works in both internal and appdata storage modes.

PHP layer (watchdog.php, partnership.php, fallback.php, monitor.php, snapshot.php,
config.php): all state reads switched to STATE_DIR constant; remote state reads use
the new vv_remote_state_cmd() helper which resolves the remote's SCRIPTS_DIR via
their varaverk.cfg before building the path.

conf_sync.sh: fixed SCRIPTS_ROOT → SCRIPTS_DIR bug on MY_CONF path; added
_remote_scripts_dir() to resolve partner's SCRIPTS_DIR before SCP pull.

fallback.php page: added controls card (PARTNERSHIP_ENABLED, FALLBACK_ENABLED,
FALLBACK_RSYNC_ENABLED toggles), status grid, and settings card.

README and Manual updated for System_Essentials, Watchdogs, Fallback, Rsync,
Media, Monitors, Orchestrators, Partnership: added new scripts (conf_sync,
conf_cache_save/restore, conf_cache_watchdog, play_state_sync, start_webhook_listener,
upgrade_webhook_handler), corrected all stale /boot/config/ state file paths to
$STATE_DIR/$DATA_DIR, noted webgui/php_fpm/mover/user_scripts scripts moved to
Plugin/unraid/System_Essentials, fixed start_webhook_listener.sh header (Node.js,
not PHP -S).
2026-06-19 19:32:39 -04:00

13 KiB

━━━━━ WATCHDOGS ━━━━━

Four tiers that run every 15 minutes through watchdog_orchestrator.sh, each with a clear lane: reduce system pressure → heal containers → check system components → reboot if nothing else worked. The orchestrator calls them in order, once per cron cycle. System component checks (storage, WebGUI, network, conf cache) run inside a thin system_watchdog.sh orchestrator that can also be run standalone.


━━━ THE PROBLEM THAT BUILT THIS ━━━

Container Memory Leaks Going Undetected for Days Emby's transcode session handling occasionally leaks memory. SABnzbd's Python process expands slowly across downloads. Neither crashes dramatically — they just consume more RAM until the system starts swapping. Docker reports both containers as Up 14 days. Nothing alerts. By the time someone notices, the system has been degraded for hours.

Fix: docker_watchdog.sh — hard per-container memory ceilings. When a container exceeds its limit the watchdog restarts it immediately. No strikes, no waiting. A memory leak is not a transient spike.

Containers That Look Running But Aren't Responding Docker reports a container as Up while its application layer has been frozen for hours. The reverse proxy forwards traffic to a service that returns nothing. Users see a broken page. Docker sees a healthy container.

Fix: docker_watchdog.sh — HTTP health checks on the actual service port every cycle. Two consecutive non-responses trigger a restart. Process running and service responding are not the same thing.

System Pressure Causing Docker Watchdog to Undo Itself Resource pressure builds. RAM drops. docker_watchdog.sh tries to restart a container into a system that's already swapping — the restarted container fails immediately and goes on the skip list. The real problem (RAM pressure) is never addressed.

Fix: resource_watchdog.sh runs first in the orchestrator sequence. At Level 1 it throttles downloaders. At Level 2 it pauses non-critical containers. At Level 3 it stops heavy services and signals docker_watchdog to defer all restarts. By the time docker_watchdog runs, the system has breathing room to actually heal.

Runaway Log Files Filling a Pool Before Anyone Notices A game server container was offline for a year, restarted for a weekend, and wrote 130GB of logs to the appdata pool. The pool grew 13% in days. No alert fired — nothing was watching for growth at the data level, only at the container level.

Fix: storage_watchdog.sh — growth-rate scan of every container's appdata directory every cycle. No per-container configuration required. Runaway growth gets three cycles to be confirmed, then alerts and (optionally) truncates logs automatically.

Rootfs at 99% With SSH Failing Silently Rootfs fills. SSH stops accepting new connections. Docker can't write log files. State files fail silently. The server is functionally dead but still technically running. Nothing in the container layer can detect or recover from this — it requires a reboot.

Fix: stability_watchdog.sh — watches the server itself: RAM, CPU, disk, kernel, daemon health. Only script in the stack authorized to reboot. Runs last in the orchestrator sequence so container healing and pressure reduction always get a chance first.

Partner Conf Backup Going Stale During a Long Outage The partner goes offline overnight. On the next reboot (planned maintenance), the persistent conf backup saved at shutdown is stale — it reflects the state from before the outage, not the most recent live values. conf_cache_restore.sh loads it faithfully, but fallback.sh may be working with old tier delays or container names.

Fix: conf_cache_watchdog.sh — refreshes the persistent backup from the RAM cache every 15 minutes while the partner is offline, keeping it current throughout the outage.


━━━ WHAT THIS FOLDER DOES ━━━

Four tiers. Fixed execution order via watchdog_orchestrator.sh.

Pressure reduction    resource_watchdog.sh   — throttle/pause/stop before healing fails
Container healing     docker_watchdog.sh     — memory, CPU, HTTP, required containers
System components     system_watchdog.sh     — thin orchestrator: system component health
  └─ System/          storage_watchdog.sh    — pool growth rate + runaway log detection
                      network_watchdog.sh    — internet, DDNS sync, Tailscale, NPM proxy
                      conf_cache_watchdog.sh — maintain persistent partner conf backup
Last resort           stability_watchdog.sh  — reboot only when nothing else can recover

Plugin/unraid/Watchdogs/System/webgui_watchdog.sh is also called by system_watchdog.sh but lives in the plugin tree because it calls Unraid-specific service commands (nginx, php-fpm, emhttp via platform adapter). See Plugin/unraid/Watchdogs/System/.

The execution order is the design. Resource pressure is reduced before docker_watchdog attempts restarts — containers restarted into a pressure-bound system just fail again. System component checks run after containers are healed. Stability watchdog runs last — reboot is always the last option, not the first.

None of these scripts run standalone loops. Each is a single-pass script called once every 15 minutes by Orchestrators/watchdog_orchestrator.sh. The orchestrator handles startup grace, overlap protection, heartbeat, and sequencing. system_watchdog.sh can also be run standalone to check all system component watchdogs at once.


━━━ RELATIONSHIP TO OTHER FOLDERS ━━━

Orchestrators/
  watchdog_orchestrator.sh ──────────────────► resource_watchdog.sh    (1st — every 15 min)
                            ──────────────────► docker_watchdog.sh     (2nd)
                            ──────────────────► system_watchdog.sh     (3rd — thin orchestrator)
                            ──────────────────► stability_watchdog.sh  (4th — last resort)

Watchdogs/System/              ← called by system_watchdog.sh each cycle
  storage_watchdog.sh          pool growth rate + runaway log detection
  network_watchdog.sh          internet reachability, DDNS sync, Tailscale, NPM proxy
  conf_cache_watchdog.sh       maintain persistent partner conf backup during outages

Plugin/unraid/Watchdogs/System/
  webgui_watchdog.sh           WebGUI availability — nginx → php-fpm → emhttp escalation
                               (Unraid-specific platform calls — lives in plugin tree)

Plugin/unraid/System_Essentials/
  unraid_api_key_renew.sh      re-registers the Varaverk plugin API key each watchdog cycle
                               (inserted between docker_watchdog and stability_watchdog)

Tools/
  watchdog_skip_list_manager.sh ◄────────────── docker_watchdog.sh writes skip list
    (operator utility — inspect + clear after fixing a crash-looping container)

Docker_Essentials/
  All container lifecycle scripts (daily restart, updates, network) — unaffected.
  docker_watchdog.sh coordinates with them via shared state, not direct calls.

System_Essentials/
  conf_cache_save.sh / conf_cache_restore.sh — at-stop/at-start bookends for the
  persistent backup that conf_cache_watchdog.sh keeps current in between.

watchdog_orchestrator.sh stays in Orchestrators/ — it's a job runner, not a watchdog. watchdog_skip_list_manager.sh stays in Tools/ — it's an operator utility, not a watchdog.


━━━ SCRIPTS IN THIS FOLDER ━━━

Script Role Called By
resource_watchdog.sh Three-level pressure reduction — throttle, pause, stop watchdog_orchestrator.sh — 1st every 15 min
docker_watchdog.sh Two-tier container healing — memory, CPU, HTTP, required watchdog_orchestrator.sh — 2nd every 15 min
system_watchdog.sh Thin orchestrator — runs SYSTEM_WATCHDOG_SCRIPTS in sequence watchdog_orchestrator.sh — 3rd every 15 min
stability_watchdog.sh Last-resort server watchdog — reboots when healing has failed watchdog_orchestrator.sh — 4th every 15 min
System/storage_watchdog.sh Pool growth rate + runaway log detection and remediation system_watchdog.sh — every 15 min
System/network_watchdog.sh Internet reachability, DDNS sync, Tailscale, NPM proxy system_watchdog.sh — every 15 min
System/conf_cache_watchdog.sh Refresh persistent partner conf backup while partner is offline system_watchdog.sh — every 15 min

Plugin/unraid/Watchdogs/System/webgui_watchdog.sh is also called by system_watchdog.sh but is not in this folder. watchdog_orchestrator.sh is in Orchestrators/. watchdog_skip_list_manager.sh is in Tools/. Neither is a watchdog.


━━━ HOW THE SCRIPTS RELATE ━━━

Every 15 minutes — watchdog_orchestrator.sh fires:

  Step 1 — resource_watchdog.sh
    │  RAM/load OK → pass through (no action)
    │  Level 1 (soft):   throttle SABnzbd + qBit download speeds
    │  Level 2 (medium): further throttle + docker pause non-critical containers
    │  Level 3 (hard):   docker stop optional services
    │                    writes mem_shutdown_active=true → RW_STATE_FILE
    │                                  ↓
  Step 2 — docker_watchdog.sh
    │  reads RW_STATE_FILE — if mem_shutdown_active=true: skip all restarts
    │
    │  Tier 1 — explicit per-container checks (configured in host*.conf):
    │    memory hard limits  → immediate restart (no strikes)
    │    CPU high sustained  → 2-strike restart
    │    HTTP non-response   → 2-strike restart
    │    required stopped    → restart (dependency ordering respected)
    │
    │  Tier 2 — global scan of all running containers:
    │    unhealthy / OOM / crashloop / dead / non-zero exit → restart
    │    N restarts in window → skip list + critical notify → human required
    │
    │  (skip list management) → Tools/watchdog_skip_list_manager.sh
    │
  Step 3 — system_watchdog.sh (thin orchestrator)
    │  runs SYSTEM_WATCHDOG_SCRIPTS from master.conf sequentially:
    │
    │  └─ storage_watchdog.sh
    │       growth rate scan: du -sm appdata/* → compare to previous cycle baseline
    │         growth > WATCHDOG_APPDATA_GROWTH_GB → 3-strike warn → alert
    │       log file scan: find *.log > WATCHDOG_APPDATA_LOG_MAX_GB
    │         oversize log found → 3-strike warn → truncate (if enabled) or alert
    │
    │  └─ Plugin/unraid/Watchdogs/System/webgui_watchdog.sh
    │       curl check → WebGUI responding → exit 0 (silent)
    │       not responding → nginx restart → wait → recheck
    │       still down    → php-fpm restart → wait → recheck
    │       still down    → emhttp restart → wait → recheck
    │       all three failed → critical notify, manual intervention needed
    │
    │  └─ network_watchdog.sh
    │       internet check → fail = alert + skip remaining (prevents false positives)
    │       DDNS: public IP vs dig @1.1.1.1 → mismatch = restart Cloudflare DDNS container
    │       Tailscale: status Running → pass; not running = notify (no auto-restart)
    │       NPM proxy: curl external URL → 2-strike system → restart NginxProxyManager
    │
    │  └─ conf_cache_watchdog.sh
    │       remote online  → remove persistent backup (conf_sync gets fresh on next boot)
    │       remote offline → refresh backup from RAM cache → backup stays current
    │       silent when remote is online and no backup exists (normal state)
    │
    │  (+ Plugin/unraid/System_Essentials/unraid_api_key_renew.sh — between steps 3 and 4)
    │
  Step 4 — stability_watchdog.sh
       checks the server itself — RAM, CPU temp, rootfs, FDs, kernel, daemon
       Tier 1 CRITICAL  → immediate reboot (no strikes)
       Tier 2 URGENT    → reboot if OOM confirmed
       Tier 3 STANDARD  → N consecutive failures → reboot
       Abort conditions → ZFS unhealthy / parity running / mover active
       Rate limit       → max N reboots per window before switching to notify

State file coordination between scripts:

State File Written By Read By Purpose
RW_STATE_FILE resource_watchdog.sh docker_watchdog.sh mem_shutdown_active flag — defer restarts during RAM emergency
SYS_WATCHDOG_STATE_FILE stability_watchdog.sh docker_watchdog.sh watchdog_cycle heartbeat — stale guard (2hr timeout)
WATCHDOG_STATE_FILE docker_watchdog.sh itself CPU/HTTP strike counts per container
DOCKER_WATCHDOG_FAILED_FILE docker_watchdog.sh watchdog_skip_list_manager.sh Container skip list
STORAGE_WATCHDOG_STATE_FILE System/storage_watchdog.sh itself Growth + log strike counts
WATCHDOG_APPDATA_GROWTH_FILE System/storage_watchdog.sh itself Per-container size baseline for growth rate
NETWORK_WATCHDOG_NPM_STATE_FILE System/network_watchdog.sh itself NPM proxy strike count

All state files are in $STATE_DIR (adapts to storage mode). See master.conf for actual variable values.