diff --git a/Docker_Essentials/Manual-Docker_Essentials.md b/Docker_Essentials/Manual-Docker_Essentials.md index 1e11232..3e130ae 100644 --- a/Docker_Essentials/Manual-Docker_Essentials.md +++ b/Docker_Essentials/Manual-Docker_Essentials.md @@ -6,235 +6,9 @@ Configuration reference, setup procedures, and operational workflows. For folder overview and design philosophy see `README-Docker_Essentials.md`. For per-script detail see the script headers directly. ---- - -## ━━━ WATCHDOG CONFIGURATION ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ - -All watchdog configuration lives in `host*.conf` (per-container lists) and -`master.conf` (shared thresholds and toggles). `detect_hosts()` aliases all -`HOST1_` / `HOST2_` prefixed vars to their unprefixed names at runtime — scripts -always read the right values for the server they're running on. - ---- - -### ── Memory Hard Limits ──────────────────────────────────────────────────────── - -```bash -# host1.conf -# Format: "ContainerName:LimitInMB" -# Immediate restart when exceeded — no strike system. Memory leaks are not spikes. -# -# Sizing guidance: -# Check normal peak: docker stats ContainerName -# Set limit at ~150-200% of normal peak -# Emby peaks ~12GB under heavy transcode load → 18GB gives headroom -# without triggering on legitimate load spikes -# -HOST1_WATCHDOG_CONTAINERS=( - "Emby:18432" # 18GB — peaks ~12GB under heavy transcode load - "LidaTube:6144" # 6GB — YouTube downloader, grows with large queues - "Tdarr:6144" # 6GB — video transcoder, memory-intensive by nature - "Code-Server:1024" # 1GB — IDE, should be light; 1GB is generous -) -``` - -A soft warning fires at `SOFT_MEM_THRESHOLD=80` percent of the hard limit — early -visibility into a container approaching its ceiling before a restart is triggered. - ---- - -### ── CPU Thresholds ──────────────────────────────────────────────────────────── - -```bash -# master.conf -# CPU is normalised against total core count — meaningful regardless of hardware. -# -# Why normalised: -# 85% on one core of a 16-core machine = 5.3% normalised → ignore it -# 85% normalised on a 16-core machine = 13.6 cores worth → runaway process -# -# CPU uses a STRIKE SYSTEM — not immediate restart like memory. -# Brief spikes are normal (Tdarr encoding, Emby transcoding, SABnzbd unpacking). -# The strike system ignores spikes and acts on sustained high usage. -# -# Strike 1: above HARD_CPU_THRESHOLD this cycle → warn, increment strike -# Strike 2: above threshold next cycle → restart, reset counter -# Recovery: drops below threshold any cycle → reset counter to 0 -# -SOFT_CPU_THRESHOLD=50 # warn at 50% normalised — informational only -HARD_CPU_THRESHOLD=85 # strike at 85% normalised -CPU_FAIL_LIMIT=2 # consecutive strikes before restart -``` - ---- - -### ── HTTP Health Checks ──────────────────────────────────────────────────────── - -```bash -# host1.conf -# Format: "ContainerName:http://host:port/optional-path" -# Hits the actual service endpoint on every watchdog cycle. -# "Container running" and "service responding" are not the same thing. -# -# Uses a STRIKE SYSTEM — network hiccups and brief restarts happen. -# Two consecutive non-responses before acting prevents false positives. -# -# The path can be a lightweight health endpoint or the root URL. -# Docker's own HEALTHCHECK requires the image to define it — most don't. -# These checks work regardless of what the image defines. -# -HOST1_WATCHDOG_CONTAINER_URLS=( - "Emby:http://localhost:8096" # Emby WebUI root — fast to respond - "NginxProxyManager:http://localhost:81" # NPM admin interface -) - -# master.conf -CURL_TIMEOUT=5 # seconds before non-response counts as a failure -RESP_FAIL_LIMIT=2 # consecutive failures before restart -``` - ---- - -### ── Required Containers ─────────────────────────────────────────────────────── - -```bash -# host1.conf -# Containers that must always be running. -# Found stopped → watchdog attempts restart every cycle until running or skip-listed. -# Uses the STRIKE SYSTEM — one miss might be mid-restart. -# Persistent failure → skip list → critical notification. -# -# These are the containers whose absence breaks everything else: -# -HOST1_WATCHDOG_REQUIRED_CONTAINERS=( - "NginxProxyManager" # reverse proxy — all external traffic routes through this - "Authelia" # SSO authentication — all protected services need it - "Mariadb-Authelia" # Authelia database — must be up before Authelia starts - "Redis-Authelia" # Authelia session cache — same startup dependency -) -``` - ---- - -### ── Dependency Ordering ─────────────────────────────────────────────────────── - -```bash -# host1.conf -# Format: "DependentContainer:dependency1 dependency2" -# Multiple dependencies space-separated. All must be running before dependent restarts. -# -# When a container and its dependency are both down: -# → restart the dependency first -# → skip the dependent this cycle -# → next cycle: dependency healthy → dependent restarts cleanly -# -# Without this: Authelia starts, can't connect to MariaDB (still starting), -# exits immediately, strike 1. Next cycle: same, strike 2. Skip list. -# MariaDB was fine the whole time. -# -HOST1_WATCHDOG_DEPENDENCIES=( - "Authelia:Mariadb-Authelia Redis-Authelia" - "Authelia-Secondary:Mariadb-Authelia Redis-Authelia-Secondary" - "NextCloud:Postgres-NextCloud" -) -``` - -The same dependency configuration is used by `docker_daily_restart.sh` and -`docker_weekly_restart.sh` — configure once, applies everywhere. - ---- - -### ── Startup Grace Period ────────────────────────────────────────────────────── - -```bash -# master.conf -# Suppress restart actions for N seconds after array start. -# Checks still run and log — only restart actions are suppressed. -# Clock starts from when the watchdog process itself starts. -# -# Without this: false-positive restarts fire in the first minutes after -# every array start while containers are still initialising. -# -WATCHDOG_STARTUP_GRACE=600 # 10 minutes -``` - ---- - -### ── Tier 2 Global Scan ──────────────────────────────────────────────────────── - -Tier 2 scans every running container when `WATCHDOG_SCAN_ALL=true`. No per-container -configuration required — it's the catch-all for everything not explicitly in Tier 1. - -```bash -# master.conf -WATCHDOG_SCAN_ALL=true # enable global scan -WATCHDOG_RESTART_UNHEALTHY=true # act on Docker HEALTHCHECK failures -WATCHDOG_NOTIFY_OOM=true # detect and notify kernel OOM kills -WATCHDOG_NOTIFY_CRASHLOOP=true # detect escalating restart counts -WATCHDOG_RESTART_DEAD=true # recover containers in dead state -WATCHDOG_RESTART_CRASHED=true # restart containers that exited non-zero -WATCHDOG_CRASH_LIMIT=5 # RestartCount above this → restart + skip list - -# Containers excluded from Tier 2 entirely. -# Use for containers you intentionally stop/start manually, or containers that -# have benign non-zero exits as part of their normal operation. -WATCHDOG_SCAN_IGNORE=( - "my-one-shot-container" # runs and exits normally — not a crash -) -``` - -Each toggle is independent — disable any check that produces false positives in your -environment without affecting the others. - ---- - -### ── Restart Loop Protection ────────────────────────────────────────────────── - -```bash -# master.conf -# N restarts within a rolling window → skip list + critical notification. -# The watchdog stops touching the container entirely. -# Skip list lives on /boot/config/ — survives reboots intentionally. -# A container bad enough to be skip-listed is still broken after a reboot. -# -WATCHDOG_CONTAINER_RESTART_LIMIT=3 # restarts in the window before skip list -WATCHDOG_CONTAINER_RESTART_WINDOW=1 # rolling window in hours -WATCHDOG_CONTAINER_RESTART_LOG="/boot/config/container_restart_history.db" -SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db" -``` - -**Auto-clear:** The watchdog checks the skip list every cycle and removes any container -it finds running. If the container recovers on its own, monitoring resumes automatically. -Manual clear is only needed when the container is stuck stopped. - ---- - -### ── Notification Batching ──────────────────────────────────────────────────── - -```bash -# master.conf -# All events from one cycle collected → one notification at end of cycle. -# -# Why: a shared database going down can cascade 10+ containers failing -# simultaneously. Without batching: 10 individual pings. With batching: -# one summary listing all affected containers. Actionable vs overwhelming. -# -WATCHDOG_BATCH_NOTIFY=true -``` - ---- - -### ── State Files Reference ──────────────────────────────────────────────────── - -| File | Config Var | Location | Resets | Purpose | -|------|-----------|----------|--------|---------| -| Strike counts | `WATCHDOG_STATE_FILE` | `/tmp/` | On reboot | Per-container CPU/HTTP strike counters | -| Skip list | `SYS_WATCHDOG_FAILED_FILE` | `/boot/config/` | Never (manual / auto-clear) | Containers that exhausted restart attempts | -| Restart history | `WATCHDOG_CONTAINER_RESTART_LOG` | `/boot/config/` | Auto-purge after window | Restart loop detection data | -| Shared state | `SYS_WATCHDOG_STATE_FILE` | `/tmp/` | On reboot | RAM emergency flag + cycle heartbeat from `system_watchdog.sh` | - -`/tmp/` resets on reboot — correct, strike counts before a reboot are meaningless after it. -`/boot/config/` survives reboots — correct, a skip-listed container is still broken after a reboot. +> **Watchdog configuration has moved.** `docker_watchdog.sh` is now in `Watchdogs/`. +> Memory limits, CPU thresholds, HTTP health checks, dependency ordering, skip list +> recovery, and all watchdog config vars are in `Watchdogs/Manual-Watchdogs.md`. --- @@ -344,18 +118,6 @@ Everything gets updated at least once per week with no explicit configuration. ```bash # Per-host — varies between HOST1 and HOST2 -# Tier 1 — memory limits ("ContainerName:LimitInMB") -HOST1_WATCHDOG_CONTAINERS=() - -# Tier 1 — HTTP health check endpoints ("ContainerName:http://host:port") -HOST1_WATCHDOG_CONTAINER_URLS=() - -# Tier 1 — must always be running -HOST1_WATCHDOG_REQUIRED_CONTAINERS=() - -# Tier 1 + 2 — dependency ordering ("Dependent:dep1 dep2") -HOST1_WATCHDOG_DEPENDENCIES=() - # Daily restart list (also drives docker_update.sh normal mode) HOST1_DAILY_RESTART_CONTAINERS=() @@ -367,6 +129,9 @@ HOST1_NETWORK_CONNECT_NETWORKS=() # Containers to connect to every configured network HOST1_NETWORK_CONNECT_CONTAINERS=() + +# Watchdog config (memory limits, HTTP checks, required, dependencies): +# → see Watchdogs/Manual-Watchdogs.md ``` --- @@ -376,47 +141,16 @@ HOST1_NETWORK_CONNECT_CONTAINERS=() ```bash # Shared — applies to both servers -# ── Watchdog timing ──────────────────────────────────────────────────────── -DOCKER_WATCHDOG_INTERVAL=900 # seconds between cycles (15 minutes) -WATCHDOG_STARTUP_GRACE=600 # seconds before restarts begin after boot -CONTAINER_DELAY=15 # seconds between dependency + dependent restart - -# ── Memory ───────────────────────────────────────────────────────────────── -SOFT_MEM_THRESHOLD=80 # warn at % of hard limit (no restart) - -# ── CPU ──────────────────────────────────────────────────────────────────── -SOFT_CPU_THRESHOLD=50 # warn threshold — normalised % -HARD_CPU_THRESHOLD=85 # strike threshold — normalised % -CPU_FAIL_LIMIT=2 # consecutive strikes before restart - -# ── HTTP health check ────────────────────────────────────────────────────── -CURL_TIMEOUT=5 # seconds before curl times out -RESP_FAIL_LIMIT=2 # consecutive failures before restart - -# ── Restart loop protection ──────────────────────────────────────────────── -WATCHDOG_CONTAINER_RESTART_LIMIT=3 # restarts before skip list -WATCHDOG_CONTAINER_RESTART_WINDOW=1 # rolling window in hours - -# ── Tier 2 global scan ───────────────────────────────────────────────────── -WATCHDOG_SCAN_ALL=true -WATCHDOG_SCAN_IGNORE=() -WATCHDOG_RESTART_UNHEALTHY=true -WATCHDOG_NOTIFY_OOM=true -WATCHDOG_NOTIFY_CRASHLOOP=true -WATCHDOG_CRASH_LIMIT=5 -WATCHDOG_RESTART_DEAD=true -WATCHDOG_RESTART_CRASHED=true - -# ── Notifications ────────────────────────────────────────────────────────── -WATCHDOG_BATCH_NOTIFY=true - -# ── Container updates ────────────────────────────────────────────────────── +# ── Container updates ────────────────────────────────────────────────── DAILY_CONTAINER_UPDATES=true WEEKLY_REMAINING_UPDATES=true -# ── Retry behaviour (shared by restart scripts) ──────────────────────────── +# ── Retry behaviour (shared by restart scripts) ──────────────────────── RETRY_COUNT=3 # retry attempts before marking failed SLEEP=5 # seconds between retry attempts + +# Watchdog thresholds (CPU, memory, HTTP, restart loop): +# → see Watchdogs/Manual-Watchdogs.md ``` --- @@ -425,47 +159,22 @@ SLEEP=5 # seconds between retry attempts ### ── Adding a Container to Monitoring ──────────────────────────────────────── -Adding a container is purely additive — add the relevant lines to `host*.conf`. +Adding a container to watchdog monitoring is purely additive — add lines to `host*.conf`. No script changes. `detect_hosts()` picks up the new config on the next watchdog cycle. +See `Watchdogs/Manual-Watchdogs.md` for the full procedure. + +To add a container to **daily restarts** (and daily image updates): ```bash -# host1.conf — example: adding "MyApp" to full Tier 1 + daily restarts - -# 1. Memory hard limit — size at ~150-200% of normal peak (check: docker stats MyApp) -HOST1_WATCHDOG_CONTAINERS=( - ...existing... - "MyApp:2048" # 2GB ceiling -) - -# 2. HTTP health check -HOST1_WATCHDOG_CONTAINER_URLS=( - ...existing... - "MyApp:http://localhost:8080/health" # or root URL if no /health endpoint -) - -# 3. Required — add only if absence breaks other services -HOST1_WATCHDOG_REQUIRED_CONTAINERS=( - ...existing... - "MyApp" -) - -# 4. Dependency — add if MyApp needs another container up first -HOST1_WATCHDOG_DEPENDENCIES=( - ...existing... - "MyApp:MyApp-Database" -) - -# 5. Daily restart — add if MyApp degrades over time +# host1.conf HOST1_DAILY_RESTART_CONTAINERS=( ...existing... "MyApp" # also adds it to the daily image update ) ``` -Tier 2 picks up MyApp automatically — no configuration needed. It will be included in -the global health scan from the next cycle onward. +To **exclude** a container from Tier 2 global scan (e.g. one-shot that exits normally): -To **exclude** MyApp from Tier 2 (e.g. it's a one-shot container that exits normally): ```bash # master.conf WATCHDOG_SCAN_IGNORE=( @@ -477,45 +186,13 @@ WATCHDOG_SCAN_IGNORE=( ### ── Skip List Recovery ──────────────────────────────────────────────────────── -Used when `docker_watchdog.sh` has skip-listed a container after exhausting restart -attempts. The watchdog stops touching it and sends a critical notification. Human -intervention required. +When `docker_watchdog.sh` skip-lists a container, use `Tools/watchdog_skip_list_manager.sh`. +Full procedure in `Watchdogs/Manual-Watchdogs.md`. ```bash -# Step 1 — understand the situation -# Shows skip list contents, container states, restart history -watchdog_skip_list_manager.sh --status - -# Step 2 — fix the underlying problem first -# Check logs: docker logs ContainerName --tail 100 -# Check disk: df -h /mnt/user -# Check database: docker exec ContainerName sqlite3 /path/to.db ".tables" -# Fix before clearing — clearing without fixing just resets the counter - -# Step 3 — clear the container from the skip list + restart history -# Clearing history is important: the counter carries over otherwise and -# the container hits the limit again almost immediately on any startup trouble -watchdog_skip_list_manager.sh --clear ContainerName - -# Step 4 — start the container manually -# Confirms your fix worked before handing back to the watchdog -docker start ContainerName - -# Step 5 — monitoring resumes automatically -# Next watchdog cycle: container seen running → removed from skip list -# Restart history clean. Back to normal. -``` - -> ⚠️ **If `docker_watchdog.sh` is currently running when you clear the skip list**, it -> may re-add the container on its very next cycle if the container is still in a bad -> state. The script detects this and warns you. Fix the root cause **before** clearing. - -Other skip list actions: -```bash -watchdog_skip_list_manager.sh # status (default) -watchdog_skip_list_manager.sh --clear-all # clear everything -watchdog_skip_list_manager.sh --clear-all --force # non-interactive -watchdog_skip_list_manager.sh --clear-all --dry-run # preview what would clear +Tools/watchdog_skip_list_manager.sh --status # see skip list + container states +Tools/watchdog_skip_list_manager.sh --clear MyApp # clear after fixing root cause +Tools/watchdog_skip_list_manager.sh --clear-all # clear everything ``` --- @@ -530,25 +207,9 @@ All scripts support these standard flags: | `--status` | Show current config, container states, and relevant runtime info, then exit. | | `--log` | Verbose mode — adds per-container banners, action lines, pull output, and list details. | -**Output tiers:** Scripts have two output levels. Without `--log`, each script processes -silently and always concludes with a summary block: identity, duration, counts, and -a status line. Per-container detail — individual container names, pull Status lines, -skip reasons — only appears with `--log`. Warnings and errors are always visible -regardless of `--log`. - -`docker_watchdog.sh` is the exception — it runs 96 cycles/day as a background daemon -and is silent on clean cycles by design. Its summary block only fires when there are -restarts or warnings. Use `--log` to see per-cycle detail on clean cycles. - -### `docker_watchdog.sh --status` shows: -Strike counts for all monitored containers, current skip list contents, whether grace -period is active and how long remains, whether RAM emergency deferral is active, last -cycle timing. - -### `docker_watchdog.sh --dry-run` shows: -A full watchdog cycle without restarting anything. Shows what the watchdog would do -based on current container states. Useful for verifying configuration before enabling -automatic restarts. +**Output tiers:** Without `--log`, each script processes silently and concludes with a +summary block: identity, duration, counts, and a status line. Per-container detail only +appears with `--log`. Warnings and errors are always visible regardless of `--log`. ### `docker_update.sh --remainder` Switches to remainder mode — updates all running containers not in the managed daily/weekly diff --git a/Docker_Essentials/README-Docker_Essentials.md b/Docker_Essentials/README-Docker_Essentials.md index 2abee83..ec5254f 100644 --- a/Docker_Essentials/README-Docker_Essentials.md +++ b/Docker_Essentials/README-Docker_Essentials.md @@ -154,24 +154,16 @@ its history after you've fixed the problem. No manual file editing required. ``` Docker_Essentials/ ← acts on containers (this folder) unRAID_Essentials/ ← acts on the server itself +Watchdogs/ ← reactive monitoring + last-resort stability Monitors/ ← observes, measures, reports Rsync/ ← moves data between servers ``` -Five distinct responsibilities, each handled by dedicated scripts: +> `docker_watchdog.sh` has moved to `Watchdogs/`. Container healing, memory limits, +> HTTP health checks, and skip list management are documented in +> `Watchdogs/README-Watchdogs.md` and `Watchdogs/Manual-Watchdogs.md`. ---- - -### 🔁 Reactive Healing — `docker_watchdog.sh` - -Continuous two-tier monitoring that catches problems as they happen and acts on them -immediately. Runs as a background process started at array start. Every 15 minutes it -checks the full stack and fixes what it can — silently when everything is fine, visibly -when something needs attention. - -Two tiers because different containers need different monitoring strategies: -- **Tier 1** — explicit per-container configuration with specific thresholds -- **Tier 2** — global scan of everything that's running with catch-all health checks +Four distinct responsibilities in this folder, each handled by dedicated scripts: --- @@ -212,34 +204,22 @@ or in-progress downloads. --- -## ━━━ RELATIONSHIP TO SYSTEM WATCHDOG ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ +## ━━━ RELATIONSHIP TO WATCHDOGS ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ -Two watchdogs run simultaneously. They are designed to work together, not compete: +`docker_watchdog.sh` has moved to `Watchdogs/` and is now one of four coordinated +single-pass scripts called every minute by `Orchestrators/watchdog_orchestrator.sh`. ``` -system_watchdog.sh ← watches the server: RAM, CPU, disk, kernel, daemon health -docker_watchdog.sh ← watches the containers: memory, CPU, HTTP response, crashes +Watchdogs/resource_watchdog.sh ← reduces pressure before healing attempts +Watchdogs/docker_watchdog.sh ← heals containers (reads resource_watchdog state) +Watchdogs/storage_watchdog.sh ← pool growth + runaway log detection +Watchdogs/system_watchdog.sh ← last resort — reboots when healing has failed ``` -**The coordination problem:** During a RAM emergency, `system_watchdog.sh` stops -non-essential containers to recover free memory. Without coordination, `docker_watchdog.sh` -would see stopped containers on its next cycle and restart them — directly undoing the -RAM recovery. The two watchdogs would fight indefinitely. RAM would never recover. -The system would eventually hit the reboot threshold anyway, having accomplished nothing. - -**The solution:** A shared state file at `SYS_WATCHDOG_STATE_FILE`. When -`system_watchdog.sh` triggers a RAM emergency shutdown it writes -`mem_shutdown_active=true`. `docker_watchdog.sh` reads this flag at the start of every -cycle and defers all container restart logic until it clears. Health URL checks for -excluded containers (DNS, auth, Emby, Dispatcharr — the ones that stayed running) still -run. Everything else stands down. - -**The stale state guard:** `system_watchdog.sh` writes `watchdog_cycle=N` to the state -file on every cycle — this keeps the file's modification time current. `docker_watchdog.sh` -checks how long ago the state file was modified. If it's more than 2 hours old while -`mem_shutdown_active=true` is set, `system_watchdog.sh` has likely stopped running. -`docker_watchdog.sh` logs a warning and resumes normal operation — it won't be silenced -indefinitely by a stale flag from a process that's no longer running. +Scripts in this folder (daily restart, updates, network connect) are unaffected — +they run on their own schedules via the maintenance orchestrators and are not part +of the every-minute watchdog cycle. See `Watchdogs/README-Watchdogs.md` for the +full coordination model between all four watchdogs. --- @@ -247,7 +227,6 @@ indefinitely by a stale flag from a process that's no longer running. | Script | Role | When It Runs | |--------|------|-------------| -| `docker_watchdog.sh` | Two-tier self-healing container monitor | Continuous background loop — array start | | `docker_daily_restart.sh` | Nightly proactive restart of degradation-prone containers | 1am via `daily_sync_maintenance.sh` | | `docker_weekly_restart.sh` | Weekly restart of less-critical services | 2:30am Sunday via `weekly_sync_maintenance.sh` | | `docker_update.sh` | Container image updates — daily list + weekly remainder mode | Daily before restart; weekly end of window | @@ -265,18 +244,12 @@ indefinitely by a stale flag from a process that's no longer running. ``` Array starts │ - ├── docker_network_connect.sh ────── run once at start - │ ensure networks + connections exist - │ silent if correct, notify if creating - │ - └── docker_watchdog.sh ──────────── continuous background loop (every 15min) - Tier 1: memory, CPU, HTTP, required containers - Tier 2: global unhealthy / OOM / crash / dead scan - reads system_watchdog state (RAM emergency deferral) - │ - │ (on skip list event → operator uses) - └── Tools/watchdog_skip_list_manager.sh - inspect state, clear after fixing root cause + └── docker_network_connect.sh ────── run once at start + ensure networks + connections exist + silent if correct, notify if creating + +Every minute (Orchestrators/watchdog_orchestrator.sh): + └── Watchdogs/docker_watchdog.sh ── see Watchdogs/README-Watchdogs.md Daily maintenance window (1am): daily_sync_maintenance.sh diff --git a/README.md b/README.md index f32d38e..829393c 100644 --- a/README.md +++ b/README.md @@ -448,8 +448,8 @@ Failover coverage: serves Gmer4Lfe.us via DDNS Monitoring: - Both servers system_watchdog.sh system_watchdog.sh - docker_watchdog.sh docker_watchdog.sh + Both servers resource/docker/storage/ resource/docker/storage/ + system_watchdog.sh system_watchdog.sh failover.sh failover.sh Sunday morning coffee report Sunday morning coffee report ``` @@ -476,10 +476,13 @@ Unraid_Scripts/ ├── Failover/ ← Mutual automatic failover — continuous background process │ README: README-Failover.md │ -├── Docker_Essentials/ ← Container lifecycle: watchdog, restarts, networks +├── Watchdogs/ ← All watchdog scripts: resource, docker, storage, system +│ README: README-Watchdogs.md +│ +├── Docker_Essentials/ ← Container lifecycle: restarts, updates, networks │ README: README-Docker_Essentials.md │ -├── unRAID_Essentials/ ← Server-level: system watchdog, WebGUI, log hygiene, tuning +├── unRAID_Essentials/ ← Server-level: WebGUI, log hygiene, kernel tuning │ README: README-Unraid_Essentials.md │ ├── Media/ ← Library health + behavior-driven discovery: permissions, junk cleanup, orphan removal, weekly arr adds @@ -539,10 +542,15 @@ At Startup of Array: → php_fpm_max_children.sh WebGUI tuning before first request → ramdisk_setup.sh create ramdisk before Emby starts → docker_network_connect.sh connect containers to extra networks - → system_watchdog.sh continuous — last resort server stability - → docker_watchdog.sh continuous — two-tier container healing → failover.sh continuous — mutual failover state machine +Every minute: + watchdog_orchestrator.sh fires each watchdog in sequence + → resource_watchdog.sh reduce pressure before healing attempts + → docker_watchdog.sh two-tier container healing + → storage_watchdog.sh pool growth + runaway log detection + → system_watchdog.sh last resort — reboots when all else fails + Every 3 minutes: transcode_management.sh cleanup → manager (order non-negotiable) @@ -582,9 +590,13 @@ Sunday morning block (6–11am): What actually happens on a typical day, from the ecosystem's perspective: ``` -Throughout the day: +Throughout the day (every minute via watchdog_orchestrator.sh): + resource_watchdog.sh managing: system pressure (throttle/pause/stop) + docker_watchdog.sh healing: memory leaks, HTTP failures, required containers + storage_watchdog.sh watching: appdata growth rate, runaway log files system_watchdog.sh watching: RAM, CPU temp, rootfs, kernel, daemon - docker_watchdog.sh watching: memory, CPU, HTTP health, required containers + +Throughout the day: failover.sh watching: remote server, internet connectivity transcode_management.sh managing: ramdisk ↔ SSD, session cleanup (every 3min) critical_sync_maintenance.sh keeping: auth stack + Emby current (every 15min) diff --git a/Watchdogs/Manual-Watchdogs.md b/Watchdogs/Manual-Watchdogs.md new file mode 100644 index 0000000..9ea818f --- /dev/null +++ b/Watchdogs/Manual-Watchdogs.md @@ -0,0 +1,641 @@ +# ━━━━━ WATCHDOGS — Manual ━━━━━ + +Configuration reference, operational procedures, and troubleshooting for all four +watchdog scripts. For design philosophy and script relationships see `README-Watchdogs.md`. +For the orchestrator that calls these scripts see `Orchestrators/watchdog_orchestrator.sh`. + +--- + +## ━━━ CONTENTS ━━━ + +- [Output Tiers](#output-tiers) +- [resource_watchdog.sh](#resource_watchdogsh) +- [docker_watchdog.sh](#docker_watchdogsh) +- [storage_watchdog.sh](#storage_watchdogsh) +- [system_watchdog.sh](#system_watchdogsh) +- [Full Configuration Reference](#full-configuration-reference) +- [Troubleshooting](#troubleshooting) + +--- + +## Output Tiers + +All watchdog scripts use a two-tier output model: `echo` lines are always visible; +`log` lines only appear when `--log` is passed. + +All four watchdogs are **single-pass scripts** called once per minute by the orchestrator. +Without `--log`, only state transitions, warnings, errors, and the conclusion line are +visible. Per-check detail is suppressed on clean cycles. + +`docker_watchdog.sh` is the exception — it is **silent on clean cycles by design**. +96 cycles/day means clean-cycle noise would bury real events. Its output only appears +when there are restarts, skip-list events, or RAM deferral. Use `--log` to see +per-cycle detail on clean cycles. + +--- + +## resource_watchdog.sh + +Runs first in the orchestrator sequence. Reduces system pressure before docker_watchdog +attempts any container restarts. Containers restarted into a RAM-pressured system just +fail again — this script ensures docker_watchdog has breathing room. + +### Pressure Levels + +Three escalating levels, each additive: + +```bash +# master.conf +RW_RAM_SOFT_GB=20 # Level 1 trigger — throttle downloaders +RW_RAM_MEDIUM_GB=15 # Level 2 trigger — throttle + pause containers +RW_RAM_HARD_GB=10 # Level 3 trigger — stop containers + defer docker_watchdog +RW_RAM_RECOVER_GB=25 # de-escalate only after RAM reaches this + +RW_LOAD_SOFT_MULTIPLIER=2.0 # load > 2× cpu count = level 1 +RW_LOAD_MEDIUM_MULTIPLIER=3.0 # load > 3× cpu count = level 2 + +RW_RECOVER_CYCLES=3 # consecutive under-threshold runs before de-escalating +``` + +**Level 1 (soft):** Throttle SABnzbd + qBittorrent download speeds. +**Level 2 (medium):** Further throttle + `docker pause` non-critical containers. +**Level 3 (hard):** `docker stop` optional services + write `mem_shutdown_active=true` +to `RW_STATE_FILE`. docker_watchdog.sh reads this flag and skips all restart logic +until pressure clears. Without this coordination, docker_watchdog would immediately +restart containers that resource_watchdog just stopped to free RAM. + +Recovery de-escalates one level at a time — prevents flip-flopping between states. + +### Per-Host Container Lists + +```bash +# host1.conf +HOST1_RW_PAUSE_CONTAINERS=("Tdarr" "HandBrake") # paused at level 2 +HOST1_RW_STOP_CONTAINERS=("LocalAI" "Satisfactory") # stopped at level 3 +``` + +Containers in `RW_CRITICAL_CONTAINERS` are never paused or stopped regardless of +pressure level. Default: Emby, NginxProxyManager, Authelia, Mariadb, Redis. + +### Downloader Throttle Config + +```bash +# master.conf +RW_SABNZBD_ENABLED=true +RW_SABNZBD_SPEED_SOFT="50M" # throttled at level 1 +RW_SABNZBD_SPEED_MEDIUM="10M" # throttled further at level 2 + +RW_QBIT_ENABLED=true +RW_QBIT_DL_SOFT=51200 # KB/s — level 1 +RW_QBIT_DL_MEDIUM=10240 # KB/s — level 2 + +# host1.conf (API access) +HOST1_SABNZBD_URL="http://localhost:8080" +HOST1_SABNZBD_API_KEY="your-api-key" +HOST1_QBIT_URL="http://localhost:8090" +HOST1_QBIT_USERNAME="admin" +HOST1_QBIT_PASSWORD="your-password" +``` + +### Usage + +```bash +resource_watchdog.sh # single pass (called by watchdog_orchestrator.sh) +resource_watchdog.sh --dry-run # show what would be throttled/paused/stopped +resource_watchdog.sh --status # current level, active actions, recovery cycle count +resource_watchdog.sh --log # verbose per-check output +``` + +--- + +## docker_watchdog.sh + +Runs second in the orchestrator sequence. Two-tier container healing — explicit +per-container configuration (Tier 1) plus a global catch-all scan (Tier 2). + +Reads `RW_STATE_FILE` at cycle start — if `mem_shutdown_active=true`, skips all +container restart logic (resource_watchdog is managing the situation). Health URL +checks for excluded containers still run. + +### Memory Hard Limits + +```bash +# host1.conf +# Format: "ContainerName:LimitInMB" +# Immediate restart when exceeded — no strike system. Memory leaks are not spikes. +# +# Sizing: check normal peak with "docker stats ContainerName" +# Set limit at ~150-200% of normal peak +# +HOST1_WATCHDOG_CONTAINERS=( + "Emby:18432" # 18GB — peaks ~12GB under heavy transcode load + "LidaTube:6144" # 6GB — YouTube downloader, grows with large queues + "Tdarr:6144" # 6GB — video transcoder, memory-intensive by nature + "Code-Server:1024" # 1GB — IDE, should be light; 1GB is generous +) +``` + +A soft warning fires at `SOFT_MEM_THRESHOLD=80` percent of the hard limit — early +visibility into a container approaching its ceiling before a restart is triggered. + +### CPU Thresholds + +```bash +# master.conf +# CPU is normalised against total core count. +# 85% normalised on a 16-core machine = 13.6 cores worth of a single process. +# +# CPU uses a STRIKE SYSTEM — brief spikes are normal (Tdarr, Emby transcoding, SABnzbd). +# Strike 1: above HARD_CPU_THRESHOLD → warn, increment strike +# Strike 2: above threshold → restart, reset counter +# Recovery: drops below threshold any cycle → reset to 0 +# +SOFT_CPU_THRESHOLD=50 # warn at 50% normalised — informational only +HARD_CPU_THRESHOLD=85 # strike at 85% normalised +CPU_FAIL_LIMIT=2 # consecutive strikes before restart +``` + +### HTTP Health Checks + +```bash +# host1.conf +# Format: "ContainerName:http://host:port/optional-path" +# Two consecutive non-responses trigger a restart. +# "Container running" and "service responding" are not the same thing. +# +HOST1_WATCHDOG_CONTAINER_URLS=( + "Emby:http://localhost:8096" # Emby WebUI root + "NginxProxyManager:http://localhost:81" # NPM admin interface +) + +# master.conf +CURL_TIMEOUT=5 # seconds before non-response counts as a failure +RESP_FAIL_LIMIT=2 # consecutive failures before restart +``` + +### Required Containers + +```bash +# host1.conf +# Found stopped → restart attempted every cycle until running or skip-listed. +# +HOST1_WATCHDOG_REQUIRED_CONTAINERS=( + "NginxProxyManager" + "Authelia" + "Mariadb-Authelia" + "Redis-Authelia" +) +``` + +### Dependency Ordering + +```bash +# host1.conf +# Format: "DependentContainer:dependency1 dependency2" +# All dependencies must be running before the dependent is restarted. +# Prevents Authelia crash-looping while MariaDB is still starting. +# +HOST1_WATCHDOG_DEPENDENCIES=( + "Authelia:Mariadb-Authelia Redis-Authelia" + "NextCloud:Postgres-NextCloud" +) +``` + +Same dependency config is used by `docker_daily_restart.sh` and +`docker_weekly_restart.sh` — configure once, applies everywhere. + +### Startup Grace Period + +```bash +# master.conf +# Suppress restart actions for N seconds after array start. +# Checks still run and log — only restart actions are suppressed. +# +WATCHDOG_STARTUP_GRACE=600 # 10 minutes +``` + +### Tier 2 Global Scan + +```bash +# master.conf +WATCHDOG_SCAN_ALL=true # enable global scan +WATCHDOG_RESTART_UNHEALTHY=true # act on Docker HEALTHCHECK failures +WATCHDOG_NOTIFY_OOM=true # detect and notify kernel OOM kills +WATCHDOG_NOTIFY_CRASHLOOP=true # detect escalating restart counts +WATCHDOG_RESTART_DEAD=true # recover containers in dead state +WATCHDOG_RESTART_CRASHED=true # restart containers that exited non-zero +WATCHDOG_CRASH_LIMIT=5 # RestartCount above this → restart + skip list + +# Exclude containers from Tier 2 (one-shots, manually managed, benign exits): +WATCHDOG_SCAN_IGNORE=( + "my-one-shot-container" +) +``` + +### Restart Loop Protection + +```bash +# master.conf +WATCHDOG_CONTAINER_RESTART_LIMIT=3 # restarts in the window before skip list +WATCHDOG_CONTAINER_RESTART_WINDOW=1 # rolling window in hours +``` + +After hitting the limit: skip list + critical notification. The watchdog stops +touching the container. Auto-clear: if the container recovers on its own and is +found running, it's removed from the skip list automatically. Manual clear is only +needed when the container is stuck stopped — use `Tools/watchdog_skip_list_manager.sh`. + +### Notification Batching + +```bash +# master.conf +WATCHDOG_BATCH_NOTIFY=true +# All events from one cycle → one notification at the end. +# A shared DB going down can cascade 10+ containers. Without batching: 10 pings. +# With batching: one summary listing all affected containers. +``` + +### Usage + +```bash +docker_watchdog.sh # single pass (called by watchdog_orchestrator.sh) +docker_watchdog.sh --dry-run # full cycle preview without restarting anything +docker_watchdog.sh --status # skip list, strike counts, grace period, RAM deferral state +docker_watchdog.sh --log # verbose per-cycle output +``` + +### Skip List Recovery + +```bash +# Step 1 — understand the situation +Tools/watchdog_skip_list_manager.sh --status + +# Step 2 — fix the underlying problem +# docker logs ContainerName --tail 100 +# df -h /mnt/user + +# Step 3 — clear the container +Tools/watchdog_skip_list_manager.sh --clear ContainerName + +# Step 4 — start manually (confirms fix before handing back to watchdog) +docker start ContainerName + +# Step 5 — monitoring resumes automatically on next cycle +``` + +--- + +## storage_watchdog.sh + +Runs third in the orchestrator sequence. Two independent checks per cycle: +growth rate detection (automatic, zero config) and oversize log detection. +Uses its own strike state file — independent from docker_watchdog. + +### Growth Rate Detection + +```bash +# master.conf +WATCHDOG_CHECK_APPDATA=true +WATCHDOG_APPDATA_PATHS=("/mnt/docker-unraid/appdata") +WATCHDOG_APPDATA_GROWTH_GB=2 # growth per cycle that triggers a strike +WATCHDOG_APPDATA_STRIKE_LIMIT=3 # strikes before alert +WATCHDOG_APPDATA_GROWTH_FILE="/tmp/watchdog_appdata_growth.db" # size baseline +``` + +Runs `du -sm appdata/*/` each cycle — pure inode metadata, very lightweight on NVMe. +Compares each container's current size to the baseline from the previous cycle. +Growth > `WATCHDOG_APPDATA_GROWTH_GB` per cycle increments the container's strike count. +Strike 1: warn. Strike 2: escalate. Strike 3: critical alert. +Strikes auto-clear when growth drops to zero (condition resolved). + +**Zero configuration required for new containers.** Growth rate detection covers all +containers automatically. The suppress array below is only for known-legitimate growth. + +### Growth Suppress Ceilings + +```bash +# host1.conf +# ONLY needed in specific cases — growth rate detection covers everything automatically. +# Use when a container's appdata legitimately grows fast during normal operation +# and you want to suppress false positives above a known-safe threshold. +# +declare -A HOST1_WATCHDOG_APPDATA_SIZES=( + ["Tdarr"]="25600" # 25GB — transcode cache grows during active jobs + ["7dtd"]="20480" # 20GB — game server world data, expected large +) +``` + +### Log File Detection + +```bash +# master.conf +WATCHDOG_APPDATA_LOG_MAX_GB=2 # *.log / *.log.* files above this trigger a strike +WATCHDOG_APPDATA_TRUNCATE_LOGS=false # true: truncate at strike limit; false: alert only +``` + +Scans all `*.log` and `*.log.*` files across all appdata paths. Files over the threshold +increment per-file strike counts. At strike limit: truncates in-place with `truncate -s 0` +(container keeps its file handle — space reclaimed immediately without container restart) +or sends a critical alert if truncation is disabled. + +Log strikes auto-clear when the file drops below threshold. + +### Usage + +```bash +storage_watchdog.sh # single pass (called by watchdog_orchestrator.sh) +storage_watchdog.sh --status # strikes, growth baseline age, suppress ceilings +storage_watchdog.sh --dry-run # show what would be alerted/truncated +storage_watchdog.sh --log # verbose per-container output +``` + +--- + +## system_watchdog.sh + +Runs last in the orchestrator sequence. The only script in the ecosystem authorized +to reboot. Watches the server itself — not containers, not storage. Reboots only when +healing at every other layer has failed or when the failure is non-recoverable. + +### Three-Tier Response + +**Tier 1 — CRITICAL (immediate reboot, no strikes)** +| Condition | Threshold | Why immediate | +|-----------|-----------|---------------| +| Docker daemon unresponsive | N/A | Every docker command hangs — nothing can be healed | +| rootfs usage | `SYS_WATCHDOG_ROOTFS_CRITICAL_PCT` (99%) | SSH stops; state files fail silently | +| Kernel oops/BUG in dmesg | delta > 0 | Kernel running with corrupted state | +| File descriptor exhaustion | `SYS_WATCHDOG_FD_CRITICAL_PCT` (95%) | New connections silently failing | +| /boot read-only | write test fails | Config writes silently failing | + +**Tier 2 — URGENT (bypass strikes with OOM confirmation)** +RAM below `MEM_GB` AND OOM kills this cycle ≥ `SYS_WATCHDOG_OOM_LIMIT`. +Both conditions required — RAM alone uses the standard strike system. +OOM confirms the system is dying faster than watchdogs can heal. + +**Tier 3 — STANDARD (`SYS_WATCHDOG_STRIKES` consecutive failures → reboot)** +| Check | Threshold | +|-------|-----------| +| Free RAM | `MEM_WARN_GB` → `MEM_SHUTDOWN_GB` → `MEM_GB` | +| Load average | `SYS_WATCHDOG_LOAD_MULTIPLIER` × cpu_count | +| CPU temperature | `SYS_WATCHDOG_CPU_TEMP` | +| Zombie processes | `SYS_WATCHDOG_ZOMBIES` | +| /var/log usage | `SYS_WATCHDOG_VAR_LOG_PCT` | +| /tmp usage | `SYS_WATCHDOG_TMP_PCT` | +| Array disk errors | mdstat error delta > 0 | +| NIC state | interface operstate != "up" | +| Required containers | containers in `SYS_WATCHDOG_REQUIRED_CONTAINERS` | + +### RAM Tiers + +``` +MEM_WARN_GB (10GB) → warn + notify, no action +MEM_SHUTDOWN_GB (6GB) → stop non-essential containers, wait for recovery +MEM_GB (4GB) → strike → reboot (URGENT bypass with OOM) +MEM_RECOVER_GB (30GB) → RAM must reach this before stopped containers restart +``` + +At `MEM_SHUTDOWN_GB`, all containers NOT listed in `SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED` +are stopped. Adjust in master.conf for your critical services. + +### Abort Conditions + +Prevent reboot — running them would risk data loss: + +```bash +SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # ZFS pool degraded/faulted +SYS_WATCHDOG_ABORT_ON_PARITY=true # Parity check/rebuild running +SYS_WATCHDOG_ABORT_ON_MOVER=true # Mover running +``` + +Tier 1 CRITICAL bypasses all abort conditions — an imminent crash outweighs data +safety concerns. + +### Reboot Rate Limit + +```bash +SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 # window in hours +SYS_WATCHDOG_MAX_REBOOTS=3 # max reboots within the window +``` + +If the server reboots `SYS_WATCHDOG_MAX_REBOOTS` times within the window, the watchdog +switches from rebooting to notifying only. Prevents a boot loop where the watchdog +reboots → something crashes again immediately → reboot again. + +### Usage + +```bash +system_watchdog.sh # single pass (called by watchdog_orchestrator.sh) +system_watchdog.sh --dry-run # run detection logic without rebooting +system_watchdog.sh --status # thresholds, current state, strike counts +system_watchdog.sh --log # verbose per-check output +``` + +--- + +## Full Configuration Reference + +```bash +# master.conf + +# ── Resource Watchdog ────────────────────────────────────────────────────────── +RW_ENABLED=true +RW_STATE_FILE="/tmp/resource_watchdog_state.db" + +RW_RAM_SOFT_GB=20 +RW_RAM_MEDIUM_GB=15 +RW_RAM_HARD_GB=10 +RW_RAM_RECOVER_GB=25 + +RW_LOAD_SOFT_MULTIPLIER=2.0 +RW_LOAD_MEDIUM_MULTIPLIER=3.0 +RW_RECOVER_CYCLES=3 + +RW_SABNZBD_ENABLED=true +RW_SABNZBD_SPEED_SOFT="50M" +RW_SABNZBD_SPEED_MEDIUM="10M" +RW_QBIT_ENABLED=true +RW_QBIT_DL_SOFT=51200 # KB/s +RW_QBIT_DL_MEDIUM=10240 + +RW_CRITICAL_CONTAINERS=("Emby" "NginxProxyManager" "Authelia" "Mariadb" "Redis") + +# host*.conf +HOST1_RW_PAUSE_CONTAINERS=("Tdarr" "HandBrake") +HOST1_RW_STOP_CONTAINERS=("LocalAI" "Satisfactory") +HOST1_SABNZBD_URL="http://localhost:8080" +HOST1_SABNZBD_API_KEY="your-api-key" +HOST1_QBIT_URL="http://localhost:8090" +HOST1_QBIT_USERNAME="admin" +HOST1_QBIT_PASSWORD="your-password" + +# ── Docker Watchdog ──────────────────────────────────────────────────────────── +DOCKER_WATCHDOG_INTERVAL=900 # seconds between cycles (15 minutes) +WATCHDOG_STARTUP_GRACE=600 # seconds before restarts begin after boot +CONTAINER_DELAY=15 # seconds between dependency + dependent restart + +SOFT_MEM_THRESHOLD=80 # warn at % of hard limit (no restart) + +SOFT_CPU_THRESHOLD=50 +HARD_CPU_THRESHOLD=85 +CPU_FAIL_LIMIT=2 + +CURL_TIMEOUT=5 +RESP_FAIL_LIMIT=2 + +WATCHDOG_CONTAINER_RESTART_LIMIT=3 +WATCHDOG_CONTAINER_RESTART_WINDOW=1 + +WATCHDOG_SCAN_ALL=true +WATCHDOG_SCAN_IGNORE=() +WATCHDOG_RESTART_UNHEALTHY=true +WATCHDOG_NOTIFY_OOM=true +WATCHDOG_NOTIFY_CRASHLOOP=true +WATCHDOG_CRASH_LIMIT=5 +WATCHDOG_RESTART_DEAD=true +WATCHDOG_RESTART_CRASHED=true +WATCHDOG_BATCH_NOTIFY=true + +# State files: +WATCHDOG_STATE_FILE="/tmp/watchdog_state.db" +SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db" +WATCHDOG_CONTAINER_RESTART_LOG="/boot/config/container_restart_history.db" +SYS_WATCHDOG_STATE_FILE="/tmp/sys_watchdog_state.db" + +# host*.conf +HOST1_WATCHDOG_CONTAINERS=() # "ContainerName:LimitMB" +HOST1_WATCHDOG_CONTAINER_URLS=() # "ContainerName:http://host:port" +HOST1_WATCHDOG_REQUIRED_CONTAINERS=() +HOST1_WATCHDOG_DEPENDENCIES=() # "Dependent:dep1 dep2" + +# ── Storage Watchdog ─────────────────────────────────────────────────────────── +WATCHDOG_CHECK_APPDATA=true +WATCHDOG_APPDATA_PATHS=("/mnt/docker-unraid/appdata") +WATCHDOG_APPDATA_GROWTH_GB=2 +WATCHDOG_APPDATA_LOG_MAX_GB=2 +WATCHDOG_APPDATA_TRUNCATE_LOGS=false +WATCHDOG_APPDATA_STRIKE_LIMIT=3 +WATCHDOG_APPDATA_GROWTH_FILE="/tmp/watchdog_appdata_growth.db" +STORAGE_WATCHDOG_STATE_FILE="/tmp/storage_watchdog_state.db" + +# host*.conf (optional — only for suppress ceilings) +# declare -A HOST1_WATCHDOG_APPDATA_SIZES=( +# ["Tdarr"]="25600" +# ) + +# ── System Watchdog ──────────────────────────────────────────────────────────── +SYS_WATCHDOG_STRIKE_LIMIT=2 +SYSTEM_WATCHDOG_INTERVAL=300 + +SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 +SYS_WATCHDOG_MAX_REBOOTS=3 +SYS_WATCHDOG_OOM_LIMIT=3 + +MEM_WARN_GB=10 +MEM_SHUTDOWN_GB=6 +MEM_GB=4 +MEM_RECOVER_GB=30 + +SYS_WATCHDOG_ROOTFS_CRITICAL_PCT=99 +SYS_WATCHDOG_FD_CRITICAL_PCT=95 +SYS_WATCHDOG_LOAD_MULTIPLIER=4 +SYS_WATCHDOG_CPU_TEMP=85 +SYS_WATCHDOG_ZOMBIES=20 +SYS_WATCHDOG_VAR_LOG_PCT=80 +SYS_WATCHDOG_TMP_PCT=85 + +SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true +SYS_WATCHDOG_ABORT_ON_PARITY=true +SYS_WATCHDOG_ABORT_ON_MOVER=true + +SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED=( + "NginxProxyManager" "Authelia" "Mariadb" "Redis" "Emby" "Dispatcharr" +) +SYS_WATCHDOG_REQUIRED_CONTAINERS=() + +# State files: +SYS_WATCHDOG_STATE_FILE="/tmp/sys_watchdog_state.db" +SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db" +SYS_WATCHDOG_REBOOT_LOG="/boot/config/system_watchdog_reboots.db" +SYS_WATCHDOG_OOM_FILE="/tmp/system_watchdog_oom.db" +``` + +--- + +## Troubleshooting + +### resource_watchdog Paused Containers It Shouldn't + +```bash +# Check current state: +resource_watchdog.sh --status + +# Add to RW_CRITICAL_CONTAINERS in master.conf: +RW_CRITICAL_CONTAINERS=("Emby" "NginxProxyManager" "Authelia" "Mariadb" "Redis" "MyContainer") + +# Un-pause manually if needed: +docker unpause MyContainer +``` + +### docker_watchdog Keeps Restarting a Healthy Container + +```bash +# Check what's triggering it — memory, CPU, HTTP, or required: +docker_watchdog.sh --status + +# Check CPU normalised — brief spikes should not trigger (2-strike system): +# If triggering on CPU: check if HARD_CPU_THRESHOLD is set appropriately +# for containers with legitimate burst usage (Tdarr encoding, SABnzbd unpacking) + +# Check HTTP — is the health endpoint returning 200? +curl -sf --max-time 5 http://localhost:PORT && echo "OK" || echo "FAIL" +``` + +### Container on the Skip List After Fixing the Problem + +```bash +# See skip list and container state: +Tools/watchdog_skip_list_manager.sh --status + +# Fix root cause first, then clear: +Tools/watchdog_skip_list_manager.sh --clear ContainerName + +# Start manually to confirm fix before handing back to watchdog: +docker start ContainerName +``` + +### storage_watchdog Alerting on a Container That Grows Legitimately + +```bash +# Check what triggered it: +storage_watchdog.sh --status + +# Add a suppress ceiling to host*.conf: +# declare -A HOST1_WATCHDOG_APPDATA_SIZES=( +# ["ContainerName"]="10240" # 10GB ceiling — legitimate growth, suppress below this +# ) +``` + +### system_watchdog Rebooted Unexpectedly + +```bash +# Check the reboot log (survives reboots): +cat /boot/config/system_watchdog_reboots.db +# Shows timestamp and reason for each watchdog-triggered reboot + +# Check syslog near the reboot time: +grep "system_watchdog" /var/log/syslog | tail -20 +``` + +### system_watchdog Not Responding / Watchdog Orchestrator Reports Timeout + +```bash +# All four watchdogs run as single-pass scripts — there is no background process to check. +# If the orchestrator reports a timeout, one pass took longer than expected. + +# Check the orchestrator itself: +Orchestrators/watchdog_orchestrator.sh --status + +# Run the slow watchdog directly with --log to see where it's hanging: +Watchdogs/system_watchdog.sh --log --dry-run +``` diff --git a/Watchdogs/README-Watchdogs.md b/Watchdogs/README-Watchdogs.md new file mode 100644 index 0000000..ee0bc29 --- /dev/null +++ b/Watchdogs/README-Watchdogs.md @@ -0,0 +1,176 @@ +# ━━━━━ WATCHDOGS ━━━━━ + +**Four single-pass scripts that run every minute through `watchdog_orchestrator.sh`, +each with a clear lane:** reduce system pressure → heal containers → protect storage → +reboot if nothing else worked. They never run standalone loops. The orchestrator calls +them in order, once per cron cycle. + +--- + +## ━━━ THE PROBLEM THAT BUILT THIS ━━━ + +**Container Memory Leaks Going Undetected for Days** +Emby's transcode session handling occasionally leaks memory. SABnzbd's Python process +expands slowly across downloads. Neither crashes dramatically — they just consume more +RAM until the system starts swapping. Docker reports both containers as `Up 14 days`. +Nothing alerts. By the time someone notices, the system has been degraded for hours. + +Fix: `docker_watchdog.sh` — hard per-container memory ceilings. When a container +exceeds its limit the watchdog restarts it immediately. No strikes, no waiting. +A memory leak is not a transient spike. + +**Containers That Look Running But Aren't Responding** +Docker reports a container as `Up` while its application layer has been frozen for +hours. The reverse proxy forwards traffic to a service that returns nothing. Users see +a broken page. Docker sees a healthy container. + +Fix: `docker_watchdog.sh` — HTTP health checks on the actual service port every cycle. +Two consecutive non-responses trigger a restart. Process running and service responding +are not the same thing. + +**System Pressure Causing Docker Watchdog to Undo Itself** +Resource pressure builds. RAM drops. docker_watchdog.sh tries to restart a container +into a system that's already swapping — the restarted container fails immediately +and goes on the skip list. The real problem (RAM pressure) is never addressed. + +Fix: `resource_watchdog.sh` runs first in the orchestrator sequence. At Level 1 it +throttles downloaders. At Level 2 it pauses non-critical containers. At Level 3 it +stops heavy services and signals docker_watchdog to defer all restarts. By the time +docker_watchdog runs, the system has breathing room to actually heal. + +**Runaway Log Files Filling a Pool Before Anyone Notices** +A game server container was offline for a year, restarted for a weekend, and wrote +130GB of logs to the appdata pool. The pool grew 13% in days. No alert fired — +nothing was watching for growth at the data level, only at the container level. + +Fix: `storage_watchdog.sh` — growth-rate scan of every container's appdata directory +every cycle. No per-container configuration required. Runaway growth gets three +cycles to be confirmed, then alerts and (optionally) truncates logs automatically. + +**Rootfs at 99% With SSH Failing Silently** +Rootfs fills. SSH stops accepting new connections. Docker can't write log files. State +files fail silently. The server is functionally dead but still technically running. +Nothing in the container layer can detect or recover from this — it requires a reboot. + +Fix: `system_watchdog.sh` — watches the server itself: RAM, CPU, disk, kernel, daemon +health. Only script in the stack authorized to reboot. Runs last in the orchestrator +sequence so container healing and pressure reduction always get a chance first. + +--- + +## ━━━ WHAT THIS FOLDER DOES ━━━ + +Four watchdogs. One purpose each. Fixed execution order via `watchdog_orchestrator.sh`. + +``` +Pressure reduction resource_watchdog.sh — throttle/pause/stop before healing fails +Container healing docker_watchdog.sh — memory, CPU, HTTP, required containers +Storage protection storage_watchdog.sh — pool growth rate + runaway log detection +Last resort system_watchdog.sh — reboot only when nothing else can recover +``` + +**The execution order is the design.** Resource pressure is reduced before docker_watchdog +attempts restarts — containers restarted into a pressure-bound system just fail again. +Storage is checked after containers are healed — no false alerts from containers that +were already being restarted. System watchdog runs last — reboot is always the last +option, not the first. + +**None of these scripts run standalone loops.** Each is a single-pass script called +once per minute by `Orchestrators/watchdog_orchestrator.sh`. The orchestrator handles +startup grace, overlap protection, heartbeat, and sequencing. + +--- + +## ━━━ RELATIONSHIP TO OTHER FOLDERS ━━━ + +``` +Orchestrators/ + watchdog_orchestrator.sh ──────────────────► resource_watchdog.sh (1st — every minute) + ──────────────────► docker_watchdog.sh (2nd) + ──────────────────► storage_watchdog.sh (3rd) + ──────────────────► system_watchdog.sh (4th — last resort) + +Tools/ + watchdog_skip_list_manager.sh ◄────────────── docker_watchdog.sh writes skip list + (operator utility — inspect + clear after fixing a crash-looping container) + +Docker_Essentials/ + All container lifecycle scripts (daily restart, updates, network) — unaffected. + docker_watchdog.sh coordinates with them via shared state, not direct calls. + +unRAID_Essentials/ + Server-level scripts (WebGUI restart, inotify tuning, log hygiene) — unaffected. + system_watchdog.sh runs in the same ecosystem but is independent of those scripts. +``` + +**watchdog_orchestrator.sh stays in Orchestrators/** — it's a job runner, not a watchdog. +**watchdog_skip_list_manager.sh stays in Tools/** — it's an operator utility, not a watchdog. + +--- + +## ━━━ SCRIPTS IN THIS FOLDER ━━━ + +| Script | Role | Called By | +|--------|------|-----------| +| `resource_watchdog.sh` | Three-level pressure reduction — throttle, pause, stop | `watchdog_orchestrator.sh` — 1st every minute | +| `docker_watchdog.sh` | Two-tier container healing — memory, CPU, HTTP, required | `watchdog_orchestrator.sh` — 2nd every minute | +| `storage_watchdog.sh` | Pool growth rate + runaway log detection and remediation | `watchdog_orchestrator.sh` — 3rd every minute | +| `system_watchdog.sh` | Last-resort server watchdog — reboots when healing has failed | `watchdog_orchestrator.sh` — 4th every minute | + +> `watchdog_orchestrator.sh` is in `Orchestrators/`. `watchdog_skip_list_manager.sh` +> is in `Tools/`. Neither is a watchdog — they sit at the edges of this system. + +--- + +## ━━━ HOW THE SCRIPTS RELATE ━━━ + +``` +Every minute — watchdog_orchestrator.sh fires: + + Step 1 — resource_watchdog.sh + │ RAM/load OK → pass through (no action) + │ Level 1 (soft): throttle SABnzbd + qBit download speeds + │ Level 2 (medium): further throttle + docker pause non-critical containers + │ Level 3 (hard): docker stop optional services + │ writes mem_shutdown_active=true → RW_STATE_FILE + │ ↓ + Step 2 — docker_watchdog.sh + │ reads RW_STATE_FILE — if mem_shutdown_active=true: skip all restarts + │ + │ Tier 1 — explicit per-container checks (configured in host*.conf): + │ memory hard limits → immediate restart (no strikes) + │ CPU high sustained → 2-strike restart + │ HTTP non-response → 2-strike restart + │ required stopped → restart (dependency ordering respected) + │ + │ Tier 2 — global scan of all running containers: + │ unhealthy / OOM / crashloop / dead / non-zero exit → restart + │ N restarts in window → skip list + critical notify → human required + │ + │ (skip list management) → Tools/watchdog_skip_list_manager.sh + │ + Step 3 — storage_watchdog.sh + │ growth rate scan: du -sm appdata/* → compare to previous cycle baseline + │ growth > WATCHDOG_APPDATA_GROWTH_GB → 3-strike warn → alert + │ log file scan: find *.log > WATCHDOG_APPDATA_LOG_MAX_GB + │ oversize log found → 3-strike warn → truncate (if enabled) or alert + │ + Step 4 — system_watchdog.sh + checks the server itself — RAM, CPU temp, rootfs, FDs, kernel, daemon + Tier 1 CRITICAL → immediate reboot (no strikes) + Tier 2 URGENT → reboot if OOM confirmed + Tier 3 STANDARD → N consecutive failures → reboot + Abort conditions → ZFS unhealthy / parity running / mover active + Rate limit → max N reboots per window before switching to notify +``` + +**State file coordination between scripts:** + +| State File | Written By | Read By | Purpose | +|-----------|-----------|---------|---------| +| `RW_STATE_FILE` | `resource_watchdog.sh` | `docker_watchdog.sh` | `mem_shutdown_active` flag — defer restarts during RAM emergency | +| `SYS_WATCHDOG_STATE_FILE` | `system_watchdog.sh` | `docker_watchdog.sh` | `watchdog_cycle` heartbeat — stale guard (2hr timeout) | +| `WATCHDOG_STATE_FILE` | `docker_watchdog.sh` | itself | CPU/HTTP strike counts per container | +| `SYS_WATCHDOG_FAILED_FILE` | `docker_watchdog.sh` | `watchdog_skip_list_manager.sh` | Container skip list | +| `STORAGE_WATCHDOG_STATE_FILE` | `storage_watchdog.sh` | itself | Growth + log strike counts | +| `WATCHDOG_APPDATA_GROWTH_FILE` | `storage_watchdog.sh` | itself | Per-container size baseline for growth rate | diff --git a/unRAID_Essentials/Manual-Unraid_Essentials.md b/unRAID_Essentials/Manual-Unraid_Essentials.md index ca0543a..0590b4a 100644 --- a/unRAID_Essentials/Manual-Unraid_Essentials.md +++ b/unRAID_Essentials/Manual-Unraid_Essentials.md @@ -4,13 +4,15 @@ Configuration reference, operational procedures, and troubleshooting for system-level scripts. Read the ARRAY_START_SCRIPTS order section before adding or reordering scripts at array start. +> **Watchdog scripts have moved.** `system_watchdog.sh` and `resource_watchdog.sh` +> now live in `Watchdogs/`. Their configuration reference and troubleshooting +> procedures are in `Watchdogs/Manual-Watchdogs.md`. + --- ## ━━━ CONTENTS ━━━ - [ARRAY_START_SCRIPTS Order](#array_start_scripts-order) -- [system_watchdog.sh](#system_watchdogsh) -- [resource_watchdog.sh](#resource_watchdogsh) - [webgui_restart.sh](#webgui_restartsh) - [inotify_tuning.sh](#inotify_tuningsh) - [php_fpm_max_children.sh](#php_fpm_max_childrensh) @@ -30,9 +32,9 @@ adding or reordering scripts at array start. All scripts use a two-tier output model: `echo` lines are always visible; `log` lines only appear when `--log` is passed. -**Daemon scripts** (`system_watchdog.sh`, `resource_watchdog.sh`, `webgui_restart.sh`): -run on every cycle. Without `--log`, only state transitions, warnings, errors, and -the clean-cycle conclusion line are visible. Per-check detail suppressed. +**Daemon scripts** (`webgui_restart.sh`): run on every cycle. Without `--log`, only +state transitions, warnings, errors, and the clean-cycle conclusion line are visible. +Per-check detail suppressed. **One-shot scripts** (`clear_logs.sh`, `docker_syslog_filter.sh`, `inotify_tuning.sh`, `mover_stop.sh`, `php_fpm_max_children.sh`, `rsync_stop.sh`, `server_reboot.sh`, @@ -59,8 +61,9 @@ ARRAY_START_SCRIPTS=( "php_fpm_max_children.sh" # 3 — before WebGUI is under load "ramdisk_setup.sh" # (from Transcodes/) before Emby starts ... - "system_watchdog.sh" # LAST or near-last — starts background loop ) +# Watchdogs are NOT in ARRAY_START_SCRIPTS — they run every minute via +# Orchestrators/watchdog_orchestrator.sh (separate cron entry). ``` Why inotify FIRST: If Code-Server starts before limits are raised, it inherits @@ -73,151 +76,6 @@ generates veth messages — these will appear in syslog if the filter isn't acti --- -## system_watchdog.sh - -### Three-Tier Response - -The watchdog categorizes every failure into one of three tiers: - -**Tier 1 — CRITICAL (immediate reboot, no strikes)** -| Condition | Threshold | Why immediate | -|-----------|-----------|---------------| -| Docker daemon unresponsive | N/A | Nothing can be healed; every docker command hangs | -| rootfs usage | SYS_WATCHDOG_ROOTFS_CRITICAL_PCT (99%) | SSH stops; state files fail silently | -| Kernel oops/BUG in dmesg | delta > 0 | Kernel running with corrupted state | -| File descriptor exhaustion | SYS_WATCHDOG_FD_CRITICAL_PCT (95%) | New connections silently failing | -| /boot read-only unexpectedly | write test fails | Config writes silently failing | - -**Tier 2 — URGENT (bypass strikes with OOM confirmation)** - -RAM below MEM_GB AND OOM kills this cycle >= SYS_WATCHDOG_OOM_LIMIT. -Both conditions required — RAM alone without OOM uses the standard strike system. -OOM confirms the system is dying faster than watchdogs can heal. - -**Tier 3 — STANDARD (SYS_WATCHDOG_STRIKES consecutive failures → reboot)** -| Check | Threshold | -|-------|-----------| -| Free RAM | MEM_WARN_GB → MEM_SHUTDOWN_GB → MEM_GB | -| Load average | SYS_WATCHDOG_LOAD_MULTIPLIER × cpu_count | -| CPU temperature | SYS_WATCHDOG_CPU_TEMP | -| Zombie processes | SYS_WATCHDOG_ZOMBIES | -| /var/log usage | SYS_WATCHDOG_VAR_LOG_PCT | -| /tmp usage | SYS_WATCHDOG_TMP_PCT | -| Array disk errors | mdstat error delta > 0 | -| NIC state | interface operstate != "up" | -| Required containers | containers in SYS_WATCHDOG_REQUIRED_CONTAINERS | - -### RAM Tiers - -``` -MEM_WARN_GB (10GB) → warn + notify, no action -MEM_SHUTDOWN_GB (6GB) → stop non-essential containers, wait for recovery -MEM_GB (4GB) → strike → reboot (URGENT bypass with OOM) -MEM_RECOVER_GB (30GB) → RAM must reach this before stopped containers restart -``` - -At MEM_SHUTDOWN_GB, all containers NOT listed in -`SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED` are stopped. Excluded containers by default: -NginxProxyManager, Authelia, Mariadb, Redis, Emby, Dispatcharr. Adjust in -master.conf for your critical services. - -### Abort Conditions - -These conditions prevent a reboot — running them would cause data loss: - -```bash -SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # ZFS pool degraded/faulted -SYS_WATCHDOG_ABORT_ON_PARITY=true # Parity check/rebuild running -SYS_WATCHDOG_ABORT_ON_MOVER=true # Mover running -``` - -CRITICAL tier bypasses all abort conditions — an imminent crash outweighs -data safety concerns. - -### Strike System - -The strike count tracks consecutive failures. A single recovery resets strikes -to 0. The reboot only fires after SYS_WATCHDOG_STRIKES consecutive failures on -the same check — transient spikes (a brief load burst, a momentary RAM dip) don't -trigger reboots. - -### Reboot Rate Limit - -```bash -SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 # window in hours -SYS_WATCHDOG_MAX_REBOOTS=3 # max reboots within the window -``` - -If the server has rebooted SYS_WATCHDOG_MAX_REBOOTS times within the window, -system_watchdog stops rebooting and notifies instead. This prevents a boot loop -where the watchdog reboots → something crashes again immediately → reboot again. - -### Usage - -```bash -system_watchdog.sh # start continuous monitoring loop -system_watchdog.sh --dry-run # run detection logic without rebooting -system_watchdog.sh --status # show thresholds, current state, strike counts -system_watchdog.sh --log # verbose per-cycle output -``` - -### Verify Running - -```bash -pgrep -a -f system_watchdog.sh -# Expected: shows PID and path -``` - ---- - -## resource_watchdog.sh - -### Pressure Levels - -```bash -# master.conf -RW_RAM_SOFT_GB=20 # Level 1 trigger — throttle downloaders -RW_RAM_MEDIUM_GB=15 # Level 2 trigger — throttle + pause containers -RW_RAM_HARD_GB=10 # Level 3 trigger — stop containers -RW_RAM_RECOVER_GB=25 # recover to this before un-stopping at level 3 - -RW_LOAD_SOFT_MULTIPLIER=2.0 # load > 2× cpu count = level 1 -RW_LOAD_MEDIUM_MULTIPLIER=3.0 # load > 3× cpu count = level 2 - -RW_RECOVER_CYCLES=3 # consecutive under-threshold runs before de-escalating -``` - -### Per-Host Container Lists - -```bash -# host1.conf -HOST1_RW_PAUSE_CONTAINERS=("Tdarr" "HandBrake") # paused at medium pressure -HOST1_RW_STOP_CONTAINERS=("LocalAI" "Satisfactory") # stopped at hard pressure -``` - -Containers in `RW_CRITICAL_CONTAINERS` are never paused or stopped regardless of -pressure level. Default includes: Emby, NginxProxyManager, Authelia, Mariadb, Redis. - -### docker_watchdog Coordination - -At level 3, resource_watchdog writes `mem_shutdown_active=true` to `RW_STATE_FILE`. -docker_watchdog.sh reads this flag each cycle and skips all container restart logic -while it is set. Without this coordination, docker_watchdog would immediately -restart containers that resource_watchdog just stopped to free RAM. - -The flag is cleared when level 3 pressure resolves and containers are restarted. - -### Usage - -```bash -resource_watchdog.sh # single pass (called by watchdog_orchestrator.sh) -resource_watchdog.sh --dry-run # show what would be throttled/paused/stopped -resource_watchdog.sh --status # current level, active actions, recovery cycle count -resource_watchdog.sh --log # verbose per-check output -``` - ---- - ## webgui_restart.sh ### Escalation Logic @@ -594,75 +452,12 @@ server_reboot.sh --reason="disk work" # include reason in notification ## Full Configuration Reference +> Watchdog configuration (`system_watchdog.sh`, `resource_watchdog.sh`, +> `docker_watchdog.sh`, `storage_watchdog.sh`) lives in `Watchdogs/Manual-Watchdogs.md`. + ```bash # master.conf -# ── System Watchdog ──────────────────────────────────────────────────────────── -SYS_WATCHDOG_INTERVAL=300 # seconds between check cycles -SYS_WATCHDOG_STRIKES=3 # consecutive failures before reboot -SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 # rate limit window -SYS_WATCHDOG_MAX_REBOOTS=3 # max reboots in window -SYS_WATCHDOG_OOM_LIMIT=3 # OOM kills/cycle for URGENT bypass - -MEM_WARN_GB=10 # warn + notify -MEM_SHUTDOWN_GB=6 # stop non-essential containers -MEM_GB=4 # strike → reboot -MEM_RECOVER_GB=30 # recovery threshold - -SYS_WATCHDOG_ROOTFS_CRITICAL_PCT=99 # Tier 1 trigger -SYS_WATCHDOG_FD_CRITICAL_PCT=95 # Tier 1 trigger -SYS_WATCHDOG_LOAD_MULTIPLIER=4 # Tier 3 — ×cpu_count -SYS_WATCHDOG_CPU_TEMP=85 # Tier 3 — Celsius -SYS_WATCHDOG_ZOMBIES=20 # Tier 3 — process count -SYS_WATCHDOG_VAR_LOG_PCT=80 # Tier 3 — percent full -SYS_WATCHDOG_TMP_PCT=85 # Tier 3 — percent full - -SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true -SYS_WATCHDOG_ABORT_ON_PARITY=true -SYS_WATCHDOG_ABORT_ON_MOVER=true - -SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED=( - "NginxProxyManager" "Authelia" "Mariadb" "Redis" "Emby" "Dispatcharr" -) -SYS_WATCHDOG_REQUIRED_CONTAINERS=() # containers that must be running - -# State files: -SYS_WATCHDOG_STATE_FILE="/tmp/sys_watchdog_state.db" -SYS_WATCHDOG_REBOOT_LOG="/tmp/sys_watchdog_reboots.db" -SYS_WATCHDOG_FAILED_FILE="/tmp/sys_watchdog_failed.db" -SYS_WATCHDOG_OOM_FILE="/tmp/sys_watchdog_oom.db" - -# ── Resource Watchdog ────────────────────────────────────────────────────────── -RW_ENABLED=true -RW_STATE_FILE="/tmp/resource_watchdog_state.db" - -RW_RAM_SOFT_GB=20 -RW_RAM_MEDIUM_GB=15 -RW_RAM_HARD_GB=10 -RW_RAM_RECOVER_GB=25 - -RW_LOAD_SOFT_MULTIPLIER=2.0 -RW_LOAD_MEDIUM_MULTIPLIER=3.0 -RW_RECOVER_CYCLES=3 - -RW_SABNZBD_ENABLED=true -RW_SABNZBD_SPEED_SOFT="50M" -RW_SABNZBD_SPEED_MEDIUM="10M" -RW_QBIT_ENABLED=true -RW_QBIT_DL_SOFT=51200 # KB/s -RW_QBIT_DL_MEDIUM=10240 - -RW_CRITICAL_CONTAINERS=("Emby" "NginxProxyManager" "Authelia" "Mariadb" "Redis") - -# ── Per-Host (host*.conf) ─────────────────────────────────────────────── -HOST1_RW_PAUSE_CONTAINERS=("Tdarr" "HandBrake") -HOST1_RW_STOP_CONTAINERS=("LocalAI" "Satisfactory") -HOST1_SABNZBD_URL="http://localhost:8080" -HOST1_SABNZBD_API_KEY="your-api-key" -HOST1_QBIT_URL="http://localhost:8090" -HOST1_QBIT_USERNAME="admin" -HOST1_QBIT_PASSWORD="your-password" - # ── WebGUI Watchdog ──────────────────────────────────────────────────────────── WEBGUI_URL="http://localhost" WEBGUI_TIMEOUT=5 @@ -699,43 +494,8 @@ REBOOT_VM_WAIT=30 ## Troubleshooting -### system_watchdog.sh Not Starting - -```bash -# Check if already running (acquire_lock prevents second instance): -pgrep -a -f system_watchdog.sh - -# Check state file permissions: -ls -la /tmp/sys_watchdog_*.db - -# Run with --log to see startup: -system_watchdog.sh --log -``` - -### system_watchdog Rebooted Unexpectedly - -```bash -# Check reboot log: -cat /tmp/sys_watchdog_reboots.db -# Shows timestamp and reason for each watchdog-triggered reboot - -# Check what condition triggered it: -# Look in /var/log/syslog for "system_watchdog" near the reboot time -grep "system_watchdog" /var/log/syslog | tail -20 -``` - -### resource_watchdog Paused Containers It Shouldn't Have - -```bash -# Check current state: -resource_watchdog.sh --status - -# Add the container to RW_CRITICAL_CONTAINERS in master.conf: -RW_CRITICAL_CONTAINERS=("Emby" "NginxProxyManager" "Authelia" "Mariadb" "Redis" "MyContainer") - -# Un-pause manually if needed: -docker unpause MyContainer -``` +> Watchdog troubleshooting (system_watchdog, resource_watchdog, docker_watchdog, +> storage_watchdog) is in `Watchdogs/Manual-Watchdogs.md`. ### rsync_stop Killed the Wrong Thing diff --git a/unRAID_Essentials/README-Unraid_Essentials.md b/unRAID_Essentials/README-Unraid_Essentials.md index f74dfbe..46614a7 100644 --- a/unRAID_Essentials/README-Unraid_Essentials.md +++ b/unRAID_Essentials/README-Unraid_Essentials.md @@ -71,8 +71,6 @@ SIGTERM (graceful — finishes current file), SIGKILL only if needed. ## ━━━ WHAT THIS FOLDER DOES ━━━ ``` -Last-resort stability system_watchdog.sh — reboots before crash -Pressure reduction resource_watchdog.sh — throttle/pause/stop under load WebGUI availability webgui_restart.sh — nginx → php-fpm → emhttp escalation Kernel tuning inotify_tuning.sh — file watch limits php_fpm_max_children.sh — PHP worker count @@ -84,6 +82,9 @@ Graceful operations mover_stop.sh — clean mover stop server_reboot.sh — clean reboot with pre-flight warnings ``` +> `system_watchdog.sh` and `resource_watchdog.sh` have moved to `Watchdogs/`. +> See `Watchdogs/README-Watchdogs.md` for the full watchdog suite. + --- ## ━━━ RELATIONSHIP TO OTHER FOLDERS ━━━ @@ -93,33 +94,22 @@ Orchestrators/ array_start.sh ─────────────────────────► inotify_tuning.sh (first in sequence) ─────────────────────────► docker_syslog_filter.sh (second) ─────────────────────────► php_fpm_max_children.sh - ─────────────────────────► system_watchdog.sh (background loop) - - watchdog_orchestrator.sh ───────────────► resource_watchdog.sh (every minute) weekly_maintenance.sh ──────────────────► clear_logs.sh server_reboot.sh ────────────────────────► user_scripts_stop.sh (called internally) -Docker_Essentials/ - docker_watchdog.sh ◄─── reads ──────────── resource_watchdog.sh state - (mem_shutdown_active flag) +Watchdogs/ + system_watchdog.sh and resource_watchdog.sh now live here. + See Watchdogs/README-Watchdogs.md for how they relate to each other + and to docker_watchdog.sh and storage_watchdog.sh. ``` -`system_watchdog.sh` and `docker_watchdog.sh` (in Docker_Essentials/) are -designed to work together — docker_watchdog heals containers first, -system_watchdog reboots only when healing has failed. `resource_watchdog.sh` -coordinates with docker_watchdog via the `mem_shutdown_active` state flag to -prevent docker_watchdog from restarting containers that resource_watchdog just -stopped to free RAM. - --- ## ━━━ SCRIPTS IN THIS FOLDER ━━━ | Script | Role | When It Runs | |--------|------|-------------| -| `system_watchdog.sh` | Three-tier last-resort stability watchdog | Continuous background loop via array_start.sh | -| `resource_watchdog.sh` | Pressure reduction — throttle/pause/stop under load | Every minute via watchdog_orchestrator.sh | | `webgui_restart.sh` | WebGUI availability — nginx → php-fpm → emhttp | Every 10 min via User Scripts | | `inotify_tuning.sh` | Raise inotify kernel limits | At array start — FIRST | | `php_fpm_max_children.sh` | Set PHP-FPM max worker count | At array start | @@ -137,20 +127,14 @@ stopped to free RAM. ``` Array starts │ - ├─ inotify_tuning.sh ← FIRST — kernel limits inherited at container launch - ├─ docker_syslog_filter.sh ← SECOND — before any veth interfaces are created - ├─ php_fpm_max_children.sh ← before WebGUI is under load - └─ system_watchdog.sh ← starts background loop - - -Every minute (watchdog_orchestrator.sh): - └─ resource_watchdog.sh - Level 1 (soft): throttle SABnzbd + qBit download speeds - Level 2 (medium): further throttle + docker pause non-essential containers - Level 3 (hard): docker stop optional services + set mem_shutdown_active=true - ↓ - docker_watchdog.sh reads mem_shutdown_active — defers restarts + ├─ inotify_tuning.sh ← FIRST — kernel limits inherited at container launch + ├─ docker_syslog_filter.sh ← SECOND — before any veth interfaces are created + └─ php_fpm_max_children.sh ← before WebGUI is under load +Every minute (watchdog_orchestrator.sh in Orchestrators/): + → Watchdogs/resource_watchdog.sh → Watchdogs/docker_watchdog.sh + → Watchdogs/storage_watchdog.sh → Watchdogs/system_watchdog.sh + (see Watchdogs/README-Watchdogs.md for full flow) Every 10 minutes (User Scripts): └─ webgui_restart.sh @@ -161,13 +145,11 @@ Every 10 minutes (User Scripts): Step 3: restart emhttp → recheck All failed → notify, exit 1 - Weekly (weekly_maintenance.sh): └─ clear_logs.sh System logs: clear if > LOG_MIN_SIZE_MB Docker logs: clear per-container if > LOG_DOCKER_MAX_MB - Manual operations: mover_stop.sh → wall → SIGTERM → SIGKILL → verify stopped rsync_stop.sh → detect orchestrator → kill rsync (or orchestrator+rsync)