From 2fa78a21b601315353b7985178ded944d90f2cf0 Mon Sep 17 00:00:00 2001 From: Gmer4Lfe Date: Fri, 22 May 2026 22:07:46 -0400 Subject: [PATCH] =?UTF-8?q?doc=20audit=20=E2=80=94=20fix=20stale=20watchdo?= =?UTF-8?q?g=20vars,=20fallback=5Fstate=5Freset,=20add=20monthly=5Fmainten?= =?UTF-8?q?ance=20and=20mesh=5Fmonitor=20sections?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Tools/Manual-Tools.md: fallback_state_reset.sh marked not yet built; manual workaround added; TOC entry updated; failover→fallback rename throughout - Watchdogs/Manual-Watchdogs.md: System Watchdog config block updated — remove pre-split vars (MEM_WARN_GB, MEM_SHUTDOWN_GB, MEM_GB, MEM_RECOVER_GB, SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED, SYS_WATCHDOG_REQUIRED_CONTAINERS); add SYS_WATCHDOG_MEM_GB (reboot trigger only); fix state file path names (watchdog_state.db → container_watchdog_state.db, sys_watchdog_state.db → system_watchdog_state.db) - Orchestrators/README-Orchestrators.md: add monthly_maintenance.sh section with two-gate model, configuration reference, and usage examples - Monitors/README-Monitors.md: add mesh_monitor.sh to both tables and diagram - Monitors/Manual-Monitors.md: add mesh_monitor.sh section - README-User_Script_Plug-in.md, Rsync/README-Rsync.md, Rsync/Manual-Rsync.md, README.md: emby-failover → emby-fallback (missed from prior batch) --- Monitors/Manual-Monitors.md | 39 +++++++++++++++++++++ Monitors/README-Monitors.md | 6 ++++ Orchestrators/README-Orchestrators.md | 50 +++++++++++++++++++++++++-- README-User_Script_Plug-in.md | 6 ++-- README.md | 2 +- Rsync/Manual-Rsync.md | 16 ++++----- Rsync/README-Rsync.md | 6 ++-- Tools/Manual-Tools.md | 40 ++++++++++----------- Watchdogs/Manual-Watchdogs.md | 18 +++------- 9 files changed, 131 insertions(+), 52 deletions(-) diff --git a/Monitors/Manual-Monitors.md b/Monitors/Manual-Monitors.md index cff275c..dc3001b 100644 --- a/Monitors/Manual-Monitors.md +++ b/Monitors/Manual-Monitors.md @@ -343,6 +343,45 @@ long-term history. It avoids flash drive writes entirely. --- +## ━━━ MESH MONITOR ━━━ + +At-a-glance view of the mesh — who the members are, what each node covers for the +others when they go down, and the current partnership state. Read-only. Parses conf +files directly — no SSH, no API calls, no sensitive data. + +### Sections + +``` +MEMBERS — server name, owner, contact email for each HOST* +COVERAGE — what each host covers for others, by tier +PARTNERSHIP — enabled/disabled, current owner, sync interval +``` + +### When to Use + +``` +Onboarding a new mesh member — paste the output to show them the current topology +After adding a new container to a fallback tier — verify it appears in the right tier +When troubleshooting fallback — confirm coverage arrays are configured correctly +``` + +### Usage + +```bash +mesh_monitor.sh # full mesh overview +mesh_monitor.sh --log # verbose output +mesh_monitor.sh --status # show config and exit +``` + +### Configuration + +No dedicated config keys. Reads `HOST*_OWNER`, `HOST*_OWNER_EMAIL`, and +`FALLBACK_HOST*_COVERS_HOST*_TIER*` arrays from `host*.conf`. Scales automatically +as new nodes join — add a new `HOST3_*` block and it appears in output without +changing the script. + +--- + ## ━━━ FULL CONFIGURATION REFERENCE ━━━ ### master.conf diff --git a/Monitors/README-Monitors.md b/Monitors/README-Monitors.md index e790180..6088e2f 100644 --- a/Monitors/README-Monitors.md +++ b/Monitors/README-Monitors.md @@ -108,6 +108,7 @@ folder is designed with this constraint in mind. | `weekly_health_digest.sh` | Never | reads existing state files only | | `bandwidth_monitor.sh` | One append + one trim per rsync run | Bounded — never exceeds `BANDWIDTH_LOG_RETENTION` lines | | `system_tuning_monitor.sh` | One append + one trim every 6 hours | Bounded — trimmed to `TUNING_LOG_RETENTION` days | +| `mesh_monitor.sh` | Never | read-only conf parse — no API calls, no SSH | --- @@ -123,6 +124,7 @@ folder is designed with this constraint in mind. | `bandwidth_monitor.sh` | rsync transfer history and trends | Auto (called by rsync.sh) + weekly report Sunday 11am | | `weekly_health_digest.sh` | Full ecosystem health aggregation | Daily 8am (DIGEST_PROFILE controls notify) | | `system_tuning_monitor.sh` | inotify + PHP-FPM utilisation trends | Every 6 hours | +| `mesh_monitor.sh` | Mesh membership — who covers whom, partnership state | On demand | --- @@ -143,6 +145,10 @@ Sunday morning block (after nightly maintenance): Every 6h: system_tuning_monitor.sh ──────────────────► TUNING_MONITOR_LOG +On demand: + mesh_monitor.sh ──────────────────────────── reads host*.conf only (no API, no SSH) + MEMBERS / COVERAGE / PARTNERSHIP sections + Daily 8am: weekly_health_digest.sh ── reads ──────────► FALLBACK_STATE_FILE ── reads ──────────► WATCHDOG_STATE_FILE diff --git a/Orchestrators/README-Orchestrators.md b/Orchestrators/README-Orchestrators.md index f86fa5b..20962fb 100644 --- a/Orchestrators/README-Orchestrators.md +++ b/Orchestrators/README-Orchestrators.md @@ -104,7 +104,7 @@ array_started.sh # Cron — one entry per maintenance window: */3 * * * * transcode_management.sh 0 */6 * * * arrs_failed_stalled_recovery.sh -*/30 * * * * rsync.sh emby-failover ← not an orchestrator, direct call +*/30 * * * * rsync.sh emby-fallback ← not an orchestrator, direct call 0 1 * * * daily_sync_maintenance.sh 30 2 * * 0 weekly_sync_maintenance.sh @@ -730,7 +730,7 @@ maintenance block before the 7am coffee report. # ───────────────────────────────────────────────────────────────────────────── # Two Emby syncs run in parallel — dirty and clean: # -# emby-failover dirty sync (every 30 minutes, Emby running): +# emby-fallback dirty sync (every 30 minutes, Emby running): # watch states, library deltas, user activity — continuous coverage # WAL files excluded — safe to copy while Emby writes # HOST2 always within 30 minutes of HOST1 on playback state @@ -838,6 +838,50 @@ weekly_sync_maintenance.sh --status # show configured shares, jobs, update t --- +## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ +## 📆 monthly_maintenance.sh +## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ + +Uptime-triggered orchestrator for long-running system tasks — ZFS scrub, SMART +long tests — that should only run on stable systems that have been up for at least +30 days. Called daily by cron; most invocations are silent no-ops. + +Two gates must both pass before any job runs: +1. Server uptime ≥ `MONTHLY_UPTIME_THRESHOLD_DAYS` +2. Last run ≥ `MONTHLY_RUN_INTERVAL_DAYS` ago (state file on `/boot/config/` — survives reboots) + +If either gate fails, the script exits 0 with no output. This is expected — it runs +daily and most days are no-ops. + +`--force` bypasses both gates and runs the job list immediately. Use for testing or +when a scrub was missed and the server hasn't reached the uptime threshold yet. + +### Configuration (master.conf) + +```bash +MONTHLY_MAINTENANCE_SCRIPTS=( + #"Tools/zfs_pool_scrub.sh" + #"Tools/smart_long_test.sh" +) +MONTHLY_UPTIME_THRESHOLD_DAYS=30 +MONTHLY_RUN_INTERVAL_DAYS=30 +MONTHLY_LAST_RUN_FILE="/boot/config/monthly_maintenance_last_run.db" +``` + +Scripts are commented out by default — uncomment what applies to your hardware. + +### Usage + +```bash +monthly_maintenance.sh # normal run (daily cron — silent no-op when gates not met) +monthly_maintenance.sh --force # bypass both gates — run immediately +monthly_maintenance.sh --dry-run # show what would run without running it +monthly_maintenance.sh --status # show gate state: uptime, last run, next eligible +monthly_maintenance.sh --log # verbose output from each child script +``` + +--- + ## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ ## 🧹 media_management.sh ## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ @@ -906,7 +950,7 @@ array_started.sh # ───────────────────────────────────────────────────────────────────────────── # Periodic — not an orchestrator, direct rsync call: # ───────────────────────────────────────────────────────────────────────────── -*/30 * * * * rsync.sh /mnt/user/Media_Server/Emby --profile=emby-failover +*/30 * * * * rsync.sh /mnt/user/Media_Server/Emby --profile=emby-fallback # ───────────────────────────────────────────────────────────────────────────── # Every 6 hours: diff --git a/README-User_Script_Plug-in.md b/README-User_Script_Plug-in.md index a7ca458..a287775 100644 --- a/README-User_Script_Plug-in.md +++ b/README-User_Script_Plug-in.md @@ -296,7 +296,7 @@ Background: YES ```bash #!/bin/bash /mnt/user/appdata/unraid_scripts/Rsync/rsync.sh \ - /mnt/user/Media_Server/Emby --profile=emby-failover + /mnt/user/Media_Server/Emby --profile=emby-fallback ``` **What it does:** @@ -304,7 +304,7 @@ Keeps HOST2's Emby within 30 minutes of HOST1 on watch states, play positions, a library changes — with Emby running on both sides. Not a full sync (that's weekly) — this is continuous coverage of what matters for failover. -**What syncs (emby-failover profile):** +**What syncs (emby-fallback profile):** ``` Syncs: users.db, library.db, authentication.db, config/ Skips: *.wal *.shm — WAL files being written by Emby, unsafe to copy mid-write @@ -800,7 +800,7 @@ Background: YES **What syncs:** full clean Emby mirror — all databases checkpointed, metadata, plugins, config. Both Emby instances stopped during sync. Excludes transcodes, logs, cache, crash -files. This is the full-fidelity sync (vs emby-failover which is the dirty 30-minute sync). +files. This is the full-fidelity sync (vs emby-fallback which is the dirty 30-minute sync). --- diff --git a/README.md b/README.md index ac01498..40b7728 100644 --- a/README.md +++ b/README.md @@ -580,7 +580,7 @@ Every 15 minutes: critical_sync_maintenance.sh Critical-Data + Emby dirty sync + partnership check Every 30 minutes: - rsync.sh --profile=emby-failover Emby watch state sync (separate cron entry) + rsync.sh --profile=emby-fallback Emby watch state sync (separate cron entry) Every 6 hours: arrs_failed_stalled_recovery.sh blocklist + re-search failed imports diff --git a/Rsync/Manual-Rsync.md b/Rsync/Manual-Rsync.md index 9d15c8b..d0fe276 100644 --- a/Rsync/Manual-Rsync.md +++ b/Rsync/Manual-Rsync.md @@ -342,19 +342,19 @@ PROFILES["emby_EXCLUDE_DIRS"]=( "transcodes/" "logs/" "crash*" "cache/" ) -# ── emby-failover ───────────────────────────────────────────────────────────── +# ── emby-fallback ───────────────────────────────────────────────────────────── # Every 30 minutes, Emby STAYS RUNNING — dirty sync of critical state only -PROFILES["emby-failover_CRITICAL_CONTAINER_NAMES"]=() # nothing stops -PROFILES["emby-failover_EXCLUDE_DIRS"]=( +PROFILES["emby-fallback_CRITICAL_CONTAINER_NAMES"]=() # nothing stops +PROFILES["emby-fallback_EXCLUDE_DIRS"]=( "*.wal" "*.shm" # unsafe mid-write "transcodes/" "logs/" "crash*" "cache/" ) -PROFILES["emby-failover_REMOTE_RESTART_CONTAINERS"]=("Emby") +PROFILES["emby-fallback_REMOTE_RESTART_CONTAINERS"]=("Emby") ``` ### Two Emby Profiles — Why Both Exist -**emby-failover** (every 30 minutes, Emby stays running): +**emby-fallback** (every 30 minutes, Emby stays running): - Syncs: users.db, library.db, authentication.db, config/ - Skips: \*.wal, \*.shm, transcodes/, logs/, cache/ - Why: WAL files are written while Emby runs — copying them would produce a corrupt database on HOST2 @@ -366,7 +366,7 @@ PROFILES["emby-failover_REMOTE_RESTART_CONTAINERS"]=("Emby") - Why: WAL is checkpointed on clean shutdown — safe to copy everything - Result: HOST2 gets a gold-standard Emby state once per week -The two profiles work together. emby-failover keeps HOST2 current for immediate failover. +The two profiles work together. emby-fallback keeps HOST2 current for immediate failover. emby gives HOST2 full fidelity once per week. Neither alone is sufficient. --- @@ -448,7 +448,7 @@ Watchdogs (resource, docker, system, stability) run separately via watchdog_orch # Every 30 minutes — Emby dirty sync: */30 * * * * bash /mnt/user/appdata/unraid_scripts/Rsync/rsync.sh \ - /mnt/user/Media_Server/Emby --profile=emby-failover + /mnt/user/Media_Server/Emby --profile=emby-fallback # Every 6 hours — failed import + stalled download recovery: 0 */6 * * * @@ -731,5 +731,5 @@ name doesn't match any profile key, or when testing a specific profile. ```bash rsync.sh /mnt/user/appdata-Failover/Critical-Data --profile=critical-data -rsync.sh /mnt/user/Media_Server/Emby --profile=emby-failover +rsync.sh /mnt/user/Media_Server/Emby --profile=emby-fallback ``` diff --git a/Rsync/README-Rsync.md b/Rsync/README-Rsync.md index a146733..90cde98 100644 --- a/Rsync/README-Rsync.md +++ b/Rsync/README-Rsync.md @@ -68,7 +68,7 @@ Orchestrators/ ← decides what to sync, when, and in what order critical_sync_maintenance.sh ────────────────────────────────► rsync.sh (partnership shares) Cron (direct): - */30 * * * * ──────────────────────────────────► rsync.sh --profile=emby-failover + */30 * * * * ──────────────────────────────────► rsync.sh --profile=emby-fallback Monitors/ bandwidth_monitor.sh ◄─── called by rsync.sh after each sync (--log-transfer) @@ -94,7 +94,7 @@ Profile key = directory basename lowercased. `--profile=name` overrides. | `critical-data` | Auth stack | Mariadb-Authelia, Redis-Authelia, NPM, Lldap | Authelia has restart delay | | `important-data` | NextCloud + Postgres | Postgres-NextCloud | NextCloud has restart delay | | `emby` | Full Emby mirror | Emby (both sides) | Weekly — Sunday 2:30am | -| `emby-failover` | Emby watch state delta | None | Dirty sync — Emby stays running | +| `emby-fallback` | Emby watch state delta | None | Dirty sync — Emby stays running | For full profile definitions see `Manual-Rsync.md`. @@ -116,7 +116,7 @@ Callers (Orchestrators/) ────────────────── weekly_sync_maintenance.sh rsync.sh /path/to/share [--profile=name] │ critical_sync_maintenance.sh ─────────────────────────────────────────────► │ fallback.sh (writeback) │ - Direct cron (emby-failover) │ + Direct cron (emby-fallback) │ ▼ ┌─────────────────────────────────────┐ │ 1. Infer/accept profile │ diff --git a/Tools/Manual-Tools.md b/Tools/Manual-Tools.md index 13131af..b7c8723 100644 --- a/Tools/Manual-Tools.md +++ b/Tools/Manual-Tools.md @@ -11,7 +11,7 @@ making any changes. - [emby_to_lidarr_sync.sh](#emby_to_lidarr_syncsh) - [emby_to_sonarr_sync.sh](#emby_to_sonarr_syncsh) - [emby_to_radarr_sync.sh](#emby_to_radarr_syncsh) -- [failover_state_reset.sh](#failover_state_resetsh) +- [fallback_state_reset.sh](#fallback_state_resetsh) *(not yet built — manual workaround)* - [watchdog_skip_list_manager.sh](#watchdog_skip_list_managersh) - [bulk_permissions_repair.sh](#bulk_permissions_repairsh) - [container_data_export.sh](#container_data_exportsh) @@ -141,15 +141,17 @@ bash Tools/emby_to_radarr_sync.sh --- -## failover_state_reset.sh +## fallback_state_reset.sh -Resets the fallback state file to NORMAL and clears all tier flags. State file only — -does NOT start or stop any containers. +> **Not yet built.** Use the manual workaround below. + +Planned: reset the fallback state file to NORMAL and clear all tier flags. State file +only — does NOT start or stop any containers. ### When to Use ``` -After failover_test.sh didn't complete cleanly +After fallback_test.sh didn't complete cleanly → state left in FALLBACK but containers are actually back to normal After a failed handback @@ -162,13 +164,13 @@ After a dev/debug session → state left in a non-NORMAL state from testing ``` -### Verify Before Resetting - -Run `--status` first and check each of these before writing: +### Manual Workaround ```bash +# Verify before resetting: + # Right containers on right server? -continuous_scripts_status.sh # shows failover current state +continuous_scripts_status.sh # shows fallback current state # DDNS pointing correctly? nslookup Gmer4Lfe.com # confirm it resolves to the right IP @@ -178,24 +180,20 @@ pgrep -f "fallback.sh" # empty output = not running # Both servers Tailscale connected? tailscale status # both hosts should show active + +# Check current state file: +cat /boot/config/fallback_state.db + +# Reset to NORMAL (only after confirming containers and DDNS are correct): +echo "state=NORMAL" > /boot/config/fallback_state.db ``` -Resetting during an actual failover causes fallback.sh to think everything is normal +Resetting during an actual fallback causes fallback.sh to think everything is normal and stop covering the remote — services go offline until the next detection cycle. -### Usage +### What the State File Contains ```bash -fallback_state_reset.sh --status # show current state file — always check first -fallback_state_reset.sh --dry-run # show what would be written, no write -fallback_state_reset.sh # interactive reset — prompts for YES to confirm -fallback_state_reset.sh --force # non-interactive — for scripts, no terminal -``` - -### What Gets Written - -```bash -# New state file after reset: state=NORMAL fallback_start=0 handback_strikes=0 diff --git a/Watchdogs/Manual-Watchdogs.md b/Watchdogs/Manual-Watchdogs.md index bdf1839..232efdf 100644 --- a/Watchdogs/Manual-Watchdogs.md +++ b/Watchdogs/Manual-Watchdogs.md @@ -652,10 +652,8 @@ WATCHDOG_RESTART_CRASHED=true WATCHDOG_BATCH_NOTIFY=true # State files: -WATCHDOG_STATE_FILE="/tmp/watchdog_state.db" -SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db" +WATCHDOG_STATE_FILE="/tmp/container_watchdog_state.db" WATCHDOG_CONTAINER_RESTART_LOG="/boot/config/container_restart_history.db" -SYS_WATCHDOG_STATE_FILE="/tmp/sys_watchdog_state.db" # host*.conf HOST1_WATCHDOG_CONTAINERS=() # "ContainerName:LimitMB" @@ -679,6 +677,8 @@ STORAGE_WATCHDOG_STATE_FILE="/tmp/storage_watchdog_state.db" # ) # ── System Watchdog ──────────────────────────────────────────────────────────── +# Reboot trigger only — warn/shutdown/recover RAM tiers handled by resource_watchdog.sh +# RW_RAM_HARD_GB > SYS_WATCHDOG_MEM_GB always (resource_watchdog acts before watchdog reboots) SYS_WATCHDOG_STRIKE_LIMIT=2 SYSTEM_WATCHDOG_INTERVAL=300 @@ -686,10 +686,7 @@ SYS_WATCHDOG_REBOOT_WINDOW_HRS=2 SYS_WATCHDOG_MAX_REBOOTS=3 SYS_WATCHDOG_OOM_LIMIT=3 -MEM_WARN_GB=10 -MEM_SHUTDOWN_GB=6 -MEM_GB=4 -MEM_RECOVER_GB=30 +SYS_WATCHDOG_MEM_GB=4 # strike system → reboot (last resort — below resource_watchdog hard stop) SYS_WATCHDOG_ROOTFS_CRITICAL_PCT=99 SYS_WATCHDOG_FD_CRITICAL_PCT=95 @@ -703,13 +700,8 @@ SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true SYS_WATCHDOG_ABORT_ON_PARITY=true SYS_WATCHDOG_ABORT_ON_MOVER=true -SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED=( - "NginxProxyManager" "Authelia" "Mariadb" "Redis" "Emby" "Dispatcharr" -) -SYS_WATCHDOG_REQUIRED_CONTAINERS=() - # State files: -SYS_WATCHDOG_STATE_FILE="/tmp/sys_watchdog_state.db" +SYS_WATCHDOG_STATE_FILE="/tmp/system_watchdog_state.db" SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db" SYS_WATCHDOG_REBOOT_LOG="/boot/config/system_watchdog_reboots.db" SYS_WATCHDOG_OOM_FILE="/tmp/system_watchdog_oom.db"