Storage-mode awareness pass + doc update for System_Essentials through Partnership
All state/data file paths in scripts and PHP now resolve via STATE_DIR / DATA_DIR / PERSISTENT_CONF_CACHE instead of hardcoded /boot/config/ or /tmp/ paths, so the ecosystem works in both internal and appdata storage modes. PHP layer (watchdog.php, partnership.php, fallback.php, monitor.php, snapshot.php, config.php): all state reads switched to STATE_DIR constant; remote state reads use the new vv_remote_state_cmd() helper which resolves the remote's SCRIPTS_DIR via their varaverk.cfg before building the path. conf_sync.sh: fixed SCRIPTS_ROOT → SCRIPTS_DIR bug on MY_CONF path; added _remote_scripts_dir() to resolve partner's SCRIPTS_DIR before SCP pull. fallback.php page: added controls card (PARTNERSHIP_ENABLED, FALLBACK_ENABLED, FALLBACK_RSYNC_ENABLED toggles), status grid, and settings card. README and Manual updated for System_Essentials, Watchdogs, Fallback, Rsync, Media, Monitors, Orchestrators, Partnership: added new scripts (conf_sync, conf_cache_save/restore, conf_cache_watchdog, play_state_sync, start_webhook_listener, upgrade_webhook_handler), corrected all stale /boot/config/ state file paths to $STATE_DIR/$DATA_DIR, noted webgui/php_fpm/mover/user_scripts scripts moved to Plugin/unraid/System_Essentials, fixed start_webhook_listener.sh header (Node.js, not PHP -S).
This commit is contained in:
@@ -10,6 +10,7 @@ Configurations/*.bak
|
|||||||
# ── Runtime state, data, logs ─────────────────────────────────────────────────
|
# ── Runtime state, data, logs ─────────────────────────────────────────────────
|
||||||
data/
|
data/
|
||||||
State_Files/
|
State_Files/
|
||||||
|
.cache/
|
||||||
*.log
|
*.log
|
||||||
*.lock
|
*.lock
|
||||||
|
|
||||||
|
|||||||
@@ -123,6 +123,7 @@
|
|||||||
# Requirement: ALL state files MUST use $STATE_DIR. No /tmp, no /boot/config root.
|
# Requirement: ALL state files MUST use $STATE_DIR. No /tmp, no /boot/config root.
|
||||||
DATA_DIR="${SCRIPTS_DIR}/data"
|
DATA_DIR="${SCRIPTS_DIR}/data"
|
||||||
STATE_DIR="${SCRIPTS_DIR}/State_Files"
|
STATE_DIR="${SCRIPTS_DIR}/State_Files"
|
||||||
|
PERSISTENT_CONF_CACHE="${SCRIPTS_DIR}/.cache/vv/d"
|
||||||
|
|
||||||
# ── Version Parity ──
|
# ── Version Parity ──
|
||||||
# Controls behaviour when local and remote unRAID versions differ.
|
# Controls behaviour when local and remote unRAID versions differ.
|
||||||
@@ -295,7 +296,7 @@
|
|||||||
# watchdog_orchestrator.sh — NOT launched here.
|
# watchdog_orchestrator.sh — NOT launched here.
|
||||||
ARRAY_START_SCRIPTS=(
|
ARRAY_START_SCRIPTS=(
|
||||||
"Plugin/unraid/System_Essentials/unraid_api_key_renew.sh" # re-register Varaverk API key — registry is ephemeral
|
"Plugin/unraid/System_Essentials/unraid_api_key_renew.sh" # re-register Varaverk API key — registry is ephemeral
|
||||||
"System_Essentials/conf_sync.sh" # pull partner confs + push own conf into /tmp/.vv/ RAM cache
|
"System_Essentials/conf_sync.sh" # pull partner confs + push own conf into /tmp/.cache/vv/d/ RAM cache
|
||||||
"System_Essentials/conf_cache_restore.sh" # load partner confs from persistent backup if conf_sync couldn't reach partner
|
"System_Essentials/conf_cache_restore.sh" # load partner confs from persistent backup if conf_sync couldn't reach partner
|
||||||
"Transcodes/ramdisk_setup.sh" # creates ramdisk + symlink before Emby starts
|
"Transcodes/ramdisk_setup.sh" # creates ramdisk + symlink before Emby starts
|
||||||
"System_Essentials/docker_syslog_filter.sh" # suppress veth noise before logs fill
|
"System_Essentials/docker_syslog_filter.sh" # suppress veth noise before logs fill
|
||||||
|
|||||||
@@ -10,7 +10,7 @@
|
|||||||
# EMPTY fields, never overwrites existing values unless --overwrite is passed.
|
# EMPTY fields, never overwrites existing values unless --overwrite is passed.
|
||||||
#
|
#
|
||||||
# After populating, pushes the updated conf to all partners via conf_sync.sh
|
# After populating, pushes the updated conf to all partners via conf_sync.sh
|
||||||
# so they have the fresh keys in their /tmp/.vv/ cache immediately.
|
# so they have the fresh keys in their /tmp/.cache/vv/d/ cache immediately.
|
||||||
#
|
#
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# AUTO-DETECTED FIELDS
|
# AUTO-DETECTED FIELDS
|
||||||
|
|||||||
@@ -120,6 +120,7 @@
|
|||||||
# Requirement: ALL state files MUST use $STATE_DIR. No /tmp, no /boot/config root.
|
# Requirement: ALL state files MUST use $STATE_DIR. No /tmp, no /boot/config root.
|
||||||
DATA_DIR="/boot/config/plugins/varaverk/data"
|
DATA_DIR="/boot/config/plugins/varaverk/data"
|
||||||
STATE_DIR="/boot/config/plugins/varaverk/State_Files"
|
STATE_DIR="/boot/config/plugins/varaverk/State_Files"
|
||||||
|
PERSISTENT_CONF_CACHE="/boot/config/plugins/varaverk/.cache/vv/d"
|
||||||
|
|
||||||
# ── Version Parity ──
|
# ── Version Parity ──
|
||||||
# Controls behaviour when local and remote unRAID versions differ.
|
# Controls behaviour when local and remote unRAID versions differ.
|
||||||
|
|||||||
@@ -139,7 +139,7 @@ happens after an update, and you want to know when it does.
|
|||||||
A container keeps appearing in a broken state. You SSH in and see it's stopped. You
|
A container keeps appearing in a broken state. You SSH in and see it's stopped. You
|
||||||
don't know if the watchdog tried to restart it and failed, gave up and skip-listed it,
|
don't know if the watchdog tried to restart it and failed, gave up and skip-listed it,
|
||||||
is mid-attempt right now, or hasn't noticed yet. You have to manually check the skip
|
is mid-attempt right now, or hasn't noticed yet. You have to manually check the skip
|
||||||
list file on `/boot/config/`, check the restart history file, check the state file —
|
list file in `$STATE_DIR`, check the restart history file, check the state file —
|
||||||
none of which have obvious formats.
|
none of which have obvious formats.
|
||||||
|
|
||||||
The fix: `watchdog_skip_list_manager.sh` (in `Tools/`). One command to see exactly what's
|
The fix: `watchdog_skip_list_manager.sh` (in `Tools/`). One command to see exactly what's
|
||||||
|
|||||||
+22
-16
@@ -37,11 +37,11 @@ up required before handback sequence starts. **(default: 3)**
|
|||||||
---
|
---
|
||||||
|
|
||||||
```
|
```
|
||||||
FALLBACK_STATE_FILE=/boot/config/fallback_state.db
|
FALLBACK_STATE_FILE="$STATE_DIR/fallback_state.db"
|
||||||
```
|
```
|
||||||
Path to the persistent state file. Lives on `/boot/` intentionally — survives reboots.
|
Path to the persistent state file. In `$STATE_DIR` — survives reboots whether storage
|
||||||
If the server was in FALLBACK state when it rebooted, it resumes FALLBACK on restart
|
mode is internal (boot device) or appdata (array). If the server was in FALLBACK state
|
||||||
rather than assuming everything is normal.
|
when it rebooted, it resumes FALLBACK on restart rather than assuming everything is normal.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -172,7 +172,7 @@ daily_sync_maintenance.sh uses, in the opposite direction. No separate TIER4 lis
|
|||||||
FALLBACK_ENABLED=true
|
FALLBACK_ENABLED=true
|
||||||
FALLBACK_CHECK_INTERVAL=30
|
FALLBACK_CHECK_INTERVAL=30
|
||||||
FALLBACK_HANDBACK_STRIKES=3
|
FALLBACK_HANDBACK_STRIKES=3
|
||||||
FALLBACK_STATE_FILE=/boot/config/fallback_state.db
|
FALLBACK_STATE_FILE="$STATE_DIR/fallback_state.db"
|
||||||
FALLBACK_RSYNC_ENABLED=true
|
FALLBACK_RSYNC_ENABLED=true
|
||||||
EXTERNAL_IP=8.8.8.8
|
EXTERNAL_IP=8.8.8.8
|
||||||
FALLBACK_TEST_BLOCK_WAIT=60
|
FALLBACK_TEST_BLOCK_WAIT=60
|
||||||
@@ -293,7 +293,10 @@ FALLBACK_HOST2_WRITEBACK_TIER1=(
|
|||||||
|
|
||||||
## ━━━ STATE FILE REFERENCE ━━━
|
## ━━━ STATE FILE REFERENCE ━━━
|
||||||
|
|
||||||
Location: `/boot/config/fallback_state.db` (survives reboots)
|
Location: `$STATE_DIR/fallback_state.db` (survives reboots — boot device or appdata)
|
||||||
|
|
||||||
|
> In a shell where load_config.sh is not sourced, use the full path:
|
||||||
|
> `/boot/config/plugins/varaverk/State_Files/fallback_state.db` (internal storage mode)
|
||||||
|
|
||||||
```
|
```
|
||||||
state=NORMAL # NORMAL | FALLBACK | NO_INTERNET | DARK
|
state=NORMAL # NORMAL | FALLBACK | NO_INTERNET | DARK
|
||||||
@@ -304,8 +307,8 @@ tier3_started=false # whether Tier 3 containers started
|
|||||||
tier4_started=false # whether Tier 4 containers started
|
tier4_started=false # whether Tier 4 containers started
|
||||||
```
|
```
|
||||||
|
|
||||||
View state: `cat /boot/config/fallback_state.db`
|
View state: `fallback.sh --status` (preferred — parsed output)
|
||||||
Check state: `fallback.sh --status`
|
Raw file: `cat "$STATE_DIR/fallback_state.db"` (requires STATE_DIR set, or use full path)
|
||||||
|
|
||||||
The file is managed exclusively by fallback.sh. Do not edit it while fallback.sh is
|
The file is managed exclusively by fallback.sh. Do not edit it while fallback.sh is
|
||||||
running — the next cycle will overwrite your changes. Use the Manual State Reset
|
running — the next cycle will overwrite your changes. Use the Manual State Reset
|
||||||
@@ -384,7 +387,7 @@ servers must be running it continuously for mutual coverage.
|
|||||||
pgrep -f "fallback.sh"
|
pgrep -f "fallback.sh"
|
||||||
|
|
||||||
# Check the state file
|
# Check the state file
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
```
|
```
|
||||||
|
|
||||||
Start via User Scripts plugin on both servers.
|
Start via User Scripts plugin on both servers.
|
||||||
@@ -405,7 +408,7 @@ in smart mode — a non-NORMAL state at digest time needs attention.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
fallback.sh --status # full state snapshot
|
fallback.sh --status # full state snapshot
|
||||||
cat /boot/config/fallback_state.db # raw state file
|
cat "$STATE_DIR/fallback_state.db" # raw state file
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -426,7 +429,7 @@ fallback_test.sh
|
|||||||
|
|
||||||
# Step 3 — check state after test completes
|
# Step 3 — check state after test completes
|
||||||
fallback.sh --status
|
fallback.sh --status
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
```
|
```
|
||||||
|
|
||||||
If the test doesn't complete cleanly, the state file may be left in FALLBACK. The
|
If the test doesn't complete cleanly, the state file may be left in FALLBACK. The
|
||||||
@@ -459,11 +462,14 @@ ping -c 5 [remote-tailscale-ip]
|
|||||||
**Stop fallback.sh first (via User Scripts Abort), then reset:**
|
**Stop fallback.sh first (via User Scripts Abort), then reset:**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
# Set STATE_DIR (or source load_config.sh to get it from the environment)
|
||||||
|
source /boot/config/plugins/varaverk/load_config.sh
|
||||||
|
|
||||||
# View current state
|
# View current state
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
|
|
||||||
# Write a clean NORMAL state
|
# Write a clean NORMAL state
|
||||||
cat > /boot/config/fallback_state.db << 'EOF'
|
cat > "$STATE_DIR/fallback_state.db" << 'EOF'
|
||||||
state=NORMAL
|
state=NORMAL
|
||||||
fallback_start=0
|
fallback_start=0
|
||||||
handback_strikes=0
|
handback_strikes=0
|
||||||
@@ -473,7 +479,7 @@ tier4_started=false
|
|||||||
EOF
|
EOF
|
||||||
|
|
||||||
# Verify the write
|
# Verify the write
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
```
|
```
|
||||||
|
|
||||||
Restart fallback.sh via User Scripts plugin. It will resume from NORMAL on its next cycle.
|
Restart fallback.sh via User Scripts plugin. It will resume from NORMAL on its next cycle.
|
||||||
@@ -510,7 +516,7 @@ Can this server reach the remote Tailscale IP?
|
|||||||
|
|
||||||
What does fallback.sh report?
|
What does fallback.sh report?
|
||||||
→ fallback.sh --status
|
→ fallback.sh --status
|
||||||
→ cat /boot/config/fallback_state.db
|
→ cat "$STATE_DIR/fallback_state.db"
|
||||||
```
|
```
|
||||||
|
|
||||||
### Handback Not Completing
|
### Handback Not Completing
|
||||||
@@ -559,7 +565,7 @@ If it has happened:
|
|||||||
|
|
||||||
3. Understand the state before resetting
|
3. Understand the state before resetting
|
||||||
fallback.sh --status
|
fallback.sh --status
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
|
|
||||||
4. Perform Manual State Reset above on the server in a bad state
|
4. Perform Manual State Reset above on the server in a bad state
|
||||||
|
|
||||||
|
|||||||
@@ -130,11 +130,13 @@
|
|||||||
# STATE FILES
|
# STATE FILES
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
#
|
#
|
||||||
# FALLBACK_STATE_FILE — /boot/config/fallback_state.db (survives reboots)
|
# FALLBACK_STATE_FILE — $STATE_DIR/fallback_state.db (survives reboots)
|
||||||
# Keys: state, fallback_start, handback_strikes, tier2_started,
|
# Keys: state, fallback_start, handback_strikes, tier2_started,
|
||||||
# tier3_started, tier4_started, partnership_suspended, partner_lost_at.
|
# tier3_started, tier4_started, partnership_suspended, partner_lost_at.
|
||||||
# Lives on /boot/ intentionally — if the server was in FALLBACK when it
|
# Survives reboots — $STATE_DIR is on the boot device (internal) or appdata
|
||||||
# rebooted, it resumes FALLBACK on restart.
|
# (flash). Either way the array is up before this script runs, so the file
|
||||||
|
# is always accessible. If the server was in FALLBACK when it rebooted,
|
||||||
|
# it resumes FALLBACK on restart.
|
||||||
#
|
#
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# CONFIGURATION
|
# CONFIGURATION
|
||||||
@@ -185,7 +187,7 @@
|
|||||||
# Consecutive remote-up checks required before handback begins. (default: 3)
|
# Consecutive remote-up checks required before handback begins. (default: 3)
|
||||||
#
|
#
|
||||||
# FALLBACK_STATE_FILE
|
# FALLBACK_STATE_FILE
|
||||||
# State file path — /boot/config/fallback_state.db — survives reboots.
|
# State file path — $STATE_DIR/fallback_state.db — survives reboots.
|
||||||
#
|
#
|
||||||
# FALLBACK_RSYNC_ENABLED
|
# FALLBACK_RSYNC_ENABLED
|
||||||
# Gate for writeback rsync jobs during handback. (default: true)
|
# Gate for writeback rsync jobs during handback. (default: true)
|
||||||
@@ -302,7 +304,7 @@ log "$ICON_GEAR Timeouts: docker=${DOCKER_TIMEOUT}s ssh=${SSH_TIMEOUT}s verify-w
|
|||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── STATE FILE HELPERS ────────────────────────────────────────────────────────────────────────
|
# ── STATE FILE HELPERS ────────────────────────────────────────────────────────────────────────
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# State file on /boot/config — survives reboots.
|
# State file at $FALLBACK_STATE_FILE ($STATE_DIR/fallback_state.db) — survives reboots.
|
||||||
# Format: key=value one per line.
|
# Format: key=value one per line.
|
||||||
# Keys: state, fallback_start, handback_strikes, tier2_started, tier3_started, tier4_started
|
# Keys: state, fallback_start, handback_strikes, tier2_started, tier3_started, tier4_started
|
||||||
|
|
||||||
|
|||||||
@@ -409,7 +409,7 @@ pgrep -a -f fallback.sh
|
|||||||
|
|
||||||
Check current state:
|
Check current state:
|
||||||
```bash
|
```bash
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
# state=NORMAL — both servers up
|
# state=NORMAL — both servers up
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -510,7 +510,7 @@ bash Orchestrators/daily_sync_maintenance.sh --dry-run --log
|
|||||||
### ── Check fallback state ────────────────────────────────────────────────────
|
### ── Check fallback state ────────────────────────────────────────────────────
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
# Expected:
|
# Expected:
|
||||||
# state=NORMAL
|
# state=NORMAL
|
||||||
# fallback_start=0
|
# fallback_start=0
|
||||||
@@ -600,7 +600,7 @@ The remote server is unreachable on the first check. Common causes:
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Confirm which state fallback is in
|
# Confirm which state fallback is in
|
||||||
cat /boot/config/fallback_state.db
|
cat "$STATE_DIR/fallback_state.db"
|
||||||
# If stuck in FALLBACK after remote comes back: reset state
|
# If stuck in FALLBACK after remote comes back: reset state
|
||||||
bash Tools/fallback_state_reset.sh
|
bash Tools/fallback_state_reset.sh
|
||||||
```
|
```
|
||||||
|
|||||||
+45
-4
@@ -132,8 +132,8 @@ LIDARR_IMPORT_SCAN_TIMEOUT=600 # seconds to wait for pre-flight import scan
|
|||||||
LIDARR_VERSION_MAJOR=3 # expected Lidarr major version (API safety check)
|
LIDARR_VERSION_MAJOR=3 # expected Lidarr major version (API safety check)
|
||||||
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
||||||
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
||||||
LIDARR_TRACKED_COUNT_FILE=/boot/config/lidarr_tracked_count # persistent baseline
|
LIDARR_TRACKED_COUNT_FILE="$DATA_DIR/lidarr_tracked.count" # persistent baseline
|
||||||
ARR_CLEANUP_STATS=/boot/config/arr_cleanup_stats.db # read by coffee report
|
ARR_CLEANUP_STATS="$DATA_DIR/arr_cleanup_stats.db" # read by coffee report
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -171,7 +171,7 @@ RADARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.a
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
ARR_SYNC_ENABLED=true
|
ARR_SYNC_ENABLED=true
|
||||||
ARR_SYNC_BLOCKLIST=/boot/config/arr_sync_blocklist.tsv # tombstone file
|
ARR_SYNC_BLOCKLIST="$DATA_DIR/arr_sync_blocklist.tsv" # tombstone file
|
||||||
ARR_SYNC_CONNECT_TIMEOUT=10 # SSH connect timeout in seconds
|
ARR_SYNC_CONNECT_TIMEOUT=10 # SSH connect timeout in seconds
|
||||||
ARR_SYNC_API_TIMEOUT=60 # curl API call timeout in seconds
|
ARR_SYNC_API_TIMEOUT=60 # curl API call timeout in seconds
|
||||||
DOCKER_APPDATA_BASE=/mnt/user/appdata
|
DOCKER_APPDATA_BASE=/mnt/user/appdata
|
||||||
@@ -189,7 +189,7 @@ ARR_IMPORT_RECOVERY_AGE=6 # hours — items newer than this are skipped
|
|||||||
SONARR_VERSION_MAJOR=4
|
SONARR_VERSION_MAJOR=4
|
||||||
RADARR_VERSION_MAJOR=6
|
RADARR_VERSION_MAJOR=6
|
||||||
LIDARR_VERSION_MAJOR=3
|
LIDARR_VERSION_MAJOR=3
|
||||||
ARR_RECOVERY_STATS=/boot/config/arr_recovery_stats.db # read by coffee report
|
ARR_RECOVERY_STATS="$DATA_DIR/arr_recovery_stats.db" # read by coffee report
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -274,6 +274,47 @@ Requires `HOST*_TMDB_API_KEY` in `host*.conf`.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
### Play State Sync
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PLAY_STATE_SYNC_ENABLED=true # toggle entire sync
|
||||||
|
PLAY_STATE_SYNC_LOOKBACK_DAYS=30 # history window for played items
|
||||||
|
```
|
||||||
|
|
||||||
|
Emby and Jellyfin servers configured per-host:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# host*.conf
|
||||||
|
HOST1_EMBY_URL="http://192.168.50.2:8096"
|
||||||
|
HOST1_EMBY_API_KEY="..."
|
||||||
|
HOST1_JELLYFIN_URL="" # empty = skip Jellyfin on this host
|
||||||
|
HOST1_JELLYFIN_API_KEY=""
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Upgrade Webhook
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# master.conf
|
||||||
|
WEBHOOK_PORT=9456 # 0 = disable listener
|
||||||
|
WEBHOOK_SECRET="" # auto-generated on first start if empty
|
||||||
|
WEBHOOK_UPGRADE_RSYNC_ENABLED=true # push upgraded file to remote nodes
|
||||||
|
```
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# host*.conf
|
||||||
|
HOST1_WEBHOOK_LAN_IP="192.168.50.2" # IP arrs call for webhook delivery
|
||||||
|
```
|
||||||
|
|
||||||
|
When `WEBHOOK_PORT=0`, `start_webhook_listener.sh` exits cleanly and no listener starts.
|
||||||
|
When `WEBHOOK_SECRET` is empty, a 32-byte hex secret is generated on first start and
|
||||||
|
written back to `master.conf`. Run `Tools/webhook_setup.sh` to register the URL in arrs.
|
||||||
|
|
||||||
|
Webhook log: `/var/log/varaverk/upgrade_webhook.log`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
### Orchestrator Job Order
|
### Orchestrator Job Order
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
|||||||
+48
-2
@@ -4,8 +4,9 @@ Library health, consistency, sync, and behavior-driven discovery for a multi-ser
|
|||||||
stack. Correct permissions so arrs can manage files. Junk removal so orphan detection
|
stack. Correct permissions so arrs can manage files. Junk removal so orphan detection
|
||||||
isn't confused by scene debris. Library sync so every node tracks the same content.
|
isn't confused by scene debris. Library sync so every node tracks the same content.
|
||||||
Orphan cleanup against live arr APIs so deleted content actually leaves disk. Emby
|
Orphan cleanup against live arr APIs so deleted content actually leaves disk. Emby
|
||||||
notified automatically after every deletion. Weekly discovery adds new music, movies,
|
notified automatically after every deletion. Watch state synced across Emby/Jellyfin
|
||||||
and TV shows based on what you actually play — no manual browsing required.
|
every 30 minutes. Quality upgrades propagate to all nodes immediately via webhook.
|
||||||
|
Weekly discovery adds new music, movies, and TV shows based on what you actually play.
|
||||||
|
|
||||||
> **These scripts permanently delete files.** The arr cleanup scripts are protected by
|
> **These scripts permanently delete files.** The arr cleanup scripts are protected by
|
||||||
> multiple safety layers that must all pass before anything is touched — but dry runs and
|
> multiple safety layers that must all pass before anything is touched — but dry runs and
|
||||||
@@ -91,6 +92,19 @@ free overnight recovery.
|
|||||||
`lidarr_missing_art.sh` — fetch missing album and artist artwork from fanart.tv and
|
`lidarr_missing_art.sh` — fetch missing album and artist artwork from fanart.tv and
|
||||||
fallback sources. Never overwrites existing files.
|
fallback sources. Never overwrites existing files.
|
||||||
|
|
||||||
|
**Play State Sync**
|
||||||
|
`play_state_sync.sh` — syncs watched/played state and resume positions across all
|
||||||
|
configured Emby and Jellyfin servers. Newest timestamp wins. Runs every 30 minutes
|
||||||
|
via critical_sync_maintenance.sh.
|
||||||
|
|
||||||
|
**Upgrade Propagation**
|
||||||
|
`start_webhook_listener.sh` — Node.js HTTP server that receives Sonarr/Radarr/Lidarr
|
||||||
|
OnUpgrade webhooks. Continuous; started at array start. Writes to
|
||||||
|
`/var/log/varaverk/upgrade_webhook.log`.
|
||||||
|
`upgrade_webhook_handler.sh` — triggered by the webhook listener. Pushes the upgraded
|
||||||
|
item folder to every remote node immediately, then triggers an arr library rescan on
|
||||||
|
each remote so the upgraded file is accepted without triggering a redundant quality search.
|
||||||
|
|
||||||
**Discovery**
|
**Discovery**
|
||||||
`playback_aware_lidarr_discovery.sh` — behavior-driven music discovery. Scores your
|
`playback_aware_lidarr_discovery.sh` — behavior-driven music discovery. Scores your
|
||||||
Emby play history, runs Last.fm getSimilar on top artists, adds the best matches to
|
Emby play history, runs Last.fm getSimilar on top artists, adds the best matches to
|
||||||
@@ -138,6 +152,24 @@ playback_aware_radarr_discovery.sh — score watch history → TMDB recommenda
|
|||||||
playback_aware_sonarr_discovery.sh — score episode history → TMDB TV recommendations → add to Sonarr
|
playback_aware_sonarr_discovery.sh — score episode history → TMDB TV recommendations → add to Sonarr
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**Every 30 min via `critical_sync_maintenance.sh` (CRITICAL_MAINTENANCE_SCRIPTS):**
|
||||||
|
|
||||||
|
```
|
||||||
|
play_state_sync.sh — sync watched/resume state across Emby + Jellyfin
|
||||||
|
```
|
||||||
|
|
||||||
|
**Continuous (started by `array_started.sh`):**
|
||||||
|
|
||||||
|
```
|
||||||
|
start_webhook_listener.sh — Node.js webhook server; dispatch upgrade_webhook_handler.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
**On every arr upgrade (triggered by webhook):**
|
||||||
|
|
||||||
|
```
|
||||||
|
upgrade_webhook_handler.sh — push upgraded folder to all remote nodes + trigger arr rescan
|
||||||
|
```
|
||||||
|
|
||||||
**Ad-hoc or separate schedule:**
|
**Ad-hoc or separate schedule:**
|
||||||
|
|
||||||
```
|
```
|
||||||
@@ -189,6 +221,9 @@ in `HOST*_MEDIA_PERMISSION_SHARES` and `HOST*_MEDIA_CLEAN_FOLDERS` in host*.conf
|
|||||||
| `lidarr_missing_art.sh` | Fetch missing album and artist artwork | Ad-hoc or separate schedule |
|
| `lidarr_missing_art.sh` | Fetch missing album and artist artwork | Ad-hoc or separate schedule |
|
||||||
| `radarr_tmdb_removed.sh` | Remove movies dropped from TMDb | Ad-hoc or weekly |
|
| `radarr_tmdb_removed.sh` | Remove movies dropped from TMDb | Ad-hoc or weekly |
|
||||||
| `sonarr_tvdb_removed.sh` | Remove series dropped from TVDB | Ad-hoc or weekly |
|
| `sonarr_tvdb_removed.sh` | Remove series dropped from TVDB | Ad-hoc or weekly |
|
||||||
|
| `play_state_sync.sh` | Sync watched/played state + resume positions across Emby + Jellyfin | Every 30 min via `critical_sync_maintenance.sh` |
|
||||||
|
| `start_webhook_listener.sh` | Node.js webhook server — receive arr OnUpgrade and dispatch handler | Continuous (started by `array_started.sh`) |
|
||||||
|
| `upgrade_webhook_handler.sh` | Push upgraded item folder to remote nodes + trigger arr rescan | On each arr upgrade (via webhook) |
|
||||||
| `playback_aware_lidarr_discovery.sh` | Behavior-driven music discovery — Emby plays → Last.fm similar → Lidarr | Weekly via `weekly_sync_maintenance.sh` |
|
| `playback_aware_lidarr_discovery.sh` | Behavior-driven music discovery — Emby plays → Last.fm similar → Lidarr | Weekly via `weekly_sync_maintenance.sh` |
|
||||||
| `playback_aware_radarr_discovery.sh` | Behavior-driven movie discovery — Emby watches → TMDB recommendations → Radarr | Weekly via `weekly_sync_maintenance.sh` |
|
| `playback_aware_radarr_discovery.sh` | Behavior-driven movie discovery — Emby watches → TMDB recommendations → Radarr | Weekly via `weekly_sync_maintenance.sh` |
|
||||||
| `playback_aware_sonarr_discovery.sh` | Behavior-driven TV discovery — Emby episodes → TMDB TV recommendations → Sonarr | Weekly via `weekly_sync_maintenance.sh` |
|
| `playback_aware_sonarr_discovery.sh` | Behavior-driven TV discovery — Emby episodes → TMDB TV recommendations → Sonarr | Weekly via `weekly_sync_maintenance.sh` |
|
||||||
@@ -228,6 +263,17 @@ Weekly discovery (WEEKLY_MAINTENANCE_SCRIPTS):
|
|||||||
│
|
│
|
||||||
└── each discovery script fires arr search immediately after successful add
|
└── each discovery script fires arr search immediately after successful add
|
||||||
|
|
||||||
|
Every 30 min (critical_sync_maintenance.sh CRITICAL_MAINTENANCE_SCRIPTS):
|
||||||
|
play_state_sync.sh ─── newest timestamp wins → watched/resume state synced
|
||||||
|
across all configured Emby + Jellyfin servers
|
||||||
|
|
||||||
|
Continuous (started by array_started.sh):
|
||||||
|
start_webhook_listener.sh ── Node.js HTTP server listens on WEBHOOK_PORT
|
||||||
|
│ arr OnUpgrade fires webhook → POST to http://HOST_LAN_IP:WEBHOOK_PORT/webhook?key=SECRET
|
||||||
|
└── upgrade_webhook_handler.sh
|
||||||
|
├── rsync upgraded folder → all remote nodes immediately
|
||||||
|
└── trigger arr library rescan on each remote (accept new file, no quality search)
|
||||||
|
|
||||||
Ad-hoc enrichment:
|
Ad-hoc enrichment:
|
||||||
lidarr_missing_art.sh ─────── discovers missing artwork → fetches from fanart.tv
|
lidarr_missing_art.sh ─────── discovers missing artwork → fetches from fanart.tv
|
||||||
radarr_tmdb_removed.sh ────── status="deleted" → remove from Radarr + add exclusion
|
radarr_tmdb_removed.sh ────── status="deleted" → remove from Radarr + add exclusion
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
#
|
#
|
||||||
# PURPOSE
|
# PURPOSE
|
||||||
# ─────────────────────────────────────────────────────────────────────────────
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
# Starts a standalone PHP HTTP server that receives Sonarr/Radarr/Lidarr
|
# Starts a standalone Node.js HTTP server that receives Sonarr/Radarr/Lidarr
|
||||||
# OnUpgrade webhooks and dispatches upgrade_webhook_handler.sh.
|
# OnUpgrade webhooks and dispatches upgrade_webhook_handler.sh.
|
||||||
#
|
#
|
||||||
# Runs outside Unraid nginx — no session auth required. The shared secret in
|
# Runs outside Unraid nginx — no session auth required. The shared secret in
|
||||||
@@ -13,9 +13,12 @@
|
|||||||
#
|
#
|
||||||
# http://<HOST_LAN_IP>:<WEBHOOK_PORT>/webhook?key=<WEBHOOK_SECRET>
|
# http://<HOST_LAN_IP>:<WEBHOOK_PORT>/webhook?key=<WEBHOOK_SECRET>
|
||||||
#
|
#
|
||||||
# Runs as a continuous script started by array_started.sh. Execs php -S which
|
# Runs as a continuous script started by array_started.sh. Execs node which
|
||||||
# replaces this process — the PID stays the same for array_started.sh's check.
|
# replaces this process — the PID stays the same for array_started.sh's check.
|
||||||
#
|
#
|
||||||
|
# Uses Node.js instead of php -S: php -S on Unraid PHP 8.4 silently drops
|
||||||
|
# POST request bodies, making webhook payloads arrive empty.
|
||||||
|
#
|
||||||
# If WEBHOOK_SECRET is empty in master.conf: generates and saves one, then starts.
|
# If WEBHOOK_SECRET is empty in master.conf: generates and saves one, then starts.
|
||||||
# If WEBHOOK_PORT is 0: exits cleanly (disables the listener).
|
# If WEBHOOK_PORT is 0: exits cleanly (disables the listener).
|
||||||
#
|
#
|
||||||
|
|||||||
@@ -151,14 +151,14 @@ BACKUP_VERIFY_MIN_SIZE="1M" # skip files smaller than this
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# master.conf
|
# master.conf
|
||||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db" # survives reboots
|
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db" # survives reboots
|
||||||
BANDWIDTH_LOG_RETENTION=90 # days — file stays bounded, never grows unbounded
|
BANDWIDTH_LOG_RETENTION=90 # days — file stays bounded, never grows unbounded
|
||||||
BANDWIDTH_WARN_GB=50 # flag transfers or daily totals exceeding this
|
BANDWIDTH_WARN_GB=50 # flag transfers or daily totals exceeding this
|
||||||
```
|
```
|
||||||
|
|
||||||
**Why `/boot/config/`**: The log needs to survive reboots to build a useful history.
|
**Why `$DATA_DIR`**: The log needs to survive reboots to build a useful history.
|
||||||
`/boot/config/` is on the USB flash drive, which survives reboots and is backed up
|
`$DATA_DIR` (`$SCRIPTS_DIR/data/`) is on the boot device (internal) or appdata (flash),
|
||||||
by unRAID's flash backup. The log is bounded by `BANDWIDTH_LOG_RETENTION` so it never
|
both of which survive reboots. The log is bounded by `BANDWIDTH_LOG_RETENTION` so it never
|
||||||
grows unbounded.
|
grows unbounded.
|
||||||
|
|
||||||
**`BANDWIDTH_WARN_GB`**: Set to a value that represents "unexpectedly large" for your
|
**`BANDWIDTH_WARN_GB`**: Set to a value that represents "unexpectedly large" for your
|
||||||
@@ -401,7 +401,7 @@ BACKUP_VERIFY_SAMPLE=10
|
|||||||
BACKUP_VERIFY_MIN_SIZE="1M"
|
BACKUP_VERIFY_MIN_SIZE="1M"
|
||||||
|
|
||||||
# bandwidth_monitor.sh
|
# bandwidth_monitor.sh
|
||||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db"
|
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db"
|
||||||
BANDWIDTH_LOG_RETENTION=90
|
BANDWIDTH_LOG_RETENTION=90
|
||||||
BANDWIDTH_WARN_GB=50
|
BANDWIDTH_WARN_GB=50
|
||||||
|
|
||||||
|
|||||||
@@ -849,7 +849,7 @@ long tests — that should only run on stable systems that have been up for at l
|
|||||||
|
|
||||||
Two gates must both pass before any job runs:
|
Two gates must both pass before any job runs:
|
||||||
1. Server uptime ≥ `MONTHLY_UPTIME_THRESHOLD_DAYS`
|
1. Server uptime ≥ `MONTHLY_UPTIME_THRESHOLD_DAYS`
|
||||||
2. Last run ≥ `MONTHLY_RUN_INTERVAL_DAYS` ago (state file on `/boot/config/` — survives reboots)
|
2. Last run ≥ `MONTHLY_RUN_INTERVAL_DAYS` ago (state file in `$STATE_DIR` — survives reboots)
|
||||||
|
|
||||||
If either gate fails, the script exits 0 with no output. This is expected — it runs
|
If either gate fails, the script exits 0 with no output. This is expected — it runs
|
||||||
daily and most days are no-ops.
|
daily and most days are no-ops.
|
||||||
@@ -866,7 +866,7 @@ MONTHLY_MAINTENANCE_SCRIPTS=(
|
|||||||
)
|
)
|
||||||
MONTHLY_UPTIME_THRESHOLD_DAYS=30
|
MONTHLY_UPTIME_THRESHOLD_DAYS=30
|
||||||
MONTHLY_RUN_INTERVAL_DAYS=30
|
MONTHLY_RUN_INTERVAL_DAYS=30
|
||||||
MONTHLY_LAST_RUN_FILE="/boot/config/monthly_maintenance_last_run.db"
|
MONTHLY_LAST_RUN_FILE="$STATE_DIR/monthly_maintenance_last_run.db"
|
||||||
```
|
```
|
||||||
|
|
||||||
Scripts are commented out by default — uncomment what applies to your hardware.
|
Scripts are commented out by default — uncomment what applies to your hardware.
|
||||||
|
|||||||
@@ -6,7 +6,7 @@
|
|||||||
# Schedule: 0 */4 * * * (every 4 hours)
|
# Schedule: 0 */4 * * * (every 4 hours)
|
||||||
#
|
#
|
||||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||||
# 1. conf_sync.sh --pull-only — refresh partner conf cache in RAM (/tmp/.vv/)
|
# 1. conf_sync.sh --pull-only — refresh partner conf cache in RAM (/tmp/.cache/vv/d/)
|
||||||
# 2. arr_sync.sh — sync Lidarr/Sonarr/Radarr libraries across all nodes
|
# 2. arr_sync.sh — sync Lidarr/Sonarr/Radarr libraries across all nodes
|
||||||
# 3. Rsync window (optional) — INTERMEDIATE_SYNC_SHARES, if any configured
|
# 3. Rsync window (optional) — INTERMEDIATE_SYNC_SHARES, if any configured
|
||||||
# 4. INTERMEDIATE_MAINTENANCE_SCRIPTS — artwork fetch and any future 4-hour jobs
|
# 4. INTERMEDIATE_MAINTENANCE_SCRIPTS — artwork fetch and any future 4-hour jobs
|
||||||
|
|||||||
@@ -574,7 +574,7 @@ Partnership/partnership_manager.sh --status # shows both sides via SSH
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Check the offline counter:
|
# Check the offline counter:
|
||||||
cat /boot/config/partnership_offline_days.db
|
cat "$STATE_DIR/partnership_offline_days.db"
|
||||||
|
|
||||||
# Extended Tailscale outage may have incremented the counter.
|
# Extended Tailscale outage may have incremented the counter.
|
||||||
# Check Tailscale peer visibility:
|
# Check Tailscale peer visibility:
|
||||||
@@ -591,9 +591,9 @@ Partnership/partnership_manager.sh --onboard
|
|||||||
## ━━━ STATE FILES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
## ━━━ STATE FILES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
/boot/config/partnership_HOST1.db # HOST1 writes only
|
$STATE_DIR/partnership_HOST1.db # HOST1 writes only
|
||||||
/boot/config/partnership_HOST2.db # HOST2 writes only
|
$STATE_DIR/partnership_HOST2.db # HOST2 writes only
|
||||||
/boot/config/partnership_blocklist.db # hostname|timestamp|reason
|
$STATE_DIR/partnership_blocklist.db # hostname|timestamp|reason
|
||||||
|
|
||||||
# Example state file:
|
# Example state file:
|
||||||
state=ACTIVE
|
state=ACTIVE
|
||||||
|
|||||||
@@ -93,11 +93,11 @@
|
|||||||
# STATE FILES
|
# STATE FILES
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
#
|
#
|
||||||
# /boot/config/partnership_HOST1.db — HOST1 writes, HOST2 reads via SSH
|
# $STATE_DIR/partnership_<hostname>.db — each host writes its own, partner reads via SSH
|
||||||
# /boot/config/partnership_HOST2.db — HOST2 writes, HOST1 reads via SSH
|
# $STATE_DIR/partnership_blocklist.db — hostname|timestamp|reason, persists until cleared
|
||||||
# /boot/config/partnership_blocklist.db — hostname|timestamp|reason, persists until cleared
|
# $STATE_DIR/partnership_offline_days.db — cumulative offline day counter
|
||||||
#
|
#
|
||||||
# On /boot/config — survives reboots, available before array starts, minimal flash wear.
|
# All in STATE_DIR — survives reboots (on /boot in internal mode, appdata in flash mode).
|
||||||
#
|
#
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# CONFIGURATION
|
# CONFIGURATION
|
||||||
|
|||||||
@@ -274,7 +274,7 @@ platform_get_templates_dir() {
|
|||||||
# Writes the path to the persistent Varaverk setup/wizard state database.
|
# Writes the path to the persistent Varaverk setup/wizard state database.
|
||||||
# ──────────────────────────────────────────────────────────────────────────────────────────────
|
# ──────────────────────────────────────────────────────────────────────────────────────────────
|
||||||
platform_setup_db_path() {
|
platform_setup_db_path() {
|
||||||
echo "/boot/config/varaverk_setup.db"
|
echo "${STATE_DIR}/varaverk_setup.db"
|
||||||
}
|
}
|
||||||
|
|
||||||
# ──────────────────────────────────────────────────────────────────────────────────────────────
|
# ──────────────────────────────────────────────────────────────────────────────────────────────
|
||||||
|
|||||||
@@ -13,7 +13,7 @@ $ramUsedMb = $ramTotalMb - $res['ram_free_mb'];
|
|||||||
$ramPct = $ramTotalMb > 0 ? (int)round($ramUsedMb / $ramTotalMb * 100) : 0;
|
$ramPct = $ramTotalMb > 0 ? (int)round($ramUsedMb / $ramTotalMb * 100) : 0;
|
||||||
|
|
||||||
// Fallback state (fast file read, no exec)
|
// Fallback state (fast file read, no exec)
|
||||||
$fbRaw = @file_get_contents('/tmp/fallback_state.db') ?: '';
|
$fbRaw = @file_get_contents(STATE_DIR . '/fallback_state.db') ?: '';
|
||||||
$fbData = vv_parse_kv_db($fbRaw);
|
$fbData = vv_parse_kv_db($fbRaw);
|
||||||
$fallbackState = $fbData['state'] ?? 'UNKNOWN';
|
$fallbackState = $fbData['state'] ?? 'UNKNOWN';
|
||||||
|
|
||||||
|
|||||||
@@ -34,8 +34,7 @@ function vv_setup_state_write(array $data): void {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Push the setup state file to all remote hosts via scp.
|
// Push the setup state file to all remote hosts via scp.
|
||||||
// Unlike master.conf push, this does NOT require the plugin to be installed on the remote —
|
// Reads the remote's varaverk.cfg to find their actual SCRIPTS_DIR (handles appdata mode).
|
||||||
// it only needs SSH to be reachable, and pushes to /boot/config/ (always available).
|
|
||||||
function vv_push_setup_state(): void {
|
function vv_push_setup_state(): void {
|
||||||
if (!file_exists(VV_SETUP_STATE_FILE)) return;
|
if (!file_exists(VV_SETUP_STATE_FILE)) return;
|
||||||
$myHostId = vv_detect_host();
|
$myHostId = vv_detect_host();
|
||||||
@@ -56,9 +55,20 @@ function vv_push_setup_state(): void {
|
|||||||
if (!$ip) continue;
|
if (!$ip) continue;
|
||||||
$sshBase = 'ssh -i ' . escapeshellarg($sshKey)
|
$sshBase = 'ssh -i ' . escapeshellarg($sshKey)
|
||||||
. ' -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@' . $ip;
|
. ' -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@' . $ip;
|
||||||
// Ensure the target dir exists (it always should on Unraid, but be safe)
|
|
||||||
shell_exec($sshBase . ' "mkdir -p /boot/config" 2>/dev/null');
|
// Get remote SCRIPTS_DIR from varaverk.cfg — handles appdata mode on remote.
|
||||||
$dest = escapeshellarg('root@' . $ip . ':/boot/config/varaverk_setup.db');
|
// Falls back to the default install path if varaverk.cfg is absent (pre-install).
|
||||||
|
$cfgRaw = trim(shell_exec($sshBase . ' "cat /boot/config/plugins/varaverk/varaverk.cfg 2>/dev/null"') ?: '');
|
||||||
|
$remoteSD = '/boot/config/plugins/varaverk';
|
||||||
|
foreach (explode("\n", $cfgRaw) as $line) {
|
||||||
|
if (str_starts_with(trim($line), 'SCRIPTS_DIR=')) {
|
||||||
|
$remoteSD = trim(substr(trim($line), strlen('SCRIPTS_DIR=')), '"\'');
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
$remoteStatePath = $remoteSD . '/State_Files/varaverk_setup.db';
|
||||||
|
shell_exec($sshBase . ' "mkdir -p ' . escapeshellarg(dirname($remoteStatePath)) . '" 2>/dev/null');
|
||||||
|
$dest = escapeshellarg('root@' . $ip . ':' . $remoteStatePath);
|
||||||
exec('scp -i ' . escapeshellarg($sshKey)
|
exec('scp -i ' . escapeshellarg($sshKey)
|
||||||
. ' -o ConnectTimeout=10 -o StrictHostKeyChecking=no'
|
. ' -o ConnectTimeout=10 -o StrictHostKeyChecking=no'
|
||||||
. ' ' . escapeshellarg(VV_SETUP_STATE_FILE) . ' ' . $dest . ' 2>&1');
|
. ' ' . escapeshellarg(VV_SETUP_STATE_FILE) . ' ' . $dest . ' 2>&1');
|
||||||
@@ -391,6 +401,15 @@ function vv_auto_create_api_key(string $hostId, string $confFile): array {
|
|||||||
return ['ok' => true, 'key_preview' => $key ? substr($key, 0, 8) . '...' . substr($key, -4) : 'registered'];
|
return ['ok' => true, 'key_preview' => $key ? substr($key, 0, 8) . '...' . substr($key, -4) : 'registered'];
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Build a bash command that reads a state file from the REMOTE host's State_Files/.
|
||||||
|
// Reads the remote's varaverk.cfg to resolve their SCRIPTS_DIR (may differ from ours
|
||||||
|
// when the remote is in appdata mode). Falls back to the internal plugin path.
|
||||||
|
function vv_remote_state_cmd(string $filename): string {
|
||||||
|
$fn = basename($filename);
|
||||||
|
return 'sd=$(grep -m1 SCRIPTS_DIR= /boot/config/plugins/varaverk/varaverk.cfg 2>/dev/null'
|
||||||
|
. ' | cut -d\'"\' -f2); cat "${sd:-/boot/config/plugins/varaverk}/State_Files/' . $fn . '" 2>/dev/null';
|
||||||
|
}
|
||||||
|
|
||||||
// Local LAN IP via routing table — static-cached per request.
|
// Local LAN IP via routing table — static-cached per request.
|
||||||
// Previously duplicated in include/docker_folders.php and inline in include/docker.php.
|
// Previously duplicated in include/docker_folders.php and inline in include/docker.php.
|
||||||
function vv_local_ip(): string {
|
function vv_local_ip(): string {
|
||||||
|
|||||||
@@ -18,12 +18,14 @@ function vv_fb_scalar(string $raw, string $varname): string {
|
|||||||
|
|
||||||
function vv_fb_parse_state(string $text): array {
|
function vv_fb_parse_state(string $text): array {
|
||||||
$out = [
|
$out = [
|
||||||
'state' => 'UNKNOWN',
|
'state' => 'UNKNOWN',
|
||||||
'fallback_start' => 0,
|
'fallback_start' => 0,
|
||||||
'handback_strikes' => 0,
|
'handback_strikes' => 0,
|
||||||
'tier2_started' => false,
|
'tier2_started' => false,
|
||||||
'tier3_started' => false,
|
'tier3_started' => false,
|
||||||
'tier4_started' => false,
|
'tier4_started' => false,
|
||||||
|
'partnership_suspended' => false,
|
||||||
|
'partner_lost_at' => 0,
|
||||||
];
|
];
|
||||||
foreach (explode("\n", $text) as $line) {
|
foreach (explode("\n", $text) as $line) {
|
||||||
$line = trim($line);
|
$line = trim($line);
|
||||||
@@ -31,24 +33,26 @@ function vv_fb_parse_state(string $text): array {
|
|||||||
[$k, $v] = array_pad(explode('=', $line, 2), 2, '');
|
[$k, $v] = array_pad(explode('=', $line, 2), 2, '');
|
||||||
$k = trim($k); $v = trim($v, '"\'');
|
$k = trim($k); $v = trim($v, '"\'');
|
||||||
switch ($k) {
|
switch ($k) {
|
||||||
case 'state': $out['state'] = $v; break;
|
case 'state': $out['state'] = $v; break;
|
||||||
case 'fallback_start': $out['fallback_start'] = (int)$v; break;
|
case 'fallback_start': $out['fallback_start'] = (int)$v; break;
|
||||||
case 'handback_strikes': $out['handback_strikes'] = (int)$v; break;
|
case 'handback_strikes': $out['handback_strikes'] = (int)$v; break;
|
||||||
case 'tier2_started': $out['tier2_started'] = $v === 'true'; break;
|
case 'tier2_started': $out['tier2_started'] = $v === 'true'; break;
|
||||||
case 'tier3_started': $out['tier3_started'] = $v === 'true'; break;
|
case 'tier3_started': $out['tier3_started'] = $v === 'true'; break;
|
||||||
case 'tier4_started': $out['tier4_started'] = $v === 'true'; break;
|
case 'tier4_started': $out['tier4_started'] = $v === 'true'; break;
|
||||||
|
case 'partnership_suspended': $out['partnership_suspended'] = $v === 'true'; break;
|
||||||
|
case 'partner_lost_at': $out['partner_lost_at'] = (int)$v; break;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return $out;
|
return $out;
|
||||||
}
|
}
|
||||||
|
|
||||||
function vv_fb_local_state(): array {
|
function vv_fb_local_state(): array {
|
||||||
$path = '/boot/config/fallback_state.db';
|
$path = STATE_DIR . '/fallback_state.db';
|
||||||
return vv_fb_parse_state(file_exists($path) ? file_get_contents($path) : '');
|
return vv_fb_parse_state(file_exists($path) ? file_get_contents($path) : '');
|
||||||
}
|
}
|
||||||
|
|
||||||
function vv_fb_remote_state(string $ip, string $sshKey): array {
|
function vv_fb_remote_state(string $ip, string $sshKey): array {
|
||||||
$out = vv_pt_ssh($ip, $sshKey, 'cat /boot/config/fallback_state.db 2>/dev/null');
|
$out = vv_pt_ssh($ip, $sshKey, vv_remote_state_cmd('fallback_state.db'));
|
||||||
return vv_fb_parse_state($out);
|
return vv_fb_parse_state($out);
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -94,8 +98,13 @@ function vv_fb_all(): array {
|
|||||||
$currentHost = vv_detect_host();
|
$currentHost = vv_detect_host();
|
||||||
$hosts = vv_fb_known_hosts();
|
$hosts = vv_fb_known_hosts();
|
||||||
$tsPeers = vv_pt_ts_peers();
|
$tsPeers = vv_pt_ts_peers();
|
||||||
$handbackReq = (int)(vv_fb_scalar(vv_read_conf_raw('master.conf'), 'FALLBACK_HANDBACK_STRIKES') ?: 3);
|
$masterRaw = vv_read_conf_raw('master.conf');
|
||||||
$fbEnabled = vv_fb_scalar(vv_read_conf_raw('master.conf'), 'FALLBACK_ENABLED') === 'true';
|
$handbackReq = (int)(vv_fb_scalar($masterRaw, 'FALLBACK_HANDBACK_STRIKES') ?: 3);
|
||||||
|
$fbEnabled = vv_fb_scalar($masterRaw, 'FALLBACK_ENABLED') === 'true';
|
||||||
|
$ptEnabled = vv_fb_scalar($masterRaw, 'PARTNERSHIP_ENABLED') === 'true';
|
||||||
|
$rsyncEnabled = vv_fb_scalar($masterRaw, 'FALLBACK_RSYNC_ENABLED') !== 'false';
|
||||||
|
$checkInterval = (int)(vv_fb_scalar($masterRaw, 'FALLBACK_CHECK_INTERVAL') ?: 30);
|
||||||
|
$suspendAfter = (int)(vv_fb_scalar($masterRaw, 'FALLBACK_PARTNERSHIP_SUSPEND_AFTER') ?: 120);
|
||||||
|
|
||||||
// Read all host conf raws upfront
|
// Read all host conf raws upfront
|
||||||
$raws = [];
|
$raws = [];
|
||||||
@@ -162,9 +171,13 @@ function vv_fb_all(): array {
|
|||||||
}
|
}
|
||||||
|
|
||||||
return [
|
return [
|
||||||
'ts' => time(),
|
'ts' => time(),
|
||||||
'fb_enabled' => $fbEnabled,
|
'fb_enabled' => $fbEnabled,
|
||||||
'handback_req' => $handbackReq,
|
'partnership_enabled' => $ptEnabled,
|
||||||
'nodes' => $nodes,
|
'fb_rsync_enabled' => $rsyncEnabled,
|
||||||
|
'handback_req' => $handbackReq,
|
||||||
|
'check_interval' => $checkInterval,
|
||||||
|
'suspend_after' => $suspendAfter,
|
||||||
|
'nodes' => $nodes,
|
||||||
];
|
];
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -58,7 +58,7 @@ function vv_fallback_state(): array {
|
|||||||
$reqStrikes = (int)($vars['FALLBACK_HANDBACK_STRIKES'] ?? 3);
|
$reqStrikes = (int)($vars['FALLBACK_HANDBACK_STRIKES'] ?? 3);
|
||||||
$suspendAfter = (int)($vars['FALLBACK_PARTNERSHIP_SUSPEND_AFTER'] ?? 120);
|
$suspendAfter = (int)($vars['FALLBACK_PARTNERSHIP_SUSPEND_AFTER'] ?? 120);
|
||||||
|
|
||||||
$stateFile = '/tmp/fallback_state.db';
|
$stateFile = STATE_DIR . '/fallback_state.db';
|
||||||
if (!file_exists($stateFile)) {
|
if (!file_exists($stateFile)) {
|
||||||
return ['state' => 'UNKNOWN', 'enabled' => $enabled, 'check_interval' => $interval,
|
return ['state' => 'UNKNOWN', 'enabled' => $enabled, 'check_interval' => $interval,
|
||||||
'handback_strikes' => 0, 'handback_strikes_required' => $reqStrikes,
|
'handback_strikes' => 0, 'handback_strikes_required' => $reqStrikes,
|
||||||
@@ -164,7 +164,7 @@ function vv_watchdog_summary(): array {
|
|||||||
usort($restarts, fn($a, $b) => $b['ts'] - $a['ts']);
|
usort($restarts, fn($a, $b) => $b['ts'] - $a['ts']);
|
||||||
|
|
||||||
// Reboots (12 h)
|
// Reboots (12 h)
|
||||||
$rebootRaw = @file_get_contents('/boot/config/system_watchdog_reboots.db') ?: '';
|
$rebootRaw = @file_get_contents(STATE_DIR . '/system_watchdog_reboots.db') ?: '';
|
||||||
$rbootCutoff = time() - 43200;
|
$rbootCutoff = time() - 43200;
|
||||||
$reboots = 0;
|
$reboots = 0;
|
||||||
foreach (explode("\n", trim($rebootRaw)) as $line) {
|
foreach (explode("\n", trim($rebootRaw)) as $line) {
|
||||||
|
|||||||
@@ -10,7 +10,7 @@ require_once __DIR__ . '/common.php'; // vv_system_info(), vv_docker_containers(
|
|||||||
function vv_pt_config(): array {
|
function vv_pt_config(): array {
|
||||||
$v = vv_conf_vars();
|
$v = vv_conf_vars();
|
||||||
$offlineDays = null;
|
$offlineDays = null;
|
||||||
$odFile = '/boot/config/partnership_offline_days.db';
|
$odFile = STATE_DIR . '/partnership_offline_days.db';
|
||||||
if (file_exists($odFile)) {
|
if (file_exists($odFile)) {
|
||||||
$raw = trim(@file_get_contents($odFile) ?: '');
|
$raw = trim(@file_get_contents($odFile) ?: '');
|
||||||
if (is_numeric($raw)) $offlineDays = (int)$raw;
|
if (is_numeric($raw)) $offlineDays = (int)$raw;
|
||||||
@@ -209,12 +209,11 @@ function vv_pt_nodes(): array {
|
|||||||
|
|
||||||
// Fallback state
|
// Fallback state
|
||||||
$fbState = 'UNKNOWN';
|
$fbState = 'UNKNOWN';
|
||||||
$fbPath = '/boot/config/fallback_state.db';
|
|
||||||
if ($isMe) {
|
if ($isMe) {
|
||||||
$fb = vv_pt_read_db($fbPath);
|
$fb = vv_pt_read_db(STATE_DIR . '/fallback_state.db');
|
||||||
$fbState = $fb['state'] ?? 'UNKNOWN';
|
$fbState = $fb['state'] ?? 'UNKNOWN';
|
||||||
} elseif ($ts['online'] && $ts['ip'] && $mySshKey) {
|
} elseif ($ts['online'] && $ts['ip'] && $mySshKey) {
|
||||||
$out = vv_pt_ssh($ts['ip'], $mySshKey, 'cat /boot/config/fallback_state.db 2>/dev/null');
|
$out = vv_pt_ssh($ts['ip'], $mySshKey, vv_remote_state_cmd('fallback_state.db'));
|
||||||
if ($out) {
|
if ($out) {
|
||||||
$fb = [];
|
$fb = [];
|
||||||
foreach (explode("\n", $out) as $line) {
|
foreach (explode("\n", $out) as $line) {
|
||||||
@@ -226,7 +225,7 @@ function vv_pt_nodes(): array {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Partnership DB — local only (each server writes its own)
|
// Partnership DB — local only (each server writes its own)
|
||||||
$dbPath = "/boot/config/partnership_{$hostname}.db";
|
$dbPath = STATE_DIR . "/partnership_{$hostname}.db";
|
||||||
$ptDb = vv_pt_read_db($dbPath);
|
$ptDb = vv_pt_read_db($dbPath);
|
||||||
|
|
||||||
// System info
|
// System info
|
||||||
|
|||||||
@@ -127,14 +127,14 @@ function vv_wd_parse_network_state(string $raw): array {
|
|||||||
// ── Local state files ─────────────────────────────────────────────────────────
|
// ── Local state files ─────────────────────────────────────────────────────────
|
||||||
|
|
||||||
function vv_wd_local_states(string $restartLogPath): array {
|
function vv_wd_local_states(string $restartLogPath): array {
|
||||||
$rwRaw = @file_get_contents('/tmp/resource_watchdog_state.db') ?: '';
|
$rwRaw = @file_get_contents(STATE_DIR . '/resource_watchdog_state.db') ?: '';
|
||||||
$dockRaw = @file_get_contents('/tmp/container_watchdog_state.db') ?: '';
|
$dockRaw = @file_get_contents(STATE_DIR . '/container_watchdog_state.db') ?: '';
|
||||||
$skipRaw = @file_get_contents('/boot/config/system_watchdog_failed.db') ?: '';
|
$skipRaw = @file_get_contents(STATE_DIR . '/docker_watchdog_failed.db') ?: '';
|
||||||
$sysRaw = @file_get_contents('/tmp/system_watchdog_state.db') ?: '';
|
$sysRaw = @file_get_contents(STATE_DIR . '/system_watchdog_state.db') ?: '';
|
||||||
$rebootRaw = @file_get_contents('/boot/config/system_watchdog_reboots.db')?: '';
|
$rebootRaw = @file_get_contents(STATE_DIR . '/system_watchdog_reboots.db') ?: '';
|
||||||
$restartRaw= @file_get_contents($restartLogPath) ?: '';
|
$restartRaw= @file_get_contents($restartLogPath) ?: '';
|
||||||
$storRaw = @file_get_contents('/tmp/storage_watchdog_state.db') ?: '';
|
$storRaw = @file_get_contents(STATE_DIR . '/storage_watchdog_state.db') ?: '';
|
||||||
$netWdRaw = @file_get_contents('/tmp/network_watchdog_state.db') ?: '';
|
$netWdRaw = @file_get_contents(STATE_DIR . '/network_watchdog_state.db') ?: '';
|
||||||
|
|
||||||
$rw = vv_wd_parse_kv($rwRaw);
|
$rw = vv_wd_parse_kv($rwRaw);
|
||||||
$dock = vv_wd_parse_kv($dockRaw);
|
$dock = vv_wd_parse_kv($dockRaw);
|
||||||
|
|||||||
@@ -1,4 +1,5 @@
|
|||||||
<style>
|
<style>
|
||||||
|
/* ── Existing status styles ── */
|
||||||
.vv-fb-active { background:#1a1200;border:1px solid #5a3800;border-radius:6px;padding:12px 14px; }
|
.vv-fb-active { background:#1a1200;border:1px solid #5a3800;border-radius:6px;padding:12px 14px; }
|
||||||
.vv-fb-active-h { display:flex;align-items:baseline;gap:10px;margin-bottom:8px; }
|
.vv-fb-active-h { display:flex;align-items:baseline;gap:10px;margin-bottom:8px; }
|
||||||
.vv-fb-badge { font-size:11px;font-weight:bold;letter-spacing:.06em;padding:2px 7px;border-radius:3px;flex-shrink:0; }
|
.vv-fb-badge { font-size:11px;font-weight:bold;letter-spacing:.06em;padding:2px 7px;border-radius:3px;flex-shrink:0; }
|
||||||
@@ -6,6 +7,7 @@
|
|||||||
.vv-fb-badge.norm { background:#1a2a1a;color:#4caf50; }
|
.vv-fb-badge.norm { background:#1a2a1a;color:#4caf50; }
|
||||||
.vv-fb-badge.dark { background:#2a1a2a;color:#9c27b0; }
|
.vv-fb-badge.dark { background:#2a1a2a;color:#9c27b0; }
|
||||||
.vv-fb-badge.nonet{ background:#1a1a2a;color:#5c7cfa; }
|
.vv-fb-badge.nonet{ background:#1a1a2a;color:#5c7cfa; }
|
||||||
|
.vv-fb-badge.susp { background:#2a1a00;color:#888; }
|
||||||
.vv-fb-meta { display:flex;gap:18px;flex-wrap:wrap;margin-bottom:10px; }
|
.vv-fb-meta { display:flex;gap:18px;flex-wrap:wrap;margin-bottom:10px; }
|
||||||
.vv-fb-meta-item{ display:flex;flex-direction:column;gap:1px; }
|
.vv-fb-meta-item{ display:flex;flex-direction:column;gap:1px; }
|
||||||
.vv-fb-meta-val { font-size:17px;font-weight:bold;color:#ffb74d; }
|
.vv-fb-meta-val { font-size:17px;font-weight:bold;color:#ffb74d; }
|
||||||
@@ -31,17 +33,109 @@
|
|||||||
.vv-fb-state-dot { width:6px;height:6px;border-radius:50%;flex-shrink:0;margin-top:3px; }
|
.vv-fb-state-dot { width:6px;height:6px;border-radius:50%;flex-shrink:0;margin-top:3px; }
|
||||||
.vv-fb-sep { border:none;border-top:1px solid #222;margin:8px 0; }
|
.vv-fb-sep { border:none;border-top:1px solid #222;margin:8px 0; }
|
||||||
.vv-fb-disabled { grid-column:1/-1;color:#3a3a3a;font-size:12px;padding:20px 0;text-align:center; }
|
.vv-fb-disabled { grid-column:1/-1;color:#3a3a3a;font-size:12px;padding:20px 0;text-align:center; }
|
||||||
|
|
||||||
|
/* ── Controls + settings card ── */
|
||||||
|
.vv-fb-card { background:#161616;border:1px solid #2a2a2a;border-radius:6px;padding:14px 16px;margin-bottom:14px; }
|
||||||
|
.vv-fb-card-hdr { font-size:11px;font-weight:700;color:#666;text-transform:uppercase;letter-spacing:.07em;margin-bottom:12px; }
|
||||||
|
.vv-fb-ctrl-row { display:flex;justify-content:space-between;align-items:center;gap:12px;padding:5px 0; }
|
||||||
|
.vv-fb-ctrl-lbl { font-size:12px;color:#888; }
|
||||||
|
.vv-fb-ctrl-sub { font-size:10px;color:#3a3a3a;margin-top:2px; }
|
||||||
|
.vv-fb-tog { width:32px;height:18px;border-radius:9px;background:#222;border:1px solid #333;
|
||||||
|
position:relative;transition:background .15s,border-color .15s;flex-shrink:0;cursor:pointer; }
|
||||||
|
.vv-fb-tog.on { background:#1a3a1a;border-color:#2d5a2d; }
|
||||||
|
.vv-fb-tog::after { content:'';position:absolute;top:2px;left:2px;width:12px;height:12px;
|
||||||
|
border-radius:50%;background:#555;transition:left .15s,background .15s; }
|
||||||
|
.vv-fb-tog.on::after { left:16px;background:#4caf50; }
|
||||||
|
.vv-fb-set-row { display:flex;justify-content:space-between;align-items:center;padding:5px 0; }
|
||||||
|
.vv-fb-set-lbl { font-size:11px;color:#555; }
|
||||||
|
.vv-fb-set-inp { background:#0d0d0d;border:1px solid #2a2a2a;border-radius:3px;color:#888;
|
||||||
|
font-size:12px;padding:4px 8px;outline:none;width:72px;font-family:monospace;
|
||||||
|
text-align:right;box-sizing:border-box; }
|
||||||
|
.vv-fb-set-inp:focus { border-color:#444; }
|
||||||
|
.vv-fb-set-unit { font-size:10px;color:#3a3a3a;min-width:56px; }
|
||||||
|
.vv-fb-save-btn { background:#1a1a1a;border:1px solid #333;color:#888;font-size:11px;
|
||||||
|
padding:5px 14px;border-radius:3px;cursor:pointer; }
|
||||||
|
.vv-fb-save-btn:hover { border-color:#555;color:#ccc; }
|
||||||
|
.vv-fb-save-btn:disabled { opacity:.4;cursor:default; }
|
||||||
</style>
|
</style>
|
||||||
|
|
||||||
|
<!-- Controls card -->
|
||||||
|
<div class="vv-fb-card">
|
||||||
|
<div class="vv-fb-card-hdr">Controls</div>
|
||||||
|
|
||||||
|
<div class="vv-fb-ctrl-row">
|
||||||
|
<div>
|
||||||
|
<div class="vv-fb-ctrl-lbl">Partnership</div>
|
||||||
|
<div class="vv-fb-ctrl-sub">Master gate — disabling stops all cross-server operations</div>
|
||||||
|
</div>
|
||||||
|
<div class="vv-fb-tog" id="vv-fb-pt-tog" onclick="vvFbToggle(this,'PARTNERSHIP_ENABLED')"></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr class="vv-fb-sep">
|
||||||
|
|
||||||
|
<div class="vv-fb-ctrl-row">
|
||||||
|
<div>
|
||||||
|
<div class="vv-fb-ctrl-lbl">Fallback</div>
|
||||||
|
<div class="vv-fb-ctrl-sub">Mutual container failover between nodes</div>
|
||||||
|
</div>
|
||||||
|
<div class="vv-fb-tog" id="vv-fb-en-tog" onclick="vvFbToggle(this,'FALLBACK_ENABLED')"></div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr class="vv-fb-sep">
|
||||||
|
|
||||||
|
<div class="vv-fb-ctrl-row">
|
||||||
|
<div>
|
||||||
|
<div class="vv-fb-ctrl-lbl">Rsync on handback</div>
|
||||||
|
<div class="vv-fb-ctrl-sub">Writeback rsync when the covered host recovers and containers return</div>
|
||||||
|
</div>
|
||||||
|
<div class="vv-fb-tog" id="vv-fb-rsync-tog" onclick="vvFbToggle(this,'FALLBACK_RSYNC_ENABLED')"></div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Status section -->
|
||||||
<div style="display:flex;justify-content:space-between;align-items:center;margin-bottom:12px;padding:0 2px;">
|
<div style="display:flex;justify-content:space-between;align-items:center;margin-bottom:12px;padding:0 2px;">
|
||||||
<span style="font-size:13px;font-weight:bold;color:#888;text-transform:uppercase;letter-spacing:.06em;">FallBack</span>
|
<span style="font-size:13px;font-weight:bold;color:#888;text-transform:uppercase;letter-spacing:.06em;">Status</span>
|
||||||
<span style="font-size:11px;color:#3a3a3a;" id="vv-fb-ts"></span>
|
<span style="font-size:11px;color:#3a3a3a;" id="vv-fb-ts"></span>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
<div id="vv-fb-grid" style="display:grid;grid-template-columns:repeat(8,1fr);gap:12px;">
|
<div id="vv-fb-grid" style="display:grid;grid-template-columns:repeat(8,1fr);gap:12px;margin-bottom:14px;">
|
||||||
<div style="grid-column:1/-1;color:#444;font-size:12px;padding:16px 0;text-align:center;">Loading…</div>
|
<div style="grid-column:1/-1;color:#444;font-size:12px;padding:16px 0;text-align:center;">Loading…</div>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
|
<!-- Settings card -->
|
||||||
|
<div class="vv-fb-card">
|
||||||
|
<div class="vv-fb-card-hdr">Settings</div>
|
||||||
|
|
||||||
|
<div class="vv-fb-set-row">
|
||||||
|
<span class="vv-fb-set-lbl">Check interval</span>
|
||||||
|
<div style="display:flex;align-items:center;gap:8px;">
|
||||||
|
<input class="vv-fb-set-inp" id="vv-fb-interval" type="number" min="5" max="300">
|
||||||
|
<span class="vv-fb-set-unit">seconds</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="vv-fb-set-row">
|
||||||
|
<span class="vv-fb-set-lbl">Handback strikes</span>
|
||||||
|
<div style="display:flex;align-items:center;gap:8px;">
|
||||||
|
<input class="vv-fb-set-inp" id="vv-fb-strikes" type="number" min="1" max="20">
|
||||||
|
<span class="vv-fb-set-unit">consecutive</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="vv-fb-set-row">
|
||||||
|
<span class="vv-fb-set-lbl">Partnership suspend after</span>
|
||||||
|
<div style="display:flex;align-items:center;gap:8px;">
|
||||||
|
<input class="vv-fb-set-inp" id="vv-fb-suspend" type="number" min="0" max="1440">
|
||||||
|
<span class="vv-fb-set-unit">minutes</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div style="display:flex;justify-content:flex-end;align-items:center;gap:10px;margin-top:10px;padding-top:8px;border-top:1px solid #1e1e1e;">
|
||||||
|
<span id="vv-fb-set-fb" style="font-size:11px;"></span>
|
||||||
|
<button class="vv-fb-save-btn" id="vv-fb-save-btn" onclick="vvFbSaveSettings()">Save</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
<script>
|
<script>
|
||||||
(function() {
|
(function() {
|
||||||
|
|
||||||
@@ -78,6 +172,7 @@ function _stateBadge(st) {
|
|||||||
OFFLINE: ['dark', 'OFFLINE'],
|
OFFLINE: ['dark', 'OFFLINE'],
|
||||||
UNREACHABLE: ['dark', 'UNREACHABLE'],
|
UNREACHABLE: ['dark', 'UNREACHABLE'],
|
||||||
UNKNOWN: ['dark', 'UNKNOWN'],
|
UNKNOWN: ['dark', 'UNKNOWN'],
|
||||||
|
SUSPENDED: ['susp', 'SUSPENDED'],
|
||||||
};
|
};
|
||||||
const [cls, label] = map[st] || ['dark', st];
|
const [cls, label] = map[st] || ['dark', st];
|
||||||
return `<span class="vv-fb-badge ${cls}">${label}</span>`;
|
return `<span class="vv-fb-badge ${cls}">${label}</span>`;
|
||||||
@@ -87,7 +182,7 @@ function _stateDot(st) {
|
|||||||
const col = {
|
const col = {
|
||||||
NORMAL:'#4caf50', FALLBACK:'#ffb74d',
|
NORMAL:'#4caf50', FALLBACK:'#ffb74d',
|
||||||
NO_INTERNET:'#5c7cfa', DARK:'#9c27b0',
|
NO_INTERNET:'#5c7cfa', DARK:'#9c27b0',
|
||||||
OFFLINE:'#555', UNREACHABLE:'#555', UNKNOWN:'#333',
|
OFFLINE:'#555', UNREACHABLE:'#555', UNKNOWN:'#333', SUSPENDED:'#444',
|
||||||
}[st] || '#333';
|
}[st] || '#333';
|
||||||
return `<span class="vv-fb-state-dot" style="background:${col}"></span>`;
|
return `<span class="vv-fb-state-dot" style="background:${col}"></span>`;
|
||||||
}
|
}
|
||||||
@@ -102,7 +197,6 @@ function _activeCard(nodes, handbackReq) {
|
|||||||
const cov = covering.covers;
|
const cov = covering.covers;
|
||||||
const covered = cov ? cov.hostname : '?';
|
const covered = cov ? cov.hostname : '?';
|
||||||
|
|
||||||
// All containers that should be running at current tier
|
|
||||||
let expected = [...(cov?.tier1 || [])];
|
let expected = [...(cov?.tier1 || [])];
|
||||||
if (tier >= 2) expected = expected.concat(cov?.tier2 || []);
|
if (tier >= 2) expected = expected.concat(cov?.tier2 || []);
|
||||||
if (tier >= 3) expected = expected.concat(cov?.tier3 || []);
|
if (tier >= 3) expected = expected.concat(cov?.tier3 || []);
|
||||||
@@ -169,9 +263,19 @@ function _tierSection(tiers, activeTier, delays) {
|
|||||||
}).join('');
|
}).join('');
|
||||||
}
|
}
|
||||||
|
|
||||||
function _nodeCard(node) {
|
function _ptStatus(st) {
|
||||||
|
if (!st) return '';
|
||||||
|
if (st.partnership_suspended) return _stateBadge('SUSPENDED');
|
||||||
|
if (st.partner_lost_at && st.partner_lost_at > 0) {
|
||||||
|
const minGone = Math.floor((Date.now() / 1000 - st.partner_lost_at) / 60);
|
||||||
|
return `<span class="vv-fb-badge susp">GRACE ${minGone}m</span>`;
|
||||||
|
}
|
||||||
|
return '';
|
||||||
|
}
|
||||||
|
|
||||||
|
function _nodeCard(node, suspendAfter) {
|
||||||
const st = node.state || {};
|
const st = node.state || {};
|
||||||
const state = st.state || 'UNKNOWN';
|
const state = st.partnership_suspended ? 'SUSPENDED' : (st.state || 'UNKNOWN');
|
||||||
const cov = node.covers;
|
const cov = node.covers;
|
||||||
const active = _activeTier(state === 'FALLBACK' ? st : null);
|
const active = _activeTier(state === 'FALLBACK' ? st : null);
|
||||||
|
|
||||||
@@ -183,6 +287,8 @@ function _nodeCard(node) {
|
|||||||
? _tierSection(cov, active, cov.delays)
|
? _tierSection(cov, active, cov.delays)
|
||||||
: '<div style="color:#3a3a3a;font-size:11px;">No coverage configured</div>';
|
: '<div style="color:#3a3a3a;font-size:11px;">No coverage configured</div>';
|
||||||
|
|
||||||
|
const ptBadge = node.is_me ? _ptStatus(st) : '';
|
||||||
|
|
||||||
return `<div class="vv-card vv-fb-node">
|
return `<div class="vv-card vv-fb-node">
|
||||||
<div class="vv-fb-node-h">
|
<div class="vv-fb-node-h">
|
||||||
${_stateDot(state)}
|
${_stateDot(state)}
|
||||||
@@ -190,6 +296,7 @@ function _nodeCard(node) {
|
|||||||
<span class="vv-fb-node-nm">${node.hostname}</span>
|
<span class="vv-fb-node-nm">${node.hostname}</span>
|
||||||
${covTarget}
|
${covTarget}
|
||||||
<span style="flex:1"></span>
|
<span style="flex:1"></span>
|
||||||
|
${ptBadge}
|
||||||
${_stateBadge(state)}
|
${_stateBadge(state)}
|
||||||
</div>
|
</div>
|
||||||
<hr class="vv-fb-sep">
|
<hr class="vv-fb-sep">
|
||||||
@@ -197,30 +304,47 @@ function _nodeCard(node) {
|
|||||||
</div>`;
|
</div>`;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function _setToggles(data) {
|
||||||
|
const pairs = [
|
||||||
|
['vv-fb-pt-tog', !!data.partnership_enabled],
|
||||||
|
['vv-fb-en-tog', !!data.fb_enabled],
|
||||||
|
['vv-fb-rsync-tog', !!data.fb_rsync_enabled],
|
||||||
|
];
|
||||||
|
pairs.forEach(([id, on]) => {
|
||||||
|
const el = document.getElementById(id);
|
||||||
|
if (el) el.classList.toggle('on', on);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function _setInputs(data) {
|
||||||
|
const fields = [
|
||||||
|
['vv-fb-interval', data.check_interval ?? 30],
|
||||||
|
['vv-fb-strikes', data.handback_req ?? 3],
|
||||||
|
['vv-fb-suspend', data.suspend_after ?? 120],
|
||||||
|
];
|
||||||
|
fields.forEach(([id, val]) => {
|
||||||
|
const el = document.getElementById(id);
|
||||||
|
if (el && el !== document.activeElement) el.value = val;
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
function _render(data) {
|
function _render(data) {
|
||||||
if (!data.fb_enabled) {
|
_setToggles(data);
|
||||||
document.getElementById('vv-fb-grid').innerHTML =
|
_setInputs(data);
|
||||||
'<div class="vv-fb-disabled">FALLBACK_ENABLED=false — fallback monitoring is disabled</div>';
|
|
||||||
return;
|
const grid = document.getElementById('vv-fb-grid');
|
||||||
|
|
||||||
|
if (!data.partnership_enabled) {
|
||||||
|
grid.innerHTML = '<div class="vv-fb-disabled">PARTNERSHIP_ENABLED=false — all cross-server operations disabled</div>';
|
||||||
|
} else if (!data.fb_enabled) {
|
||||||
|
grid.innerHTML = '<div class="vv-fb-disabled">FALLBACK_ENABLED=false — fallback monitoring is disabled</div>';
|
||||||
|
} else {
|
||||||
|
const nodes = data.nodes || [];
|
||||||
|
let html = _activeCard(nodes, data.handback_req || 3);
|
||||||
|
for (const node of nodes) html += _nodeCard(node, data.suspend_after || 120);
|
||||||
|
grid.innerHTML = html || '<div class="vv-fb-disabled">No nodes configured.</div>';
|
||||||
}
|
}
|
||||||
|
|
||||||
const nodes = data.nodes || [];
|
|
||||||
let html = '';
|
|
||||||
|
|
||||||
// Top: active fallback card (if any)
|
|
||||||
html += _activeCard(nodes, data.handback_req || 3);
|
|
||||||
|
|
||||||
// Per-node cards
|
|
||||||
for (const node of nodes) {
|
|
||||||
html += _nodeCard(node);
|
|
||||||
}
|
|
||||||
|
|
||||||
if (!html) {
|
|
||||||
html = '<div class="vv-fb-disabled">No nodes configured.</div>';
|
|
||||||
}
|
|
||||||
|
|
||||||
document.getElementById('vv-fb-grid').innerHTML = html;
|
|
||||||
|
|
||||||
const ts = data.ts
|
const ts = data.ts
|
||||||
? new Date(data.ts * 1000).toLocaleString([], {
|
? new Date(data.ts * 1000).toLocaleString([], {
|
||||||
month:'numeric', day:'numeric', year:'numeric',
|
month:'numeric', day:'numeric', year:'numeric',
|
||||||
@@ -239,6 +363,55 @@ function vvFbLoad() {
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
window.vvFbToggle = function(track, key) {
|
||||||
|
const on = !track.classList.contains('on');
|
||||||
|
track.classList.toggle('on', on);
|
||||||
|
const fd = new FormData();
|
||||||
|
fd.append('id', 'fallback');
|
||||||
|
fd.append('changes', JSON.stringify([{ file: 'master.conf', key, value: on ? 'true' : 'false', type: 'scalar' }]));
|
||||||
|
fetch('/plugins/varaverk/api/confform.php', { method: 'POST', body: fd })
|
||||||
|
.then(r => r.json())
|
||||||
|
.then(d => { if (!d.ok) track.classList.toggle('on', !on); })
|
||||||
|
.catch(() => track.classList.toggle('on', !on));
|
||||||
|
};
|
||||||
|
|
||||||
|
window.vvFbSaveSettings = function() {
|
||||||
|
const ivEl = document.getElementById('vv-fb-interval');
|
||||||
|
const stEl = document.getElementById('vv-fb-strikes');
|
||||||
|
const suEl = document.getElementById('vv-fb-suspend');
|
||||||
|
const fb = document.getElementById('vv-fb-set-fb');
|
||||||
|
const btn = document.getElementById('vv-fb-save-btn');
|
||||||
|
|
||||||
|
const interval = parseInt(ivEl.value, 10);
|
||||||
|
const strikes = parseInt(stEl.value, 10);
|
||||||
|
const suspend = parseInt(suEl.value, 10);
|
||||||
|
|
||||||
|
if ([interval, strikes, suspend].some(n => isNaN(n) || n < 0)) {
|
||||||
|
fb.style.color = '#ef5350'; fb.textContent = 'Invalid values'; return;
|
||||||
|
}
|
||||||
|
|
||||||
|
btn.disabled = true; btn.textContent = 'Saving…'; fb.textContent = '';
|
||||||
|
const fd = new FormData();
|
||||||
|
fd.append('id', 'fallback');
|
||||||
|
fd.append('changes', JSON.stringify([
|
||||||
|
{ file: 'master.conf', key: 'FALLBACK_CHECK_INTERVAL', value: String(interval), type: 'scalar' },
|
||||||
|
{ file: 'master.conf', key: 'FALLBACK_HANDBACK_STRIKES', value: String(strikes), type: 'scalar' },
|
||||||
|
{ file: 'master.conf', key: 'FALLBACK_PARTNERSHIP_SUSPEND_AFTER', value: String(suspend), type: 'scalar' },
|
||||||
|
]));
|
||||||
|
fetch('/plugins/varaverk/api/confform.php', { method: 'POST', body: fd })
|
||||||
|
.then(r => r.json())
|
||||||
|
.then(d => {
|
||||||
|
btn.disabled = false; btn.textContent = 'Save';
|
||||||
|
fb.style.color = d.ok ? '#4caf50' : '#ef5350';
|
||||||
|
fb.textContent = d.ok ? 'Saved ✓' : (d.error || 'Failed');
|
||||||
|
if (d.ok) setTimeout(() => { fb.textContent = ''; }, 3000);
|
||||||
|
})
|
||||||
|
.catch(() => {
|
||||||
|
btn.disabled = false; btn.textContent = 'Save';
|
||||||
|
fb.style.color = '#ef5350'; fb.textContent = 'Request failed';
|
||||||
|
});
|
||||||
|
};
|
||||||
|
|
||||||
vvFbLoad();
|
vvFbLoad();
|
||||||
setInterval(vvFbLoad, 30000);
|
setInterval(vvFbLoad, 30000);
|
||||||
|
|
||||||
|
|||||||
@@ -715,8 +715,8 @@ Change `--dry-run` to nothing and run during a maintenance window.
|
|||||||
> forced kill of fallback.sh:
|
> forced kill of fallback.sh:
|
||||||
> ```bash
|
> ```bash
|
||||||
> # Verify first — right containers on right server, DDNS correct, fallback.sh not running
|
> # Verify first — right containers on right server, DDNS correct, fallback.sh not running
|
||||||
> cat /boot/config/fallback_state.db
|
> cat "$STATE_DIR/fallback_state.db"
|
||||||
> echo "state=NORMAL" > /boot/config/fallback_state.db
|
> echo "state=NORMAL" > "$STATE_DIR/fallback_state.db"
|
||||||
> ```
|
> ```
|
||||||
> Resets state only — does NOT start or stop any containers.
|
> Resets state only — does NOT start or stop any containers.
|
||||||
|
|
||||||
|
|||||||
@@ -621,7 +621,7 @@ SLEEP=60 # seconds between retries
|
|||||||
ROOTFS_WARN_PCT=75 # abort if remote rootfs above this %
|
ROOTFS_WARN_PCT=75 # abort if remote rootfs above this %
|
||||||
|
|
||||||
# Shared with Monitors/
|
# Shared with Monitors/
|
||||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db"
|
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db"
|
||||||
BANDWIDTH_LOG_RETENTION=90
|
BANDWIDTH_LOG_RETENTION=90
|
||||||
BANDWIDTH_WARN_GB=50
|
BANDWIDTH_WARN_GB=50
|
||||||
|
|
||||||
|
|||||||
@@ -4,23 +4,24 @@ Configuration reference, operational procedures, and troubleshooting for
|
|||||||
system-level scripts. Read the ARRAY_START_SCRIPTS order section before
|
system-level scripts. Read the ARRAY_START_SCRIPTS order section before
|
||||||
adding or reordering scripts at array start.
|
adding or reordering scripts at array start.
|
||||||
|
|
||||||
> **Watchdog scripts have moved.** `stability_watchdog.sh` and `resource_watchdog.sh`
|
> **Platform-specific scripts** (`webgui_watchdog.sh`, `php_fpm_max_children.sh`,
|
||||||
> now live in `Watchdogs/`. Their configuration reference and troubleshooting
|
> `mover_stop.sh`, `user_scripts_stop.sh`) and their configuration live in
|
||||||
> procedures are in `Watchdogs/Manual-Watchdogs.md`.
|
> `Plugin/unraid/System_Essentials/`. Watchdog configuration (`stability_watchdog.sh`,
|
||||||
|
> `resource_watchdog.sh`, `docker_watchdog.sh`, `System/storage_watchdog.sh`) lives
|
||||||
|
> in `Watchdogs/Manual-Watchdogs.md`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ━━━ CONTENTS ━━━
|
## ━━━ CONTENTS ━━━
|
||||||
|
|
||||||
- [ARRAY_START_SCRIPTS Order](#array_start_scripts-order)
|
- [ARRAY_START_SCRIPTS / ARRAY_STOP_SCRIPTS Order](#array_start_scripts--array_stop_scripts-order)
|
||||||
- [webgui_watchdog.sh](#webgui_watchdogsh)
|
- [conf_sync.sh](#conf_syncsh)
|
||||||
|
- [conf_cache_save.sh](#conf_cache_savesh)
|
||||||
|
- [conf_cache_restore.sh](#conf_cache_restoresh)
|
||||||
- [inotify_tuning.sh](#inotify_tuningsh)
|
- [inotify_tuning.sh](#inotify_tuningsh)
|
||||||
- [php_fpm_max_children.sh](#php_fpm_max_childrensh)
|
|
||||||
- [docker_syslog_filter.sh](#docker_syslog_filtersh)
|
- [docker_syslog_filter.sh](#docker_syslog_filtersh)
|
||||||
- [clear_logs.sh](#clear_logssh)
|
- [clear_logs.sh](#clear_logssh)
|
||||||
- [mover_stop.sh](#mover_stopsh)
|
|
||||||
- [rsync_stop.sh](#rsync_stopsh)
|
- [rsync_stop.sh](#rsync_stopsh)
|
||||||
- [user_scripts_stop.sh](#user_scripts_stopsh)
|
|
||||||
- [server_reboot.sh](#server_rebootsh)
|
- [server_reboot.sh](#server_rebootsh)
|
||||||
- [Full Configuration Reference](#full-configuration-reference)
|
- [Full Configuration Reference](#full-configuration-reference)
|
||||||
- [Troubleshooting](#troubleshooting)
|
- [Troubleshooting](#troubleshooting)
|
||||||
@@ -32,109 +33,155 @@ adding or reordering scripts at array start.
|
|||||||
All scripts use a two-tier output model: `echo` lines are always visible; `log`
|
All scripts use a two-tier output model: `echo` lines are always visible; `log`
|
||||||
lines only appear when `--log` is passed.
|
lines only appear when `--log` is passed.
|
||||||
|
|
||||||
**Daemon scripts** (`webgui_watchdog.sh`): run on every cycle. Without `--log`, only
|
**One-shot scripts** (`conf_sync.sh`, `conf_cache_save.sh`, `conf_cache_restore.sh`,
|
||||||
state transitions, warnings, errors, and the clean-cycle conclusion line are visible.
|
`clear_logs.sh`, `docker_syslog_filter.sh`, `inotify_tuning.sh`, `rsync_stop.sh`,
|
||||||
Per-check detail suppressed.
|
`server_reboot.sh`): without `--log`, section headers, per-step results, and the
|
||||||
|
|
||||||
**One-shot scripts** (`clear_logs.sh`, `docker_syslog_filter.sh`, `inotify_tuning.sh`,
|
|
||||||
`mover_stop.sh`, `php_fpm_max_children.sh`, `rsync_stop.sh`, `server_reboot.sh`,
|
|
||||||
`user_scripts_stop.sh`): without `--log`, section headers, per-step results, and the
|
|
||||||
final summary are visible. Per-item detail suppressed.
|
final summary are visible. Per-item detail suppressed.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ARRAY_START_SCRIPTS Order
|
## ARRAY_START_SCRIPTS / ARRAY_STOP_SCRIPTS Order
|
||||||
|
|
||||||
> **The order of scripts in ARRAY_START_SCRIPTS matters for three of these
|
> **Order matters.** Scripts that set up conditions other scripts depend on must run
|
||||||
> scripts.** Getting it wrong causes subtle failures that don't show up
|
> first. Getting the order wrong causes subtle failures.
|
||||||
> immediately.
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# master.conf
|
# master.conf (current order — from ARRAY_START_SCRIPTS)
|
||||||
ARRAY_START_SCRIPTS=(
|
ARRAY_START_SCRIPTS=(
|
||||||
"inotify_tuning.sh" # 1 — FIRST: kernel limits must be set before
|
"Plugin/unraid/System_Essentials/unraid_api_key_renew.sh" # plugin API key — ephemeral registry
|
||||||
# any container starts. Containers inherit
|
"System_Essentials/conf_sync.sh" # 1st cross-server op — pull partner confs into RAM cache
|
||||||
# inotify limits at launch, not dynamically.
|
"System_Essentials/conf_cache_restore.sh" # fill gaps if partner was down at boot
|
||||||
"docker_syslog_filter.sh" # 2 — SECOND: before any veth interfaces are
|
"Transcodes/ramdisk_setup.sh" # ramdisk before Emby starts
|
||||||
# created. If a container starts first, its
|
"System_Essentials/docker_syslog_filter.sh" # before veth interfaces are created
|
||||||
# veth creation is already in syslog.
|
"Plugin/unraid/System_Essentials/php_fpm_max_children.sh" # WebGUI tuning
|
||||||
"php_fpm_max_children.sh" # 3 — before WebGUI is under load
|
"System_Essentials/inotify_tuning.sh" # before docker_network_connect and continuous scripts
|
||||||
"ramdisk_setup.sh" # (from Transcodes/) before Emby starts
|
"Docker_Essentials/docker_network_connect.sh" # ensure networks + container connections
|
||||||
...
|
"Tools/claude_startup.sh" # persist Claude data to appdata; re-symlink on boot
|
||||||
|
"Media/start_webhook_listener.sh" # arr upgrade webhook — continuous
|
||||||
|
"Fallback/fallback.sh" # mutual failover — continuous
|
||||||
|
)
|
||||||
|
|
||||||
|
ARRAY_STOP_SCRIPTS=(
|
||||||
|
"System_Essentials/conf_cache_save.sh" # FIRST: snapshot RAM cache while fresh
|
||||||
|
"Plugin/unraid/System_Essentials/user_scripts_stop.sh"
|
||||||
|
"Fallback/fallback.sh --stop"
|
||||||
|
"System_Essentials/rsync_stop.sh --rsync-only"
|
||||||
|
"Plugin/unraid/System_Essentials/mover_stop.sh"
|
||||||
|
"Docker_Essentials/docker_container_stop.sh"
|
||||||
)
|
)
|
||||||
# Watchdogs are NOT in ARRAY_START_SCRIPTS — they run every 15 minutes via
|
|
||||||
# Orchestrators/watchdog_orchestrator.sh (separate cron entry).
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Why inotify FIRST: If Code-Server starts before limits are raised, it inherits
|
Why conf_sync before conf_cache_restore: the sync runs first to get the freshest
|
||||||
the old low limits. The limits are kernel-wide — a restart of Code-Server picks
|
possible partner conf. The restore only fills in for confs that the sync couldn't
|
||||||
up the new values, but it's a manual step. Avoid by running inotify_tuning.sh first.
|
fetch — it is a fallback, not the primary path.
|
||||||
|
|
||||||
Why docker_syslog_filter SECOND: The filter must be in place before any container
|
Why docker_syslog_filter before inotify: the filter must be in place before any
|
||||||
starts creating veth interfaces. The first container start after array start
|
container creates veth interfaces. inotify must be set before docker_network_connect
|
||||||
generates veth messages — these will appear in syslog if the filter isn't active.
|
and the continuous scripts (start_webhook_listener, fallback), which are the first
|
||||||
|
steps that actually touch containers.
|
||||||
|
|
||||||
|
Why conf_cache_save is FIRST in ARRAY_STOP_SCRIPTS: the RAM cache at
|
||||||
|
`/tmp/.cache/vv/d/` is wiped on reboot. Saving it must happen before anything
|
||||||
|
else shuts down or changes state.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## webgui_watchdog.sh
|
## conf_sync.sh
|
||||||
|
|
||||||
### Escalation Logic
|
### Modes
|
||||||
|
|
||||||
```
|
```
|
||||||
curl $WEBGUI_URL → 200 OK → exit 0 (silent)
|
conf_sync.sh Full sync: pull from all partners + push to all partners
|
||||||
|
conf_sync.sh --push-only Push own conf to all partners (fast — for conf-save hook)
|
||||||
Not responding:
|
conf_sync.sh --pull-only Pull partner confs only (intermediate orchestrator)
|
||||||
1. platform_restart_service nginx
|
conf_sync.sh --dry-run Show what would happen, no changes
|
||||||
wait WEBGUI_NGINX_WAIT (15s) → recheck
|
conf_sync.sh --log Verbose output
|
||||||
→ recovered: notify, exit 0
|
|
||||||
|
|
||||||
2. platform_restart_service php-fpm
|
|
||||||
wait WEBGUI_PHP_WAIT (10s) → recheck
|
|
||||||
→ recovered: notify, exit 0
|
|
||||||
|
|
||||||
3. platform_restart_service emhttp
|
|
||||||
wait WEBGUI_EMHTTP_WAIT (30s) → recheck
|
|
||||||
→ recovered: notify, exit 0
|
|
||||||
|
|
||||||
All three failed → notify warning, exit 1
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Configuration
|
### What It Syncs
|
||||||
|
|
||||||
|
- **Pull**: reads the partner's `Configurations/${partner_id}.conf` from their disk
|
||||||
|
via SCP → writes to local `/tmp/.cache/vv/d/${partner_id}.conf`
|
||||||
|
- **Push**: sends own `Configurations/${my_id}.conf` to partner's
|
||||||
|
`/tmp/.cache/vv/d/${my_id}.conf` via SCP
|
||||||
|
- **Own conf in local cache**: copies own conf to `/tmp/.cache/vv/d/${my_id}.conf`
|
||||||
|
on full sync (so the cache has a complete snapshot of all confs)
|
||||||
|
|
||||||
|
Only partner confs are sourced from cache — `load_config.sh` always reads own conf
|
||||||
|
from disk to avoid sourcing a stale cached copy.
|
||||||
|
|
||||||
|
### Remote SCRIPTS_DIR Resolution
|
||||||
|
|
||||||
|
The pull path reads the partner's `/boot/config/plugins/varaverk/varaverk.cfg` to
|
||||||
|
find their actual `SCRIPTS_DIR` before building the SCP path. This handles the case
|
||||||
|
where the partner is in appdata storage mode and their conf is at
|
||||||
|
`/mnt/user/appdata/Varaverk/Configurations/` rather than the internal path.
|
||||||
|
|
||||||
|
### PARTNERSHIP_ENABLED Gate
|
||||||
|
|
||||||
|
conf_sync.sh calls `require_partnership` — if `PARTNERSHIP_ENABLED=false`, it exits
|
||||||
|
silently with 0. The conf cache will be empty for partner confs while partnership
|
||||||
|
is disabled.
|
||||||
|
|
||||||
|
### If Conf Pull Fails at Boot
|
||||||
|
|
||||||
|
If the partner is unreachable, `conf_sync.sh` logs a warning and exits 1.
|
||||||
|
`conf_cache_restore.sh` then runs (next in ARRAY_START_SCRIPTS) and loads the
|
||||||
|
persistent backup from `$PERSISTENT_CONF_CACHE` if available.
|
||||||
|
|
||||||
|
A notification fires if any partner fails — check partner reachability via Tailscale.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## conf_cache_save.sh
|
||||||
|
|
||||||
|
### What It Does
|
||||||
|
|
||||||
|
At array stop, copies all partner confs from `/tmp/.cache/vv/d/` to
|
||||||
|
`$PERSISTENT_CONF_CACHE`. Own conf is skipped (always on disk). The backup survives
|
||||||
|
the reboot and is used by `conf_cache_restore.sh` at next array start if the sync
|
||||||
|
can't reach the partner.
|
||||||
|
|
||||||
|
`conf_cache_watchdog.sh` (in `Watchdogs/System/`) refreshes this backup every
|
||||||
|
15 minutes while the partner is offline — keeping it current even during extended
|
||||||
|
outages.
|
||||||
|
|
||||||
|
### PARTNERSHIP_ENABLED Gate
|
||||||
|
|
||||||
|
Exits silently when `PARTNERSHIP_ENABLED=false`. No backup is written.
|
||||||
|
|
||||||
|
### Usage
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
WEBGUI_URL="http://localhost" # URL to check
|
conf_cache_save.sh # save partner confs from RAM cache (called at array stop)
|
||||||
WEBGUI_TIMEOUT=5 # curl timeout in seconds
|
conf_cache_save.sh --dry-run # show what would be saved
|
||||||
WEBGUI_NGINX_WAIT=15 # seconds after nginx restart before recheck
|
conf_cache_save.sh --log # verbose output
|
||||||
WEBGUI_PHP_WAIT=10 # seconds after php-fpm restart before recheck
|
|
||||||
WEBGUI_EMHTTP_WAIT=30 # seconds after emhttp restart before recheck
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### WebGUI Frozen — Manual Recovery
|
---
|
||||||
|
|
||||||
|
## conf_cache_restore.sh
|
||||||
|
|
||||||
|
### What It Does
|
||||||
|
|
||||||
|
At array start (after `conf_sync.sh`), checks which partner confs are missing from
|
||||||
|
`/tmp/.cache/vv/d/`. For each missing conf, loads it from `$PERSISTENT_CONF_CACHE`
|
||||||
|
if a backup exists there.
|
||||||
|
|
||||||
|
Always removes the persistent backup when done — whether used or not. On a normal
|
||||||
|
reboot where the partner was reachable, the sync already populated the cache and the
|
||||||
|
backup is cleaned up without being used.
|
||||||
|
|
||||||
|
### PARTNERSHIP_ENABLED Gate
|
||||||
|
|
||||||
|
Exits silently when `PARTNERSHIP_ENABLED=false`.
|
||||||
|
|
||||||
|
### Usage
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Check which services are running:
|
conf_cache_restore.sh # restore missing confs from backup (called at array start)
|
||||||
webgui_watchdog.sh --status
|
conf_cache_restore.sh --dry-run # show what would be restored
|
||||||
|
conf_cache_restore.sh --log # verbose output
|
||||||
# Try manual restart sequence (mirrors what the script calls):
|
|
||||||
# Source the ecosystem first to get platform functions:
|
|
||||||
source /boot/config/plugins/varaverk/load_config.sh
|
|
||||||
platform_restart_service nginx
|
|
||||||
# wait 15s, then:
|
|
||||||
curl -sf --max-time 5 http://localhost >/dev/null && echo "OK" || echo "still down"
|
|
||||||
|
|
||||||
# If nginx did not fix it, php-fpm:
|
|
||||||
platform_restart_service php-fpm
|
|
||||||
|
|
||||||
# If still down, emhttp:
|
|
||||||
platform_restart_service emhttp
|
|
||||||
|
|
||||||
# Raw equivalents (no source needed — paste directly into terminal):
|
|
||||||
# /etc/rc.d/rc.nginx restart
|
|
||||||
# /etc/rc.d/rc.php-fpm restart
|
|
||||||
# /usr/local/sbin/emhttp stop && /usr/local/sbin/emhttp start
|
|
||||||
|
|
||||||
# If all three failed:
|
|
||||||
server_reboot.sh --status # check for active sessions first
|
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -176,48 +223,6 @@ inotify_tuning.sh --log
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## php_fpm_max_children.sh
|
|
||||||
|
|
||||||
### What It Sets
|
|
||||||
|
|
||||||
```bash
|
|
||||||
PHP_MAX_CHILDREN=250 # target pm.max_children (default: 4-8 on unRAID)
|
|
||||||
PHP_CONF="/etc/php83/php-fpm.d/www.conf"
|
|
||||||
```
|
|
||||||
|
|
||||||
250 workers × ~2MB per worker = ~500MB total. On 128GB this is trivially small.
|
|
||||||
The default of 4–8 saturates immediately under load on a busy server.
|
|
||||||
|
|
||||||
### Verify
|
|
||||||
|
|
||||||
```bash
|
|
||||||
php_fpm_max_children.sh --status
|
|
||||||
# Shows current value vs target, PHP-FPM worker count
|
|
||||||
|
|
||||||
# Manual verify:
|
|
||||||
grep "^pm.max_children" /etc/php83/php-fpm.d/www.conf
|
|
||||||
# Expected: pm.max_children = 250
|
|
||||||
```
|
|
||||||
|
|
||||||
### If WebGUI Is Slow Despite the Setting
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Check PHP-FPM worker utilization (requires system_tuning_monitor.sh in Monitors/):
|
|
||||||
# Look at the webgui_watchdog.sh escalation — step 2 (php-fpm restart) is specifically
|
|
||||||
# for worker exhaustion. If webgui_watchdog.sh is regularly hitting step 2, the
|
|
||||||
# pm.max_children value may still be too low, or there's a PHP worker leak.
|
|
||||||
|
|
||||||
# Check running worker count:
|
|
||||||
pgrep -fc php-fpm
|
|
||||||
# Compare to pm.max_children — if equal, workers are saturated
|
|
||||||
|
|
||||||
# Increase if needed:
|
|
||||||
# master.conf: PHP_MAX_CHILDREN=350
|
|
||||||
# Then: php_fpm_max_children.sh --log (will update and restart php-fpm)
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## docker_syslog_filter.sh
|
## docker_syslog_filter.sh
|
||||||
|
|
||||||
### What It Creates
|
### What It Creates
|
||||||
@@ -230,8 +235,7 @@ pgrep -fc php-fpm
|
|||||||
```
|
```
|
||||||
|
|
||||||
This drops any syslog message containing "veth" or "docker0" before it reaches
|
This drops any syslog message containing "veth" or "docker0" before it reaches
|
||||||
any output target, including the log file. The drop rule is applied at rsyslog
|
any output target, including the log file.
|
||||||
level — not at the log viewer level.
|
|
||||||
|
|
||||||
### Verify
|
### Verify
|
||||||
|
|
||||||
@@ -298,39 +302,6 @@ du -sh /var/lib/docker/containers/*/*.log 2>/dev/null | sort -rh | head -5
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## mover_stop.sh
|
|
||||||
|
|
||||||
### Stop Sequence
|
|
||||||
|
|
||||||
```
|
|
||||||
1. Check if mover is running (platform_is_mover_running) → exit cleanly if not
|
|
||||||
2. Wall message to all logged-in terminal users
|
|
||||||
3. Wait MOVER_STOP_TIMEOUT seconds (default: 30)
|
|
||||||
4. SIGTERM — mover finishes its current file operation, then stops
|
|
||||||
5. Wait 5 seconds → verify stopped
|
|
||||||
6. SIGKILL if still running — forced stop, partial files possible
|
|
||||||
7. Final verify — error if still running after SIGKILL
|
|
||||||
```
|
|
||||||
|
|
||||||
SIGTERM first because the mover can finish the file it is currently moving,
|
|
||||||
leaving no partial copies split across cache and array. SIGKILL is a last resort.
|
|
||||||
|
|
||||||
### Configuration
|
|
||||||
|
|
||||||
```bash
|
|
||||||
MOVER_STOP_TIMEOUT=30 # seconds between wall warning and SIGTERM
|
|
||||||
```
|
|
||||||
|
|
||||||
### Usage
|
|
||||||
|
|
||||||
```bash
|
|
||||||
mover_stop.sh # check and stop if running
|
|
||||||
mover_stop.sh --status # show current mover state and PID
|
|
||||||
mover_stop.sh --dry-run # show what would happen without stopping
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## rsync_stop.sh
|
## rsync_stop.sh
|
||||||
|
|
||||||
### Auto-Detection Logic
|
### Auto-Detection Logic
|
||||||
@@ -355,8 +326,8 @@ After killing rsync, the script checks all containers in `PROFILE_CRITICAL_CONTA
|
|||||||
for any that were stopped by the interrupted rsync session and restarts them.
|
for any that were stopped by the interrupted rsync session and restarts them.
|
||||||
Remote containers are left for docker_watchdog.sh to recover.
|
Remote containers are left for docker_watchdog.sh to recover.
|
||||||
|
|
||||||
Skip container recovery with `--rsync-only` — used when called by other scripts
|
Skip container recovery with `--rsync-only` — used when called by array_stopping.sh
|
||||||
that handle recovery themselves.
|
(which handles container stop separately via docker_container_stop.sh).
|
||||||
|
|
||||||
### Usage
|
### Usage
|
||||||
|
|
||||||
@@ -371,34 +342,6 @@ rsync_stop.sh --full-stop --dry-run # preview full stop
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## user_scripts_stop.sh
|
|
||||||
|
|
||||||
### Process Identification
|
|
||||||
|
|
||||||
Scans `/proc/*/cmdline` for any process whose command line contains
|
|
||||||
`/tmp/user.scripts`. The unRAID User Scripts plugin stages all scripts in
|
|
||||||
`/tmp/user.scripts/` before execution — this signature is reliable regardless of
|
|
||||||
what the script is named or how it was launched.
|
|
||||||
|
|
||||||
Script names are extracted from the path for display: you see which scripts are
|
|
||||||
being stopped, not just PIDs.
|
|
||||||
|
|
||||||
### Self-Exclusion
|
|
||||||
|
|
||||||
If this script is run via the User Scripts plugin, it would find its own PID in
|
|
||||||
the scan. It excludes both `$$` (its own PID) and `$PPID` (its parent process)
|
|
||||||
from the kill list.
|
|
||||||
|
|
||||||
### Usage
|
|
||||||
|
|
||||||
```bash
|
|
||||||
user_scripts_stop.sh # stop all User Script processes
|
|
||||||
user_scripts_stop.sh --status # show running scripts with names and elapsed time
|
|
||||||
user_scripts_stop.sh --dry-run # show what would be stopped
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## server_reboot.sh
|
## server_reboot.sh
|
||||||
|
|
||||||
### Full Shutdown Sequence
|
### Full Shutdown Sequence
|
||||||
@@ -422,11 +365,12 @@ user_scripts_stop.sh --dry-run # show what would be stopped
|
|||||||
6. /etc/rc.d/rc.libvirt stop (VM Manager)
|
6. /etc/rc.d/rc.libvirt stop (VM Manager)
|
||||||
|
|
||||||
7. Orchestrators/array_stopping.sh — safe ordered array stop:
|
7. Orchestrators/array_stopping.sh — safe ordered array stop:
|
||||||
- user_scripts_stop.sh stop background User Scripts
|
- conf_cache_save.sh snapshot partner conf RAM cache
|
||||||
- fallback.sh --stop graceful fallback teardown
|
- user_scripts_stop.sh stop background User Scripts
|
||||||
- rsync_stop.sh --rsync-only kill active rsync transfers
|
- fallback.sh --stop graceful fallback teardown
|
||||||
- mover_stop.sh stop mover
|
- rsync_stop.sh --rsync-only kill active rsync transfers
|
||||||
- docker_container_stop.sh stop all containers gracefully
|
- mover_stop.sh stop mover
|
||||||
|
- docker_container_stop.sh stop all containers gracefully
|
||||||
|
|
||||||
8. sync (flush filesystem buffers to disk)
|
8. sync (flush filesystem buffers to disk)
|
||||||
|
|
||||||
@@ -446,7 +390,6 @@ For a clean reboot when services are active:
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
rsync_stop.sh # stop any active rsync (smart mode)
|
rsync_stop.sh # stop any active rsync (smart mode)
|
||||||
mover_stop.sh # stop mover gracefully
|
|
||||||
server_reboot.sh --status # check what's still running
|
server_reboot.sh --status # check what's still running
|
||||||
server_reboot.sh --reason="planned maintenance"
|
server_reboot.sh --reason="planned maintenance"
|
||||||
```
|
```
|
||||||
@@ -464,27 +407,17 @@ server_reboot.sh --reason="disk work" # include reason in notification
|
|||||||
|
|
||||||
## Full Configuration Reference
|
## Full Configuration Reference
|
||||||
|
|
||||||
> Watchdog configuration (`stability_watchdog.sh`, `resource_watchdog.sh`,
|
|
||||||
> `docker_watchdog.sh`, `System/storage_watchdog.sh`) lives in `Watchdogs/Manual-Watchdogs.md`.
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# master.conf
|
# master.conf
|
||||||
|
|
||||||
# ── WebGUI Watchdog ────────────────────────────────────────────────────────────
|
# ── Conf Sync ──────────────────────────────────────────────────────────────────
|
||||||
WEBGUI_URL="http://localhost"
|
CONF_SYNC_ENABLED=true # toggle: false disables conf_sync.sh entirely
|
||||||
WEBGUI_TIMEOUT=5
|
|
||||||
WEBGUI_NGINX_WAIT=15
|
|
||||||
WEBGUI_PHP_WAIT=10
|
|
||||||
WEBGUI_EMHTTP_WAIT=30
|
|
||||||
|
|
||||||
# ── inotify Tuning ─────────────────────────────────────────────────────────────
|
# ── inotify Tuning ─────────────────────────────────────────────────────────────
|
||||||
INOTIFY_MAX_INSTANCES=1024
|
INOTIFY_MAX_INSTANCES=1024
|
||||||
INOTIFY_MAX_WATCHES=1048576
|
INOTIFY_MAX_WATCHES=1048576
|
||||||
INOTIFY_MAX_QUEUED_EVENTS=32768
|
INOTIFY_MAX_QUEUED_EVENTS=32768
|
||||||
|
INOTIFY_WARN_PCT=80 # warn (in reports) if instances exceed this % of limit
|
||||||
# ── PHP-FPM ────────────────────────────────────────────────────────────────────
|
|
||||||
PHP_MAX_CHILDREN=250
|
|
||||||
PHP_CONF="/etc/php83/php-fpm.d/www.conf"
|
|
||||||
|
|
||||||
# ── Syslog Filter ──────────────────────────────────────────────────────────────
|
# ── Syslog Filter ──────────────────────────────────────────────────────────────
|
||||||
FILTER_FILE="/etc/rsyslog.d/ignore-docker-veth.conf"
|
FILTER_FILE="/etc/rsyslog.d/ignore-docker-veth.conf"
|
||||||
@@ -494,20 +427,48 @@ LOG_FILES=("/var/log/syslog" "/var/log/messages" "/var/log/dmesg")
|
|||||||
LOG_MIN_SIZE_MB=10
|
LOG_MIN_SIZE_MB=10
|
||||||
LOG_DOCKER_MAX_MB=100
|
LOG_DOCKER_MAX_MB=100
|
||||||
|
|
||||||
# ── Mover Stop ─────────────────────────────────────────────────────────────────
|
|
||||||
MOVER_STOP_TIMEOUT=30
|
|
||||||
|
|
||||||
# ── Server Reboot ──────────────────────────────────────────────────────────────
|
# ── Server Reboot ──────────────────────────────────────────────────────────────
|
||||||
REBOOT_SLEEP=30
|
REBOOT_SLEEP=30
|
||||||
REBOOT_VM_WAIT=30
|
REBOOT_VM_WAIT=30
|
||||||
|
|
||||||
|
# ── State Paths (auto-adapt to storage mode) ───────────────────────────────────
|
||||||
|
# STATE_DIR and PERSISTENT_CONF_CACHE are derived from SCRIPTS_DIR in master.conf.
|
||||||
|
# They adapt to internal (/boot/config/plugins/varaverk) or appdata storage mode
|
||||||
|
# (/mnt/user/appdata/Varaverk) automatically — do not hardcode paths.
|
||||||
|
STATE_DIR="${SCRIPTS_DIR}/State_Files"
|
||||||
|
PERSISTENT_CONF_CACHE="${SCRIPTS_DIR}/.cache/vv/d"
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Troubleshooting
|
## Troubleshooting
|
||||||
|
|
||||||
> Watchdog troubleshooting (stability_watchdog, resource_watchdog, docker_watchdog,
|
### Partner Conf Not In RAM Cache After Boot
|
||||||
> System/System/storage_watchdog) is in `Watchdogs/Manual-Watchdogs.md`.
|
|
||||||
|
```bash
|
||||||
|
# Check what's in the RAM cache:
|
||||||
|
ls -la /tmp/.cache/vv/d/
|
||||||
|
|
||||||
|
# Check what's in the persistent backup:
|
||||||
|
ls -la "$PERSISTENT_CONF_CACHE/" # set SCRIPTS_DIR first or use full path
|
||||||
|
|
||||||
|
# Re-run conf sync manually:
|
||||||
|
/boot/config/plugins/varaverk/System_Essentials/conf_sync.sh --log
|
||||||
|
|
||||||
|
# If partner is reachable but pull failed, check SSH key:
|
||||||
|
ssh -i /path/to/ssh_key root@<partner-tailscale-ip> "echo ok"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Persistent Backup Is Stale or Empty
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# conf_cache_watchdog.sh refreshes the backup while partner is offline.
|
||||||
|
# Check watchdog state:
|
||||||
|
cat "$STATE_DIR/conf_cache_watchdog_state.db"
|
||||||
|
|
||||||
|
# Force a manual backup from current RAM cache:
|
||||||
|
/boot/config/plugins/varaverk/System_Essentials/conf_cache_save.sh --log
|
||||||
|
```
|
||||||
|
|
||||||
### rsync_stop Killed the Wrong Thing
|
### rsync_stop Killed the Wrong Thing
|
||||||
|
|
||||||
@@ -521,36 +482,29 @@ rsync_stop.sh --dry-run # shows smart mode decision
|
|||||||
rsync_stop.sh --full-stop --dry-run # shows full-stop decision
|
rsync_stop.sh --full-stop --dry-run # shows full-stop decision
|
||||||
```
|
```
|
||||||
|
|
||||||
### WebGUI Recovery After All Three Steps Failed
|
### inotify Exhaustion After Boot
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Check if processes are running:
|
# Verify limits are applied:
|
||||||
pgrep -x nginx && echo "nginx: yes" || echo "nginx: no"
|
sysctl fs.inotify.max_user_watches # expect 1048576
|
||||||
pgrep emhttpd && echo "emhttp: yes" || echo "emhttp: no"
|
sysctl fs.inotify.max_user_instances # expect 1024
|
||||||
pgrep -f php-fpm && echo "php-fpm: yes" || echo "php-fpm: no"
|
|
||||||
|
|
||||||
# Check recent nginx errors:
|
# If not set — run manually:
|
||||||
cat /var/log/nginx/error.log | tail -20
|
inotify_tuning.sh --log
|
||||||
|
|
||||||
# Check emhttp log:
|
# Check current usage:
|
||||||
tail -20 /var/log/syslog | grep emhttp
|
inotify_tuning.sh --status
|
||||||
|
|
||||||
# Last resort — reboot:
|
|
||||||
server_reboot.sh --reason="WebGUI unrecoverable"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### PHP-FPM Config Not Found After unRAID Update
|
### Syslog Still Noisy After Array Start
|
||||||
|
|
||||||
unRAID updates occasionally change the PHP version. If `php_fpm_max_children.sh`
|
|
||||||
errors with "config file not found":
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Find the new config path:
|
# Check filter is in place:
|
||||||
find /etc -name "www.conf" 2>/dev/null
|
docker_syslog_filter.sh --status
|
||||||
|
|
||||||
# Update PHP_CONF in master.conf:
|
# Re-apply if needed:
|
||||||
PHP_CONF="/etc/php84/php-fpm.d/www.conf" # example for php84
|
docker_syslog_filter.sh --log
|
||||||
|
|
||||||
# Verify with:
|
# Restart rsyslog to pick up the filter:
|
||||||
php_fpm_max_children.sh --status
|
/etc/rc.d/rc.rsyslogd restart
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,35 +1,18 @@
|
|||||||
# ━━━━━ SYSTEM ESSENTIALS ━━━━━
|
# ━━━━━ SYSTEM ESSENTIALS ━━━━━
|
||||||
|
|
||||||
**System-level scripts that act on the server itself — not containers,
|
**System-level scripts that act on the server itself — not containers,
|
||||||
not media, not monitoring.** Keeping the server stable under load, recovering a
|
not media, not monitoring.** Kernel limits, log hygiene, conf synchronisation
|
||||||
frozen WebGUI, tuning kernel limits, suppressing log noise, and handling graceful
|
between servers, and graceful shutdowns with proper warning sequences.
|
||||||
shutdowns with proper warning sequences.
|
|
||||||
|
> **Platform-specific scripts** (`webgui_watchdog.sh`, `php_fpm_max_children.sh`,
|
||||||
|
> `mover_stop.sh`, `user_scripts_stop.sh`) live in `Plugin/unraid/System_Essentials/`
|
||||||
|
> because they call Unraid-specific service commands and paths. This folder
|
||||||
|
> contains scripts that would run unchanged on any Linux host.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ━━━ THE PROBLEM THAT BUILT THIS ━━━
|
## ━━━ THE PROBLEM THAT BUILT THIS ━━━
|
||||||
|
|
||||||
**Server Getting Into Unstable States With No Recovery Path**
|
|
||||||
A container has a memory leak. RAM drops to 2GB. The system starts swapping. Docker
|
|
||||||
watchdog tries to restart the container — but Docker itself is barely responding.
|
|
||||||
The restart hangs. The server needs a reboot, but nothing in the ecosystem is
|
|
||||||
authorized to call one. Or: rootfs fills to 99%. SSH stops working. Docker can't
|
|
||||||
write log files. The server is functionally dead but still technically running.
|
|
||||||
|
|
||||||
Fix: `stability_watchdog.sh` — three-tier response: immediate reboot on critical
|
|
||||||
failures, OOM-confirmed bypass for RAM crises, strike system for sustained
|
|
||||||
threshold breaches. Last line of defense before a hard crash.
|
|
||||||
|
|
||||||
**WebGUI Freezing and Nobody Noticing**
|
|
||||||
The WebGUI becomes unresponsive. Nginx gets into a bad state, or PHP-FPM workers
|
|
||||||
are saturated, or emhttp has frozen. From a user perspective: dashboard doesn't
|
|
||||||
load, settings don't save, containers can't be started or stopped via the UI. No
|
|
||||||
container-level alert fires because this isn't a container problem — it's a web
|
|
||||||
server problem. By the time someone notices it may have been broken for hours.
|
|
||||||
|
|
||||||
Fix: `webgui_watchdog.sh` — checks every 10 minutes, escalates through nginx →
|
|
||||||
php-fpm → emhttp. Lightest fix first. Silent when healthy.
|
|
||||||
|
|
||||||
**50+ Containers Starting and Filling Syslog With Veth Noise**
|
**50+ Containers Starting and Filling Syslog With Veth Noise**
|
||||||
Array starts. 50+ containers come up simultaneously. Docker creates a virtual
|
Array starts. 50+ containers come up simultaneously. Docker creates a virtual
|
||||||
network interface for each one. Each interface generates multiple syslog entries.
|
network interface for each one. Each interface generates multiple syslog entries.
|
||||||
@@ -40,68 +23,79 @@ In the first minute after array start, syslog is buried under 200–400 lines of
|
|||||||
Fix: `docker_syslog_filter.sh` — creates an rsyslog drop rule before any container
|
Fix: `docker_syslog_filter.sh` — creates an rsyslog drop rule before any container
|
||||||
starts. Applied at array start. Idempotent — silent when already correct.
|
starts. Applied at array start. Idempotent — silent when already correct.
|
||||||
|
|
||||||
**WebGUI Queuing Requests Under Load Without Explanation**
|
|
||||||
The WebGUI feels slow. Clicking a button takes 5 seconds. Nothing in the logs
|
|
||||||
explains it. The cause: PHP-FPM's `pm.max_children` defaults to 4–8 workers. With
|
|
||||||
multiple users, active plugins, and 50+ containers potentially hitting the WebGUI,
|
|
||||||
those workers saturate immediately. New requests queue behind active ones.
|
|
||||||
|
|
||||||
Fix: `php_fpm_max_children.sh` — sets `pm.max_children=250` at array start.
|
|
||||||
250 workers × ~2MB = ~500MB total. On 128GB this is trivially small.
|
|
||||||
|
|
||||||
**inotify Exhaustion Producing Unexplained Failures**
|
**inotify Exhaustion Producing Unexplained Failures**
|
||||||
When inotify limits are exhausted, containers silently stop receiving filesystem
|
When inotify limits are exhausted, containers silently stop receiving filesystem
|
||||||
events. Arrs don't detect completed downloads. VSCode shows "unable to watch for
|
events. Arrs don't detect completed downloads. VSCode shows "unable to watch for
|
||||||
file changes." Code-Server with node_modules alone can consume 100K–200K watches,
|
file changes." Code-Server with node_modules alone can consume 100K–200K watches,
|
||||||
and all containers share the same pool.
|
and all containers share the same pool.
|
||||||
|
|
||||||
Fix: `inotify_tuning.sh` — raises all three inotify limits at array start. Must
|
Fix: `inotify_tuning.sh` — raises all three inotify limits at array start, before
|
||||||
run FIRST in ARRAY_START_SCRIPTS before any containers start.
|
any container-connecting or continuous scripts start.
|
||||||
|
|
||||||
**Mover Getting Killed Mid-Transfer Leaving Files Inconsistent**
|
**Partner Conf Lost Across Reboots When Partner is Down**
|
||||||
The mover is running — moving a large batch of files from cache to array. A reboot
|
Scripts like `fallback.sh` need the partner's conf vars (credentials, container
|
||||||
is triggered. The mover stops mid-file. The file exists partially on both cache and
|
names, tier delays) to operate. The partner conf lives in a RAM cache at
|
||||||
array simultaneously. unRAID's deduplication layer is confused.
|
`/tmp/.cache/vv/d/` — wiped every reboot. At array start, `conf_sync.sh` pulls
|
||||||
|
a fresh copy from the partner. But if the partner is offline at boot time, the
|
||||||
|
pull fails and fallback has no partner vars to work with.
|
||||||
|
|
||||||
Fix: `mover_stop.sh` — warns users via wall message, waits the configured timeout,
|
Fix: `conf_cache_save.sh` + `conf_cache_restore.sh` — snapshot the RAM cache to
|
||||||
SIGTERM (graceful — finishes current file), SIGKILL only if needed.
|
`$PERSISTENT_CONF_CACHE` on array stop; reload on next start for any confs the
|
||||||
|
sync couldn't fetch.
|
||||||
|
|
||||||
|
**Mover/Rsync Interruption During Reboot**
|
||||||
|
Rsync transfers or mover runs are in progress when a reboot is triggered. Stopping
|
||||||
|
them uncleanly leaves partial files.
|
||||||
|
|
||||||
|
Fix: `server_reboot.sh` — orchestrates `array_stopping.sh` which stops rsync, mover,
|
||||||
|
and containers in the correct order before calling `/sbin/reboot`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ━━━ WHAT THIS FOLDER DOES ━━━
|
## ━━━ WHAT THIS FOLDER DOES ━━━
|
||||||
|
|
||||||
```
|
```
|
||||||
WebGUI availability webgui_watchdog.sh — nginx → php-fpm → emhttp escalation
|
|
||||||
Kernel tuning inotify_tuning.sh — file watch limits
|
Kernel tuning inotify_tuning.sh — file watch limits
|
||||||
php_fpm_max_children.sh — PHP worker count
|
|
||||||
Log hygiene docker_syslog_filter.sh — suppress veth noise at start
|
Log hygiene docker_syslog_filter.sh — suppress veth noise at start
|
||||||
clear_logs.sh — weekly log trimming
|
clear_logs.sh — size-threshold log trimming
|
||||||
Graceful operations mover_stop.sh — clean mover stop
|
Conf synchronisation conf_sync.sh — pull/push partner confs → RAM cache
|
||||||
rsync_stop.sh — smart rsync stop (orchestrator-aware)
|
conf_cache_save.sh — snapshot RAM cache → persistent at stop
|
||||||
user_scripts_stop.sh — stop running User Scripts
|
conf_cache_restore.sh — reload from snapshot at start (offline partner)
|
||||||
|
Graceful operations rsync_stop.sh — smart rsync stop (orchestrator-aware)
|
||||||
server_reboot.sh — clean reboot with pre-flight warnings
|
server_reboot.sh — clean reboot with pre-flight warnings
|
||||||
```
|
```
|
||||||
|
|
||||||
> `stability_watchdog.sh` and `resource_watchdog.sh` have moved to `Watchdogs/`.
|
|
||||||
> See `Watchdogs/README-Watchdogs.md` for the full watchdog suite.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ━━━ RELATIONSHIP TO OTHER FOLDERS ━━━
|
## ━━━ RELATIONSHIP TO OTHER FOLDERS ━━━
|
||||||
|
|
||||||
```
|
```
|
||||||
Orchestrators/
|
Orchestrators/
|
||||||
array_started.sh ─────────────────────────► inotify_tuning.sh (first in sequence)
|
array_started.sh ────────────────────────► conf_sync.sh (pulls partner confs)
|
||||||
─────────────────────────► docker_syslog_filter.sh (second)
|
────────────────────────► conf_cache_restore.sh (fills gaps if partner down)
|
||||||
─────────────────────────► php_fpm_max_children.sh
|
────────────────────────► docker_syslog_filter.sh (before containers start)
|
||||||
|
────────────────────────► inotify_tuning.sh (before docker_network_connect)
|
||||||
|
array_stopping.sh ───────────────────────► conf_cache_save.sh (first — while cache is fresh)
|
||||||
|
───────────────────────► rsync_stop.sh --rsync-only (stop transfers)
|
||||||
weekly_maintenance.sh ──────────────────► clear_logs.sh
|
weekly_maintenance.sh ──────────────────► clear_logs.sh
|
||||||
|
|
||||||
server_reboot.sh ────────────────────────► user_scripts_stop.sh (called internally)
|
Watchdogs/System/
|
||||||
|
conf_cache_watchdog.sh ─────────────────► maintains $PERSISTENT_CONF_CACHE while partner offline
|
||||||
|
(runs every 15 min via watchdog_orchestrator.sh)
|
||||||
|
|
||||||
Watchdogs/
|
System_Essentials/
|
||||||
stability_watchdog.sh and resource_watchdog.sh now live here.
|
server_reboot.sh ────────────────────────► array_stopping.sh (via Orchestrators/)
|
||||||
See Watchdogs/README-Watchdogs.md for how they relate to each other
|
────────────────────────► mover_stop.sh, user_scripts_stop.sh
|
||||||
and to docker_watchdog.sh and System/storage_watchdog.sh.
|
(via Plugin/unraid/System_Essentials/)
|
||||||
|
|
||||||
|
Plugin/unraid/System_Essentials/
|
||||||
|
php_fpm_max_children.sh — WebGUI tuning (Unraid-specific PHP paths)
|
||||||
|
mover_stop.sh — clean mover stop (Unraid mover daemon)
|
||||||
|
user_scripts_stop.sh — stop User Scripts plugin processes
|
||||||
|
unraid_api_key_renew.sh — Varaverk plugin API key renewal
|
||||||
|
|
||||||
|
Plugin/unraid/Watchdogs/System/
|
||||||
|
webgui_watchdog.sh — nginx → php-fpm → emhttp escalation
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -110,14 +104,13 @@ Watchdogs/
|
|||||||
|
|
||||||
| Script | Role | When It Runs |
|
| Script | Role | When It Runs |
|
||||||
|--------|------|-------------|
|
|--------|------|-------------|
|
||||||
| `webgui_watchdog.sh` | WebGUI availability — nginx → php-fpm → emhttp | Every minute via watchdog_orchestrator → system_watchdog |
|
| `inotify_tuning.sh` | Raise inotify kernel limits | At array start — before container ops |
|
||||||
| `inotify_tuning.sh` | Raise inotify kernel limits | At array start — FIRST |
|
|
||||||
| `php_fpm_max_children.sh` | Set PHP-FPM max worker count | At array start |
|
|
||||||
| `docker_syslog_filter.sh` | Suppress Docker veth syslog noise | At array start — before containers |
|
| `docker_syslog_filter.sh` | Suppress Docker veth syslog noise | At array start — before containers |
|
||||||
|
| `conf_sync.sh` | Sync partner confs: pull → RAM cache, push own conf to partners | At array start + every 4 h (--pull-only) |
|
||||||
|
| `conf_cache_save.sh` | Snapshot partner RAM cache → `$PERSISTENT_CONF_CACHE` | At array stop — first step |
|
||||||
|
| `conf_cache_restore.sh` | Load missing partner confs from persistent backup into RAM | At array start — after conf_sync |
|
||||||
| `clear_logs.sh` | Size-threshold log cleanup | Weekly via weekly_maintenance.sh |
|
| `clear_logs.sh` | Size-threshold log cleanup | Weekly via weekly_maintenance.sh |
|
||||||
| `mover_stop.sh` | Clean mover stop with SIGTERM → SIGKILL | Manual / before reboot |
|
| `rsync_stop.sh` | Orchestrator-aware rsync stop | Manual / called by array_stopping.sh |
|
||||||
| `rsync_stop.sh` | Orchestrator-aware rsync stop | Manual |
|
|
||||||
| `user_scripts_stop.sh` | Stop all running User Script processes | Manual / called by server_reboot.sh |
|
|
||||||
| `server_reboot.sh` | Graceful reboot with pre-flight warnings | Manual |
|
| `server_reboot.sh` | Graceful reboot with pre-flight warnings | Manual |
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -125,27 +118,28 @@ Watchdogs/
|
|||||||
## ━━━ HOW THE SCRIPTS RELATE ━━━
|
## ━━━ HOW THE SCRIPTS RELATE ━━━
|
||||||
|
|
||||||
```
|
```
|
||||||
Array starts
|
Array starts (array_started.sh, ARRAY_START_SCRIPTS):
|
||||||
│
|
│
|
||||||
├─ inotify_tuning.sh ← FIRST — kernel limits inherited at container launch
|
├─ conf_sync.sh ← SSH/SCP: pull partner confs into /tmp/.cache/vv/d/
|
||||||
├─ docker_syslog_filter.sh ← SECOND — before any veth interfaces are created
|
│ push own conf to partner's /tmp/.cache/vv/d/
|
||||||
└─ php_fpm_max_children.sh ← before WebGUI is under load
|
├─ conf_cache_restore.sh ← if partner was down: load last-known-good conf from
|
||||||
|
│ $PERSISTENT_CONF_CACHE into /tmp/.cache/vv/d/
|
||||||
|
├─ docker_syslog_filter.sh ← before any container starts (veth filter must be live)
|
||||||
|
└─ inotify_tuning.sh ← before docker_network_connect.sh and continuous scripts
|
||||||
|
|
||||||
Every minute (watchdog_orchestrator.sh in Orchestrators/):
|
Array stops (array_stopping.sh, ARRAY_STOP_SCRIPTS):
|
||||||
→ Watchdogs/resource_watchdog.sh
|
│
|
||||||
→ Watchdogs/docker_watchdog.sh
|
├─ conf_cache_save.sh ← FIRST: snapshot /tmp/.cache/vv/d/ → $PERSISTENT_CONF_CACHE
|
||||||
→ Watchdogs/system_watchdog.sh (thin orchestrator)
|
│ while RAM cache is still fresh
|
||||||
└─ Watchdogs/System/storage_watchdog.sh
|
├─ rsync_stop.sh --rsync-only ← kill active rsync, skip container recovery
|
||||||
└─ Watchdogs/System/webgui_watchdog.sh
|
└─ ...other stop scripts...
|
||||||
WebGUI OK → silent exit
|
|
||||||
Not responding:
|
Every 15 minutes (watchdog_orchestrator.sh):
|
||||||
Step 1: restart nginx → recheck
|
→ Watchdogs/System/conf_cache_watchdog.sh
|
||||||
Step 2: restart php-fpm → recheck
|
If partner is offline and persistent backup is stale → refresh from last RAM cache
|
||||||
Step 3: restart emhttp → recheck
|
|
||||||
All failed → notify, exit 1
|
Every 4 hours (intermediate_sync_maintenance.sh):
|
||||||
└─ Watchdogs/System/network_watchdog.sh
|
└─ conf_sync.sh --pull-only ← refresh partner conf in RAM without pushing own conf
|
||||||
→ Watchdogs/stability_watchdog.sh
|
|
||||||
(see Watchdogs/README-Watchdogs.md for full flow)
|
|
||||||
|
|
||||||
Weekly (weekly_maintenance.sh):
|
Weekly (weekly_maintenance.sh):
|
||||||
└─ clear_logs.sh
|
└─ clear_logs.sh
|
||||||
@@ -153,8 +147,6 @@ Weekly (weekly_maintenance.sh):
|
|||||||
Docker logs: clear per-container if > LOG_DOCKER_MAX_MB
|
Docker logs: clear per-container if > LOG_DOCKER_MAX_MB
|
||||||
|
|
||||||
Manual operations:
|
Manual operations:
|
||||||
mover_stop.sh → wall → SIGTERM → SIGKILL → verify stopped
|
rsync_stop.sh → detect orchestrator → kill rsync (or orchestrator+rsync)
|
||||||
rsync_stop.sh → detect orchestrator → kill rsync (or orchestrator+rsync)
|
server_reboot.sh → pre-flight → array_stopping.sh → reboot
|
||||||
user_scripts_stop.sh → scan /proc → SIGTERM → SIGKILL per process
|
|
||||||
server_reboot.sh → pre-flight → wall → wait → VMs → Docker → sync → reboot
|
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -30,7 +30,7 @@ detect_hosts
|
|||||||
require_partnership
|
require_partnership
|
||||||
|
|
||||||
RAM_CACHE="/tmp/.cache/vv/d"
|
RAM_CACHE="/tmp/.cache/vv/d"
|
||||||
SAVE_DIR="/boot/config/.cache/vv/d"
|
SAVE_DIR="$PERSISTENT_CONF_CACHE"
|
||||||
|
|
||||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
||||||
|
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
#
|
#
|
||||||
# PURPOSE
|
# PURPOSE
|
||||||
# ─────────────────────────────────────────────────────────────────────────────
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
# Snapshots the partner conf RAM cache to /boot/config/.cache/vv/d/ on
|
# Snapshots the partner conf RAM cache to $PERSISTENT_CONF_CACHE on
|
||||||
# array stop. Survives reboot. Used by conf_cache_restore.sh at next array start
|
# array stop. Survives reboot. Used by conf_cache_restore.sh at next array start
|
||||||
# to reload partner vars into RAM when the partner is unreachable at boot time.
|
# to reload partner vars into RAM when the partner is unreachable at boot time.
|
||||||
#
|
#
|
||||||
@@ -15,7 +15,7 @@
|
|||||||
# down, the backup fills the gap so fallback.sh has the vars it needs.
|
# down, the backup fills the gap so fallback.sh has the vars it needs.
|
||||||
#
|
#
|
||||||
# Only partner confs are saved — own conf is always on disk.
|
# Only partner confs are saved — own conf is always on disk.
|
||||||
# Location is outside the git repo and outside the main plugin folder.
|
# Path adapts to storage mode: $SCRIPTS_DIR/.cache/vv/d (internal or appdata).
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
|
|
||||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
@@ -26,7 +26,7 @@ detect_hosts
|
|||||||
require_partnership
|
require_partnership
|
||||||
|
|
||||||
RAM_CACHE="/tmp/.cache/vv/d"
|
RAM_CACHE="/tmp/.cache/vv/d"
|
||||||
SAVE_DIR="/boot/config/.cache/vv/d"
|
SAVE_DIR="$PERSISTENT_CONF_CACHE"
|
||||||
|
|
||||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be written"
|
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be written"
|
||||||
|
|
||||||
|
|||||||
@@ -11,15 +11,18 @@
|
|||||||
# On array start (default / --array-start):
|
# On array start (default / --array-start):
|
||||||
# 1. Copy own conf to local cache
|
# 1. Copy own conf to local cache
|
||||||
# 2. Pull each available partner's conf from their disk → local cache
|
# 2. Pull each available partner's conf from their disk → local cache
|
||||||
# 3. Push own conf to each available partner's /tmp/.vv/ cache
|
# 3. Push own conf to each available partner's /tmp/.cache/vv/d/ cache
|
||||||
#
|
#
|
||||||
# On conf save (--push-only):
|
# On conf save (--push-only):
|
||||||
# Fast path — push updated own conf to all partners' /tmp/.vv/ cache only.
|
# Fast path — push updated own conf to all partners' /tmp/.cache/vv/d/ only.
|
||||||
# No pulls, no local cache rebuild.
|
# No pulls, no local cache rebuild.
|
||||||
#
|
#
|
||||||
# Cache is /tmp (tmpfs) — cleared every reboot, repopulated by this script
|
# Cache is /tmp/.cache/vv/d (tmpfs) — cleared every reboot, repopulated by
|
||||||
# on next array start. Scripts source from cache for partner vars; own vars
|
# this script on next array start. Scripts source from cache for partner vars;
|
||||||
# always come from disk (load_config.sh skips cached copy of own conf).
|
# own vars always come from disk (load_config.sh skips cached copy of own conf).
|
||||||
|
#
|
||||||
|
# Pull path resolves the remote's SCRIPTS_DIR from their varaverk.cfg so it
|
||||||
|
# works whether the remote is in internal or appdata storage mode.
|
||||||
#
|
#
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# RUNTIME MODES
|
# RUNTIME MODES
|
||||||
@@ -56,9 +59,25 @@ if [[ "${CONF_SYNC_ENABLED:-true}" == false ]]; then
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
CACHE_DIR="/tmp/.cache/vv/d"
|
CACHE_DIR="/tmp/.cache/vv/d"
|
||||||
MY_CONF="$SCRIPTS_ROOT/Configurations/${MY_ID,,}.conf"
|
MY_CONF="$SCRIPTS_DIR/Configurations/${MY_ID,,}.conf"
|
||||||
SSH_TIMEOUT=10
|
SSH_TIMEOUT=10
|
||||||
|
|
||||||
|
# Reads the remote's varaverk.cfg to find their actual SCRIPTS_DIR.
|
||||||
|
# Handles the case where the remote is in appdata storage mode.
|
||||||
|
_remote_scripts_dir() {
|
||||||
|
local ip="$1"
|
||||||
|
local cfg line sd
|
||||||
|
cfg=$(timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
|
||||||
|
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
|
||||||
|
"root@${ip}" "cat /boot/config/plugins/varaverk/varaverk.cfg 2>/dev/null" 2>/dev/null) || true
|
||||||
|
while IFS= read -r line; do
|
||||||
|
[[ "$line" == SCRIPTS_DIR=* ]] || continue
|
||||||
|
sd="${line#SCRIPTS_DIR=}"; sd="${sd//\"/}"; sd="${sd//\'/}"
|
||||||
|
echo "$sd"; return
|
||||||
|
done <<< "$cfg"
|
||||||
|
echo "/boot/config/plugins/varaverk"
|
||||||
|
}
|
||||||
|
|
||||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
||||||
|
|
||||||
# ── Ensure cache dir exists ───────────────────────────────────────────────────
|
# ── Ensure cache dir exists ───────────────────────────────────────────────────
|
||||||
@@ -101,7 +120,8 @@ for host_var in $(compgen -v | grep -E '^HOST[0-9]+$' | sort); do
|
|||||||
|
|
||||||
# ── Pull: grab partner's conf from their disk → our local cache ──────────
|
# ── Pull: grab partner's conf from their disk → our local cache ──────────
|
||||||
if [[ "$PUSH_ONLY" == false ]]; then
|
if [[ "$PUSH_ONLY" == false ]]; then
|
||||||
remote_conf="${SCRIPTS_DIR}/Configurations/${partner_slot}.conf"
|
remote_sd=$(_remote_scripts_dir "$partner_ip")
|
||||||
|
remote_conf="${remote_sd}/Configurations/${partner_slot}.conf"
|
||||||
if [[ "$DRY_RUN" == true ]]; then
|
if [[ "$DRY_RUN" == true ]]; then
|
||||||
warn "DRY RUN — would pull $partner_host:$remote_conf → $CACHE_DIR/${partner_slot}.conf"
|
warn "DRY RUN — would pull $partner_host:$remote_conf → $CACHE_DIR/${partner_slot}.conf"
|
||||||
elif timeout "$SSH_TIMEOUT" scp -i "$SSH_KEY" \
|
elif timeout "$SSH_TIMEOUT" scp -i "$SSH_KEY" \
|
||||||
@@ -116,7 +136,7 @@ for host_var in $(compgen -v | grep -E '^HOST[0-9]+$' | sort); do
|
|||||||
fi
|
fi
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# ── Push: send own conf to partner's /tmp/.vv/ cache ────────────────────
|
# ── Push: send own conf to partner's /tmp/.cache/vv/d/ cache ───────────
|
||||||
if [[ "$PULL_ONLY" == true ]]; then
|
if [[ "$PULL_ONLY" == true ]]; then
|
||||||
continue
|
continue
|
||||||
fi
|
fi
|
||||||
|
|||||||
@@ -214,7 +214,7 @@ View and manage the persistent container skip list used by `docker_watchdog.sh`.
|
|||||||
|
|
||||||
```
|
```
|
||||||
docker_watchdog.sh restarts the same container N times within the rolling window
|
docker_watchdog.sh restarts the same container N times within the rolling window
|
||||||
→ container added to skip list on /boot/config/
|
→ container added to skip list in $STATE_DIR
|
||||||
→ critical notification sent
|
→ critical notification sent
|
||||||
→ watchdog stops touching it entirely
|
→ watchdog stops touching it entirely
|
||||||
|
|
||||||
|
|||||||
@@ -13,8 +13,9 @@ For the orchestrator that calls these scripts see `Orchestrators/watchdog_orches
|
|||||||
- [docker_watchdog.sh](#docker_watchdogsh)
|
- [docker_watchdog.sh](#docker_watchdogsh)
|
||||||
- [system_watchdog.sh](#system_watchdogsh)
|
- [system_watchdog.sh](#system_watchdogsh)
|
||||||
- [System/storage_watchdog.sh](#systemstorage_watchdogsh)
|
- [System/storage_watchdog.sh](#systemstorage_watchdogsh)
|
||||||
- [System/webgui_watchdog.sh](#systemwebgui_watchdogsh)
|
- [System/webgui_watchdog.sh](#systemwebgui_watchdogsh) ← Plugin/unraid/Watchdogs/System/
|
||||||
- [System/network_watchdog.sh](#systemnetwork_watchdogsh)
|
- [System/network_watchdog.sh](#systemnetwork_watchdogsh)
|
||||||
|
- [System/conf_cache_watchdog.sh](#systemconf_cache_watchdogsh)
|
||||||
- [stability_watchdog.sh](#stability_watchdogsh)
|
- [stability_watchdog.sh](#stability_watchdogsh)
|
||||||
- [Full Configuration Reference](#full-configuration-reference)
|
- [Full Configuration Reference](#full-configuration-reference)
|
||||||
- [Troubleshooting](#troubleshooting)
|
- [Troubleshooting](#troubleshooting)
|
||||||
@@ -373,6 +374,10 @@ storage_watchdog.sh --log # verbose per-container output
|
|||||||
|
|
||||||
## System/webgui_watchdog.sh
|
## System/webgui_watchdog.sh
|
||||||
|
|
||||||
|
> **Lives in `Plugin/unraid/Watchdogs/System/webgui_watchdog.sh`** — calls Unraid-specific
|
||||||
|
> service commands (`rc.nginx`, `rc.php-fpm`, `emhttp`) via the platform adapter. Called by
|
||||||
|
> `system_watchdog.sh` via `SYSTEM_WATCHDOG_SCRIPTS` in master.conf.
|
||||||
|
|
||||||
Called by `system_watchdog.sh` each cycle. Monitors WebGUI availability and escalates
|
Called by `system_watchdog.sh` each cycle. Monitors WebGUI availability and escalates
|
||||||
through three restart steps if unresponsive. Silent when healthy.
|
through three restart steps if unresponsive. Silent when healthy.
|
||||||
|
|
||||||
@@ -446,7 +451,7 @@ NETWORK_WATCHDOG_INTERNET_TIMEOUT=5
|
|||||||
NETWORK_WATCHDOG_CHECK_TAILSCALE=true
|
NETWORK_WATCHDOG_CHECK_TAILSCALE=true
|
||||||
NETWORK_WATCHDOG_NPM_TIMEOUT=10
|
NETWORK_WATCHDOG_NPM_TIMEOUT=10
|
||||||
NETWORK_WATCHDOG_NPM_STRIKE_LIMIT=2
|
NETWORK_WATCHDOG_NPM_STRIKE_LIMIT=2
|
||||||
NETWORK_WATCHDOG_NPM_STATE_FILE="/tmp/network_watchdog_state.db"
|
NETWORK_WATCHDOG_NPM_STATE_FILE="$STATE_DIR/network_watchdog_state.db"
|
||||||
|
|
||||||
# host*.conf (host-specific)
|
# host*.conf (host-specific)
|
||||||
HOST1_NETWORK_WATCHDOG_DDNS_DOMAIN="gmer4lfe.com"
|
HOST1_NETWORK_WATCHDOG_DDNS_DOMAIN="gmer4lfe.com"
|
||||||
@@ -506,6 +511,40 @@ tailscale up
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## System/conf_cache_watchdog.sh
|
||||||
|
|
||||||
|
Called by `system_watchdog.sh` each cycle. Maintains the persistent partner conf backup
|
||||||
|
at `$PERSISTENT_CONF_CACHE` while the partner is offline.
|
||||||
|
|
||||||
|
### Behaviour
|
||||||
|
|
||||||
|
**Remote online:** removes the persistent backup if one exists. It is not needed —
|
||||||
|
`conf_sync.sh` will pull a fresh copy on the next boot. Silent when backup is already absent.
|
||||||
|
|
||||||
|
**Remote offline:** copies partner confs from the RAM cache (`/tmp/.cache/vv/d/`) to
|
||||||
|
`$PERSISTENT_CONF_CACHE`. Runs every 15 minutes, so the backup stays current throughout
|
||||||
|
an extended outage. If this host reboots while the partner is still down,
|
||||||
|
`conf_cache_restore.sh` will load the backup into RAM and fallback.sh will have
|
||||||
|
valid partner vars.
|
||||||
|
|
||||||
|
Silent when remote is online and no backup exists (the normal steady state).
|
||||||
|
|
||||||
|
### Gates
|
||||||
|
|
||||||
|
- `FALLBACK_ENABLED=false` → no-op (no fallback means no need for partner vars)
|
||||||
|
- `CONF_SYNC_ENABLED=false` → no-op
|
||||||
|
- `PARTNERSHIP_ENABLED=false` → exits silently (require_partnership gate)
|
||||||
|
|
||||||
|
### Usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conf_cache_watchdog.sh # single pass (called by system_watchdog.sh)
|
||||||
|
conf_cache_watchdog.sh --dry-run # show what would be written or removed
|
||||||
|
conf_cache_watchdog.sh --log # verbose output
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## stability_watchdog.sh
|
## stability_watchdog.sh
|
||||||
|
|
||||||
Runs last in the orchestrator sequence. The only script in the ecosystem authorized
|
Runs last in the orchestrator sequence. The only script in the ecosystem authorized
|
||||||
@@ -653,9 +692,9 @@ WATCHDOG_RESTART_DEAD=true
|
|||||||
WATCHDOG_RESTART_CRASHED=true
|
WATCHDOG_RESTART_CRASHED=true
|
||||||
WATCHDOG_BATCH_NOTIFY=true
|
WATCHDOG_BATCH_NOTIFY=true
|
||||||
|
|
||||||
# State files:
|
# State files (all paths adapt to storage mode via STATE_DIR / DATA_DIR):
|
||||||
WATCHDOG_STATE_FILE="/tmp/container_watchdog_state.db"
|
WATCHDOG_STATE_FILE="$STATE_DIR/container_watchdog_state.db"
|
||||||
WATCHDOG_CONTAINER_RESTART_LOG="/boot/config/container_restart_history.db"
|
WATCHDOG_CONTAINER_RESTART_LOG="$DATA_DIR/container_restart_history.db"
|
||||||
|
|
||||||
# host*.conf
|
# host*.conf
|
||||||
HOST1_WATCHDOG_CONTAINERS=() # "ContainerName:LimitMB"
|
HOST1_WATCHDOG_CONTAINERS=() # "ContainerName:LimitMB"
|
||||||
@@ -670,8 +709,8 @@ WATCHDOG_APPDATA_GROWTH_GB=2
|
|||||||
WATCHDOG_APPDATA_LOG_MAX_GB=2
|
WATCHDOG_APPDATA_LOG_MAX_GB=2
|
||||||
WATCHDOG_APPDATA_TRUNCATE_LOGS=false
|
WATCHDOG_APPDATA_TRUNCATE_LOGS=false
|
||||||
WATCHDOG_APPDATA_STRIKE_LIMIT=3
|
WATCHDOG_APPDATA_STRIKE_LIMIT=3
|
||||||
WATCHDOG_APPDATA_GROWTH_FILE="/tmp/watchdog_appdata_growth.db"
|
WATCHDOG_APPDATA_GROWTH_FILE="$STATE_DIR/watchdog_appdata_growth.db"
|
||||||
STORAGE_WATCHDOG_STATE_FILE="/tmp/storage_watchdog_state.db"
|
STORAGE_WATCHDOG_STATE_FILE="$STATE_DIR/storage_watchdog_state.db"
|
||||||
|
|
||||||
# host*.conf (optional — only for suppress ceilings)
|
# host*.conf (optional — only for suppress ceilings)
|
||||||
# declare -A HOST1_WATCHDOG_APPDATA_SIZES=(
|
# declare -A HOST1_WATCHDOG_APPDATA_SIZES=(
|
||||||
@@ -701,11 +740,12 @@ SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true
|
|||||||
SYS_WATCHDOG_ABORT_ON_PARITY=false
|
SYS_WATCHDOG_ABORT_ON_PARITY=false
|
||||||
SYS_WATCHDOG_ABORT_ON_MOVER=false
|
SYS_WATCHDOG_ABORT_ON_MOVER=false
|
||||||
|
|
||||||
# State files (all in State_Files/ — survive reboots):
|
# State files — all in $STATE_DIR (survive reboots; adapt to storage mode):
|
||||||
SYS_WATCHDOG_STATE_FILE="State_Files/system_watchdog_state.db"
|
SYS_WATCHDOG_STATE_FILE="$STATE_DIR/system_watchdog_state.db"
|
||||||
DOCKER_WATCHDOG_FAILED_FILE="State_Files/docker_watchdog_failed.db"
|
DOCKER_WATCHDOG_FAILED_FILE="$STATE_DIR/docker_watchdog_failed.db"
|
||||||
SYS_WATCHDOG_REBOOT_LOG="State_Files/system_watchdog_reboots.db"
|
SYS_WATCHDOG_REBOOT_LOG="$STATE_DIR/system_watchdog_reboots.db"
|
||||||
SYS_WATCHDOG_OOM_FILE="State_Files/system_watchdog_oom.db"
|
SYS_WATCHDOG_OOM_FILE="$STATE_DIR/system_watchdog_oom.db"
|
||||||
|
RW_STATE_FILE="$STATE_DIR/resource_watchdog_state.db"
|
||||||
|
|
||||||
# ── Network Watchdog ───────────────────────────────────────────────────────────
|
# ── Network Watchdog ───────────────────────────────────────────────────────────
|
||||||
NETWORK_WATCHDOG_ENABLED=true
|
NETWORK_WATCHDOG_ENABLED=true
|
||||||
@@ -714,7 +754,7 @@ NETWORK_WATCHDOG_INTERNET_TIMEOUT=5
|
|||||||
NETWORK_WATCHDOG_CHECK_TAILSCALE=true
|
NETWORK_WATCHDOG_CHECK_TAILSCALE=true
|
||||||
NETWORK_WATCHDOG_NPM_TIMEOUT=10
|
NETWORK_WATCHDOG_NPM_TIMEOUT=10
|
||||||
NETWORK_WATCHDOG_NPM_STRIKE_LIMIT=2
|
NETWORK_WATCHDOG_NPM_STRIKE_LIMIT=2
|
||||||
NETWORK_WATCHDOG_NPM_STATE_FILE="/tmp/network_watchdog_state.db"
|
NETWORK_WATCHDOG_NPM_STATE_FILE="$STATE_DIR/network_watchdog_state.db"
|
||||||
|
|
||||||
# host*.conf (host-specific)
|
# host*.conf (host-specific)
|
||||||
HOST1_NETWORK_WATCHDOG_DDNS_DOMAIN="gmer4lfe.com"
|
HOST1_NETWORK_WATCHDOG_DDNS_DOMAIN="gmer4lfe.com"
|
||||||
@@ -781,8 +821,8 @@ storage_watchdog.sh --status
|
|||||||
### stability_watchdog Rebooted Unexpectedly
|
### stability_watchdog Rebooted Unexpectedly
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Check the reboot log (survives reboots):
|
# Check the reboot log (survives reboots — in State_Files/):
|
||||||
cat /boot/config/system_watchdog_reboots.db
|
cat /boot/config/plugins/varaverk/State_Files/system_watchdog_reboots.db
|
||||||
# Shows timestamp and reason for each watchdog-triggered reboot
|
# Shows timestamp and reason for each watchdog-triggered reboot
|
||||||
|
|
||||||
# Check syslog near the reboot time:
|
# Check syslog near the reboot time:
|
||||||
|
|||||||
@@ -3,8 +3,8 @@
|
|||||||
**Four tiers that run every 15 minutes through `watchdog_orchestrator.sh`,
|
**Four tiers that run every 15 minutes through `watchdog_orchestrator.sh`,
|
||||||
each with a clear lane:** reduce system pressure → heal containers → check system components →
|
each with a clear lane:** reduce system pressure → heal containers → check system components →
|
||||||
reboot if nothing else worked. The orchestrator calls them in order, once per cron cycle.
|
reboot if nothing else worked. The orchestrator calls them in order, once per cron cycle.
|
||||||
System component checks (storage, WebGUI, network) run inside a thin `system_watchdog.sh`
|
System component checks (storage, WebGUI, network, conf cache) run inside a thin
|
||||||
orchestrator that can also be run standalone.
|
`system_watchdog.sh` orchestrator that can also be run standalone.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -57,6 +57,15 @@ Fix: `stability_watchdog.sh` — watches the server itself: RAM, CPU, disk, kern
|
|||||||
health. Only script in the stack authorized to reboot. Runs last in the orchestrator
|
health. Only script in the stack authorized to reboot. Runs last in the orchestrator
|
||||||
sequence so container healing and pressure reduction always get a chance first.
|
sequence so container healing and pressure reduction always get a chance first.
|
||||||
|
|
||||||
|
**Partner Conf Backup Going Stale During a Long Outage**
|
||||||
|
The partner goes offline overnight. On the next reboot (planned maintenance), the
|
||||||
|
persistent conf backup saved at shutdown is stale — it reflects the state from before
|
||||||
|
the outage, not the most recent live values. `conf_cache_restore.sh` loads it
|
||||||
|
faithfully, but fallback.sh may be working with old tier delays or container names.
|
||||||
|
|
||||||
|
Fix: `conf_cache_watchdog.sh` — refreshes the persistent backup from the RAM cache
|
||||||
|
every 15 minutes while the partner is offline, keeping it current throughout the outage.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ━━━ WHAT THIS FOLDER DOES ━━━
|
## ━━━ WHAT THIS FOLDER DOES ━━━
|
||||||
@@ -66,13 +75,17 @@ Four tiers. Fixed execution order via `watchdog_orchestrator.sh`.
|
|||||||
```
|
```
|
||||||
Pressure reduction resource_watchdog.sh — throttle/pause/stop before healing fails
|
Pressure reduction resource_watchdog.sh — throttle/pause/stop before healing fails
|
||||||
Container healing docker_watchdog.sh — memory, CPU, HTTP, required containers
|
Container healing docker_watchdog.sh — memory, CPU, HTTP, required containers
|
||||||
System components system_watchdog.sh — thin orchestrator: storage + WebGUI health
|
System components system_watchdog.sh — thin orchestrator: system component health
|
||||||
└─ System/ storage_watchdog.sh — pool growth rate + runaway log detection
|
└─ System/ storage_watchdog.sh — pool growth rate + runaway log detection
|
||||||
webgui_watchdog.sh — WebGUI availability, nginx → php-fpm → emhttp
|
|
||||||
network_watchdog.sh — internet, DDNS sync, Tailscale, NPM proxy
|
network_watchdog.sh — internet, DDNS sync, Tailscale, NPM proxy
|
||||||
|
conf_cache_watchdog.sh — maintain persistent partner conf backup
|
||||||
Last resort stability_watchdog.sh — reboot only when nothing else can recover
|
Last resort stability_watchdog.sh — reboot only when nothing else can recover
|
||||||
```
|
```
|
||||||
|
|
||||||
|
> `Plugin/unraid/Watchdogs/System/webgui_watchdog.sh` is also called by `system_watchdog.sh`
|
||||||
|
> but lives in the plugin tree because it calls Unraid-specific service commands
|
||||||
|
> (nginx, php-fpm, emhttp via platform adapter). See `Plugin/unraid/Watchdogs/System/`.
|
||||||
|
|
||||||
**The execution order is the design.** Resource pressure is reduced before docker_watchdog
|
**The execution order is the design.** Resource pressure is reduced before docker_watchdog
|
||||||
attempts restarts — containers restarted into a pressure-bound system just fail again.
|
attempts restarts — containers restarted into a pressure-bound system just fail again.
|
||||||
System component checks run after containers are healed. Stability watchdog runs last —
|
System component checks run after containers are healed. Stability watchdog runs last —
|
||||||
@@ -96,8 +109,16 @@ Orchestrators/
|
|||||||
|
|
||||||
Watchdogs/System/ ← called by system_watchdog.sh each cycle
|
Watchdogs/System/ ← called by system_watchdog.sh each cycle
|
||||||
storage_watchdog.sh pool growth rate + runaway log detection
|
storage_watchdog.sh pool growth rate + runaway log detection
|
||||||
webgui_watchdog.sh WebGUI availability — nginx → php-fpm → emhttp escalation
|
|
||||||
network_watchdog.sh internet reachability, DDNS sync, Tailscale, NPM proxy
|
network_watchdog.sh internet reachability, DDNS sync, Tailscale, NPM proxy
|
||||||
|
conf_cache_watchdog.sh maintain persistent partner conf backup during outages
|
||||||
|
|
||||||
|
Plugin/unraid/Watchdogs/System/
|
||||||
|
webgui_watchdog.sh WebGUI availability — nginx → php-fpm → emhttp escalation
|
||||||
|
(Unraid-specific platform calls — lives in plugin tree)
|
||||||
|
|
||||||
|
Plugin/unraid/System_Essentials/
|
||||||
|
unraid_api_key_renew.sh re-registers the Varaverk plugin API key each watchdog cycle
|
||||||
|
(inserted between docker_watchdog and stability_watchdog)
|
||||||
|
|
||||||
Tools/
|
Tools/
|
||||||
watchdog_skip_list_manager.sh ◄────────────── docker_watchdog.sh writes skip list
|
watchdog_skip_list_manager.sh ◄────────────── docker_watchdog.sh writes skip list
|
||||||
@@ -106,6 +127,10 @@ Tools/
|
|||||||
Docker_Essentials/
|
Docker_Essentials/
|
||||||
All container lifecycle scripts (daily restart, updates, network) — unaffected.
|
All container lifecycle scripts (daily restart, updates, network) — unaffected.
|
||||||
docker_watchdog.sh coordinates with them via shared state, not direct calls.
|
docker_watchdog.sh coordinates with them via shared state, not direct calls.
|
||||||
|
|
||||||
|
System_Essentials/
|
||||||
|
conf_cache_save.sh / conf_cache_restore.sh — at-stop/at-start bookends for the
|
||||||
|
persistent backup that conf_cache_watchdog.sh keeps current in between.
|
||||||
```
|
```
|
||||||
|
|
||||||
**watchdog_orchestrator.sh stays in Orchestrators/** — it's a job runner, not a watchdog.
|
**watchdog_orchestrator.sh stays in Orchestrators/** — it's a job runner, not a watchdog.
|
||||||
@@ -119,14 +144,15 @@ Docker_Essentials/
|
|||||||
|--------|------|-----------|
|
|--------|------|-----------|
|
||||||
| `resource_watchdog.sh` | Three-level pressure reduction — throttle, pause, stop | `watchdog_orchestrator.sh` — 1st every 15 min |
|
| `resource_watchdog.sh` | Three-level pressure reduction — throttle, pause, stop | `watchdog_orchestrator.sh` — 1st every 15 min |
|
||||||
| `docker_watchdog.sh` | Two-tier container healing — memory, CPU, HTTP, required | `watchdog_orchestrator.sh` — 2nd every 15 min |
|
| `docker_watchdog.sh` | Two-tier container healing — memory, CPU, HTTP, required | `watchdog_orchestrator.sh` — 2nd every 15 min |
|
||||||
| `system_watchdog.sh` | Thin orchestrator — runs System/ component watchdogs in sequence | `watchdog_orchestrator.sh` — 3rd every 15 min |
|
| `system_watchdog.sh` | Thin orchestrator — runs SYSTEM_WATCHDOG_SCRIPTS in sequence | `watchdog_orchestrator.sh` — 3rd every 15 min |
|
||||||
| `stability_watchdog.sh` | Last-resort server watchdog — reboots when healing has failed | `watchdog_orchestrator.sh` — 4th every 15 min |
|
| `stability_watchdog.sh` | Last-resort server watchdog — reboots when healing has failed | `watchdog_orchestrator.sh` — 4th every 15 min |
|
||||||
| `System/storage_watchdog.sh` | Pool growth rate + runaway log detection and remediation | `system_watchdog.sh` — every 15 min |
|
| `System/storage_watchdog.sh` | Pool growth rate + runaway log detection and remediation | `system_watchdog.sh` — every 15 min |
|
||||||
| `System/webgui_watchdog.sh` | WebGUI availability — nginx → php-fpm → emhttp escalation | `system_watchdog.sh` — every 15 min |
|
|
||||||
| `System/network_watchdog.sh` | Internet reachability, DDNS sync, Tailscale, NPM proxy | `system_watchdog.sh` — every 15 min |
|
| `System/network_watchdog.sh` | Internet reachability, DDNS sync, Tailscale, NPM proxy | `system_watchdog.sh` — every 15 min |
|
||||||
|
| `System/conf_cache_watchdog.sh` | Refresh persistent partner conf backup while partner is offline | `system_watchdog.sh` — every 15 min |
|
||||||
|
|
||||||
> `watchdog_orchestrator.sh` is in `Orchestrators/`. `watchdog_skip_list_manager.sh`
|
> `Plugin/unraid/Watchdogs/System/webgui_watchdog.sh` is also called by `system_watchdog.sh`
|
||||||
> is in `Tools/`. Neither is a watchdog — they sit at the edges of this system.
|
> but is not in this folder. `watchdog_orchestrator.sh` is in `Orchestrators/`.
|
||||||
|
> `watchdog_skip_list_manager.sh` is in `Tools/`. Neither is a watchdog.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -166,7 +192,7 @@ Every 15 minutes — watchdog_orchestrator.sh fires:
|
|||||||
│ log file scan: find *.log > WATCHDOG_APPDATA_LOG_MAX_GB
|
│ log file scan: find *.log > WATCHDOG_APPDATA_LOG_MAX_GB
|
||||||
│ oversize log found → 3-strike warn → truncate (if enabled) or alert
|
│ oversize log found → 3-strike warn → truncate (if enabled) or alert
|
||||||
│
|
│
|
||||||
│ └─ webgui_watchdog.sh
|
│ └─ Plugin/unraid/Watchdogs/System/webgui_watchdog.sh
|
||||||
│ curl check → WebGUI responding → exit 0 (silent)
|
│ curl check → WebGUI responding → exit 0 (silent)
|
||||||
│ not responding → nginx restart → wait → recheck
|
│ not responding → nginx restart → wait → recheck
|
||||||
│ still down → php-fpm restart → wait → recheck
|
│ still down → php-fpm restart → wait → recheck
|
||||||
@@ -179,6 +205,13 @@ Every 15 minutes — watchdog_orchestrator.sh fires:
|
|||||||
│ Tailscale: status Running → pass; not running = notify (no auto-restart)
|
│ Tailscale: status Running → pass; not running = notify (no auto-restart)
|
||||||
│ NPM proxy: curl external URL → 2-strike system → restart NginxProxyManager
|
│ NPM proxy: curl external URL → 2-strike system → restart NginxProxyManager
|
||||||
│
|
│
|
||||||
|
│ └─ conf_cache_watchdog.sh
|
||||||
|
│ remote online → remove persistent backup (conf_sync gets fresh on next boot)
|
||||||
|
│ remote offline → refresh backup from RAM cache → backup stays current
|
||||||
|
│ silent when remote is online and no backup exists (normal state)
|
||||||
|
│
|
||||||
|
│ (+ Plugin/unraid/System_Essentials/unraid_api_key_renew.sh — between steps 3 and 4)
|
||||||
|
│
|
||||||
Step 4 — stability_watchdog.sh
|
Step 4 — stability_watchdog.sh
|
||||||
checks the server itself — RAM, CPU temp, rootfs, FDs, kernel, daemon
|
checks the server itself — RAM, CPU temp, rootfs, FDs, kernel, daemon
|
||||||
Tier 1 CRITICAL → immediate reboot (no strikes)
|
Tier 1 CRITICAL → immediate reboot (no strikes)
|
||||||
@@ -195,7 +228,9 @@ Every 15 minutes — watchdog_orchestrator.sh fires:
|
|||||||
| `RW_STATE_FILE` | `resource_watchdog.sh` | `docker_watchdog.sh` | `mem_shutdown_active` flag — defer restarts during RAM emergency |
|
| `RW_STATE_FILE` | `resource_watchdog.sh` | `docker_watchdog.sh` | `mem_shutdown_active` flag — defer restarts during RAM emergency |
|
||||||
| `SYS_WATCHDOG_STATE_FILE` | `stability_watchdog.sh` | `docker_watchdog.sh` | `watchdog_cycle` heartbeat — stale guard (2hr timeout) |
|
| `SYS_WATCHDOG_STATE_FILE` | `stability_watchdog.sh` | `docker_watchdog.sh` | `watchdog_cycle` heartbeat — stale guard (2hr timeout) |
|
||||||
| `WATCHDOG_STATE_FILE` | `docker_watchdog.sh` | itself | CPU/HTTP strike counts per container |
|
| `WATCHDOG_STATE_FILE` | `docker_watchdog.sh` | itself | CPU/HTTP strike counts per container |
|
||||||
| `SYS_WATCHDOG_FAILED_FILE` | `docker_watchdog.sh` | `watchdog_skip_list_manager.sh` | Container skip list |
|
| `DOCKER_WATCHDOG_FAILED_FILE` | `docker_watchdog.sh` | `watchdog_skip_list_manager.sh` | Container skip list |
|
||||||
| `STORAGE_WATCHDOG_STATE_FILE` | `System/storage_watchdog.sh` | itself | Growth + log strike counts |
|
| `STORAGE_WATCHDOG_STATE_FILE` | `System/storage_watchdog.sh` | itself | Growth + log strike counts |
|
||||||
| `WATCHDOG_APPDATA_GROWTH_FILE` | `System/storage_watchdog.sh` | itself | Per-container size baseline for growth rate |
|
| `WATCHDOG_APPDATA_GROWTH_FILE` | `System/storage_watchdog.sh` | itself | Per-container size baseline for growth rate |
|
||||||
| `NETWORK_WATCHDOG_NPM_STATE_FILE` | `System/network_watchdog.sh` | itself | NPM proxy strike count |
|
| `NETWORK_WATCHDOG_NPM_STATE_FILE` | `System/network_watchdog.sh` | itself | NPM proxy strike count |
|
||||||
|
|
||||||
|
All state files are in `$STATE_DIR` (adapts to storage mode). See `master.conf` for actual variable values.
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
#
|
#
|
||||||
# PURPOSE
|
# PURPOSE
|
||||||
# ─────────────────────────────────────────────────────────────────────────────
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
# Maintains the persistent partner conf backup at /boot/config/.cache/vv/d/.
|
# Maintains the persistent partner conf backup at $PERSISTENT_CONF_CACHE.
|
||||||
# Runs every 15 minutes via SYSTEM_WATCHDOG_SCRIPTS.
|
# Runs every 15 minutes via SYSTEM_WATCHDOG_SCRIPTS.
|
||||||
#
|
#
|
||||||
# When remote is OFFLINE:
|
# When remote is OFFLINE:
|
||||||
@@ -33,7 +33,7 @@ require_partnership
|
|||||||
[[ -z "${REMOTE_ID:-}" ]] && exit 0
|
[[ -z "${REMOTE_ID:-}" ]] && exit 0
|
||||||
|
|
||||||
RAM_CACHE="/tmp/.cache/vv/d"
|
RAM_CACHE="/tmp/.cache/vv/d"
|
||||||
SAVE_DIR="/boot/config/.cache/vv/d"
|
SAVE_DIR="$PERSISTENT_CONF_CACHE"
|
||||||
|
|
||||||
if ping_remote; then
|
if ping_remote; then
|
||||||
if [[ -d "$SAVE_DIR" ]]; then
|
if [[ -d "$SAVE_DIR" ]]; then
|
||||||
|
|||||||
Reference in New Issue
Block a user