Compare commits

...
4 Commits
Author SHA1 Message Date
Gmer4Lfe e8b114094a Bring script headers onto the template and close safeguard gaps
Headers claimed protections the code never had, and several destructive paths had no
guard against a collapsed config value.
2026-08-01 20:37:59 -04:00
Gmer4Lfe cdce877601 Add a second GPU card to the monitor page
The 3080 was invisible because the stats parser only ever described GPU 0.
2026-08-01 20:37:58 -04:00
Gmer4Lfe 77b24e9823 Resolve network watchdog host config after detect_hosts
MY_ID is not exported and is unset until detect_hosts runs, so the DDNS and NPM checks
were silently resolving empty config and skipping themselves every cycle.
2026-08-01 20:37:58 -04:00
Gmer4Lfe 3f5ad22632 Only apply --delete on merge-run when the pull pass completed
A capped or failed pull leaves the remote holding content the local never received, so
deleting against it destroys the only copy.
2026-08-01 20:37:58 -04:00
85 changed files with 3578 additions and 327 deletions
+23
View File
@@ -30,6 +30,29 @@
# An arr not configured on this host (e.g. Lidarr is HOST1-only) is skipped cleanly. # An arr not configured on this host (e.g. Lidarr is HOST1-only) is skipped cleanly.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Warm the Cache, Never Own It
# This script only pre-populates what arr_get_tracked_data() would fetch on demand
# anyway. Nothing depends on it having run — every consumer still writes through on a
# cold cache. It removes latency and staleness, it is not a dependency.
#
# Failure Is a No-Op, Not an Error
# An arr that never comes up in time simply leaves its cache cold, exactly as if this
# script did not exist. That is why a missed prefill is logged rather than notified —
# the fallback path is the normal path.
#
# Wait Ceiling Matched to the Trigger
# The array-start run tolerates a long wait because containers are genuinely still
# starting. The 30-minute recurring run does not, because a live fetch takes seconds
# and a long wait there would only serve to overlap the next tick.
#
# Per-Arr Independence
# Each arr is prefilled on its own. One unconfigured or slow-starting arr never
# prevents the other two from being warmed.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+93
View File
@@ -62,6 +62,94 @@
# logged in Sonarr's history as "UserInvokedSearch"). # logged in Sonarr's history as "UserInvokedSearch").
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Detect Always, Act Only on Request
# A bare run probes and reports. Deleting a file the arr believes it has is a
# destructive act, so it requires --remediate explicitly. The scan can be scheduled
# weekly and read without any risk of it removing media on its own.
#
# One Probe Result Is Not Evidence
# ffprobe can fail for reasons that have nothing to do with the file — a mid-write
# import, an NFS blip, a container restart. Corruption must be observed
# CORRUPTION_SCAN_STRIKE_LIMIT times consecutively before remediation acts, and a
# single clean re-probe resets the counter.
#
# Skip-Cache Over Re-Probing
# At 90k+ tracked files a full re-probe every run is not viable. Files unchanged by
# mtime and size since they last verified clean are skipped, so each run spends its
# time on what actually changed rather than re-proving the library from scratch.
#
# Delete the Record, Let the Arr Re-Acquire
# Remediation removes the file record and explicitly triggers a search. The arr is
# left to obtain a good copy through its normal path — this script never tries to
# repair a file in place.
#
# Explicit Search, Not the Background Cycle
# The re-search is triggered directly rather than left to the arr's own missing-search
# cycle, because that cycle skips unmonitored items entirely and would silently leave
# an unmonitored corrupt file deleted and never replaced.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS
# ==============================================================================================
#
# Root Enforcement
# docker exec into the ffprobe container requires root.
#
# Lock Acquisition
# acquire_lock prevents overlapping runs. Two instances would both probe and could
# both count a strike against the same file, reaching the limit in half the intended
# number of observations.
#
# Host Detection
# detect_hosts() aliases the arr URLs, API keys and HOST*_FFPROBE_CONTAINER.
#
# jq Dependency Check
# Fails fast if jq is missing — the tracked-file lists and every hasFile verification
# are parsed with it.
#
# Report-Only Default
# Nothing is deleted without --remediate.
#
# ffprobe Configuration Check
# Exits cleanly if FFPROBE_CONTAINER / FFPROBE_BIN are unconfigured for this host.
#
# ffprobe Container Health Check
# check_container_health() verifies the container is running and healthy before any
# probing. Every probe is a docker exec into it — if it is stopped or unhealthy every
# exec fails, every file reads as corrupt, and two consecutive runs would clear the
# strike limit and hand --remediate the whole library to delete.
#
# API Reachability + Version Gate
# check_api then check_arr_version per arr. A version mismatch skips that arr rather
# than issuing deletes against an API whose file-record endpoints may have moved.
#
# Per-Arr Isolation
# Sonarr and Radarr run sequentially, and one failing, unconfigured or version-
# mismatched arr never blocks the other.
#
# Unmapped Path Skip
# Files whose arr-side path cannot be mapped into the ffprobe container's mount
# namespace are skipped and counted, never probed through a wrong path and never
# treated as corrupt because the probe could not see them.
#
# Strike Threshold
# CORRUPTION_SCAN_STRIKE_LIMIT consecutive corrupt detections are required before
# --remediate deletes anything. A transient ffprobe failure cannot trigger a delete,
# and a clean re-probe clears the counter.
#
# Post-Delete Verification
# The DELETE response is never trusted. hasFile is re-checked and must have flipped
# false before the re-search is issued, so a failed delete never leaves the arr
# searching for something it still believes it has.
#
# Targeted Deletion
# Only the specific episodefile/moviefile record for the corrupt file is removed —
# never the series, movie, or any sibling file.
#
# ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
@@ -174,6 +262,11 @@ if [[ -z "$FFPROBE_CONTAINER" || -z "$FFPROBE_BIN" ]]; then
exit 0 exit 0
fi fi
# Every probe is a docker exec into this container. If it is stopped or unhealthy, every
# exec fails, every file reads as corrupt, and two such runs would clear the strike limit
# and hand --remediate an entire library to delete. Abort before probing anything.
check_container_health "$FFPROBE_CONTAINER" "${DOCKER_TIMEOUT:-30}" "Arr Corruption Scan"
CORRUPTION_SCAN_STATE_FILE="${CORRUPTION_SCAN_STATE_FILE:-$DATA_DIR/corruption_scan_state.tsv}" CORRUPTION_SCAN_STATE_FILE="${CORRUPTION_SCAN_STATE_FILE:-$DATA_DIR/corruption_scan_state.tsv}"
mkdir -p "$(dirname "$CORRUPTION_SCAN_STATE_FILE")" mkdir -p "$(dirname "$CORRUPTION_SCAN_STATE_FILE")"
touch "$CORRUPTION_SCAN_STATE_FILE" touch "$CORRUPTION_SCAN_STATE_FILE"
+105 -17
View File
@@ -39,7 +39,8 @@
# age gate an entry the arr cannot even name is not going to import: if the # age gate an entry the arr cannot even name is not going to import: if the
# title is in the library and monitored, clearing it lets the arr search a # title is in the library and monitored, clearing it lets the arr search a
# copy it can actually parse; if it is not in the library, nothing is # copy it can actually parse; if it is not in the library, nothing is
# tracking it and it is dead weight either way. Guarded — see safeguard 6 # tracking it and it is dead weight either way. Guarded — see Non-Empty
# Library Requirement below
# HELD — IMPORTABLE entries the arr keeps refusing (XEM-blocked, season-span # HELD — IMPORTABLE entries the arr keeps refusing (XEM-blocked, season-span
# files), and everything skipped by a guard → report only, human call # files), and everything skipped by a guard → report only, human call
# #
@@ -47,27 +48,114 @@
# there is no way to tell tracked from orphaned, and guessing means deleting active imports. # there is no way to tell tracked from orphaned, and guessing means deleting active imports.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Classify Before Acting
# Every entry is placed in exactly one class before anything is deleted, and each class
# has its own justification. Nothing is removed because it merely looked unwanted — it is
# removed because it matched a category whose deletion rationale is written down above.
#
# The Queue Is the Source of Truth
# Tracked-vs-orphaned is decided by the arr's own queue, never inferred from filenames or
# timestamps. If the queue cannot be read, the arr is skipped entirely rather than
# falling back to a weaker signal — a guess here deletes an active import.
#
# Deleting Is Recoverable, Deleting Wrong Is Not
# The classes that get deleted are ones the arr can re-acquire: junk it could never
# import, content the library already has, and entries it cannot even name. Anything
# whose loss would be permanent or ambiguous is held and reported for a human instead.
#
# Abnormal Volume Means Broken Input
# The delete cap exists because the realistic failure mode is bad input, not bad logic —
# a partial queue fetch classifies live downloads as orphans, and the only visible
# symptom is an unusually large delete total. The cap turns that into a stop.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# 1. DOWNLOAD_ORPHAN_CLEANER_ENABLED master toggle # Root Enforcement
# 2. Download dir must exist and live under /mnt/ — refuses to walk anything else # Download folders are written by container users; removing them requires root.
# 3. Queue fetch must succeed (see above) #
# 4. Age gate — nothing under DOWNLOAD_ORPHAN_AGE days is touched # Lock Acquisition
# 5. Deletion only for JUNK, parse-verified REDUNDANT, and UNMATCHED — never for # acquire_lock "wait" with an EXIT trap releasing all locks, so an interrupted run never
# IMPORTABLE or anything a guard has held # strands a lock and the daily orchestrator is never silently skipped.
# 6. UNMATCHED deletes require the arr to report a non-empty library. An empty or #
# restoring database answers every parse with "no match", which would condemn the # Host Detection
# entire download dir; the library is queried directly rather than inferred from the # detect_hosts() runs before any HOST*_-prefixed download dir is resolved.
# run's own matches, since a small batch that is legitimately all-unmatched is normal #
# once daily runs have caught up and would otherwise read as a broken database # DOWNLOAD_ORPHAN_CLEANER_ENABLED Toggle
# 7. Run total over DOWNLOAD_ORPHAN_MAX_DELETE_GB aborts the delete pass and notifies — # Master switch — exits cleanly when disabled.
# a queue fetch that returned partial data would classify live downloads as orphans, #
# and an abnormally large delete total is the visible symptom of exactly that # Download Path Restriction
# (--i-know-what-im-doing overrides, e.g. for a first run against a known backlog) # The download dir must exist and live under /mnt/. Anything else is refused rather
# than walked, so a blank or malformed path can never point the scan at the filesystem
# root or a system directory.
#
# Queue Fetch Hard Gate
# The arr is skipped entirely if its queue cannot be read. Without the queue there is no
# way to distinguish tracked from orphaned, and guessing deletes active imports.
#
# Age Gate
# Nothing under DOWNLOAD_ORPHAN_AGE days is touched, so an entry mid-import is never a
# deletion candidate regardless of how it classifies.
#
# Deletion Class Restriction
# Only JUNK, parse-verified REDUNDANT and UNMATCHED are deleted. IMPORTABLE entries and
# anything a guard has held are reported, never removed.
#
# Non-Empty Library Requirement
# UNMATCHED deletions require the arr to report a non-empty library. An empty or
# restoring database answers every parse with "no match", which would condemn the whole
# download dir. The library is queried directly rather than inferred from this run's own
# match rate — a small batch that is legitimately all-unmatched is normal once daily runs
# have caught up, and would otherwise read as a broken database.
#
# Delete Volume Cap
# A run total over DOWNLOAD_ORPHAN_MAX_DELETE_GB aborts the delete pass and notifies.
# --i-know-what-im-doing overrides it for a deliberate first run against a known backlog.
#
# Dry Run Support
# --dry-run classifies everything and reports, deleting and importing nothing.
# #
# ============================================================================================== # ==============================================================================================
# USAGE # CONFIGURATION
# ==============================================================================================
#
# host*.conf (resolved per host after detect_hosts())
#
# HOST*_SONARR_DOWNLOAD_DIR / HOST*_RADARR_DOWNLOAD_DIR
# Host-side path to the arr's completed-download folder. Absent means that arr's
# cleanup is skipped, not an error.
#
# HOST*_SONARR_DOWNLOAD_CONTAINER_DIR / HOST*_RADARR_DOWNLOAD_CONTAINER_DIR
# The same folder as the arr container sees it — used when triggering the
# DownloadedEpisodesScan / DownloadedMoviesScan path.
#
# SONARR_URL / SONARR_API_KEY / RADARR_URL / RADARR_API_KEY
# Aliased by detect_hosts(). A missing URL or key skips that arr.
#
# master.conf
#
# DOWNLOAD_ORPHAN_CLEANER_ENABLED
# Master toggle (default: true)
#
# DOWNLOAD_ORPHAN_AGE
# Days before an entry is eligible at all — younger entries may be mid-import
# (default: 7)
#
# DOWNLOAD_ORPHAN_MIN_VIDEO_MB
# An entry with no video file above this size is JUNK (default: 50)
#
# DOWNLOAD_ORPHAN_MAX_DELETE_GB
# Abort the delete pass if the run total exceeds this (default: 100)
#
# SONARR_EXTENSIONS / RADARR_EXTENSIONS
# Video extensions used to decide whether an entry contains real media
#
# ==============================================================================================
# RUNTIME MODES
# ============================================================================================== # ==============================================================================================
# #
# arr_download_orphan_cleaner.sh — daily orchestrator entry # arr_download_orphan_cleaner.sh — daily orchestrator entry
+90 -13
View File
@@ -20,22 +20,91 @@
# and running them concurrently would just contend for the same disk I/O for no benefit. # and running them concurrently would just contend for the same disk I/O for no benefit.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Lidarr, then Sonarr, then Radarr — strictly sequential. Per arr:
#
# 1. Reachability
# → check_api; unreachable skips this arr only
#
# 2. Already-scanning check
# → a rescan already active (manual, or another script) means skip rather than
# stack a second full-disk walk on top of it
#
# 3. Capture the before count
# → tracked file count read from the arr's own stats
#
# 4. Trigger the rescan command
# → RescanFolders (Lidarr) / RescanSeries (Sonarr) / RescanMovie (Radarr)
#
# 5. Poll to completion
# → bounded by ARR_FULL_RESCAN_TIMEOUT
#
# 6. Report the delta
# → before vs after tracked count, so drift that was corrected is visible
#
# ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Proactive, Not Reactive
# check_tracked_count_floor() catches stat drift reactively, at the moment some other
# script is about to act on bad numbers. This job exists so that drift is corrected on a
# schedule instead of being discovered by whichever cleanup happens to trip over it first.
#
# Sequential by Design
# Each rescan is a full-disk walk. Running three concurrently contends for the same
# spindles and finishes no sooner, so the arrs are never parallelised — the slowness is
# accepted deliberately rather than optimised into I/O thrash.
#
# Never Stack a Scan
# An already-running rescan is left alone rather than duplicated. A second concurrent
# walk of the same library doubles the I/O cost and returns nothing the first will not.
#
# Per-Arr Isolation
# One arr being down, slow, or already scanning must never prevent the other two from
# being reconciled. Partial coverage beats a skipped run.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root required — required for consistency across Arrs_Stack/, no direct fs writes here # Root Enforcement
# acquire_lock "wait" — waits for a prior run to finish rather than skipping or colliding; # Kept for consistency across Arrs_Stack/ — this script makes no direct filesystem writes.
# a full-library rescan across three arrs can run long, worth queuing #
# behind rather than silently no-op'ing # Lock Acquisition
# Reachability check — check_api before touching an arr; unreachable → skip that arr only # acquire_lock "wait" — waits for a prior run rather than skipping or colliding. A full
# Active-rescan check — skips triggering a NEW rescan if one's already active on that arr # rescan across three arrs runs long and is worth queuing behind, not silently dropping.
# (manual trigger, another script) — never stacks a duplicate scan, #
# see Tools/arr_rescan_monitor.sh for catching up that arr's cache # Host Detection
# once the pre-existing scan finishes instead of waiting for next week # detect_hosts() aliases each arr's URL and API key.
# Sequential only — never runs two arrs' rescans in parallel; each is a heavy full-disk #
# walk and concurrent walks would just contend for the same disk I/O # jq Dependency Check
# Per-arr isolation — one arr failing, timing out, or being skipped never blocks the others # Fails fast if jq is missing. Both the before/after tracked counts and the command
# --dry-run mode — reports which arrs would be rescanned, triggers nothing # payload are built with jq — without it the counts read empty and every delta would be
# reported as if nothing changed.
#
# Reachability Check
# check_api before touching an arr; unreachable skips that arr only.
#
# Active-Rescan Check
# Skips triggering a new rescan if one is already active on that arr, so a duplicate
# full-disk walk is never stacked. See Tools/arr_rescan_monitor.sh for catching that
# arr's cache up once the pre-existing scan finishes, rather than waiting a week.
#
# Sequential Only
# Two arrs' rescans never run in parallel.
#
# Per-Arr Isolation
# One arr failing, timing out, or being skipped never blocks the others.
#
# Timeout Bound
# ARR_FULL_RESCAN_TIMEOUT caps the wait per arr, so a rescan that never completes cannot
# hold the weekly window open indefinitely.
#
# Dry Run Support
# --dry-run reports which arrs would be rescanned and triggers nothing.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -72,6 +141,14 @@ if [[ "$EUID" -ne 0 ]]; then
exit 1 exit 1
fi fi
# Both the tracked-count reads and the command payload are built with jq — without it the
# counts read empty and every arr would report a zero delta as if nothing had drifted.
if ! command -v jq >/dev/null 2>&1; then
error "jq not found — required for JSON parsing"
notify "Arr full rescan failed on $(hostname) — jq not installed" "Arr Full Rescan" "warning"
exit 1
fi
acquire_lock "wait" acquire_lock "wait"
detect_hosts detect_hosts
-1
View File
@@ -88,7 +88,6 @@
# jq + curl validation — exits if either tool missing # jq + curl validation — exits if either tool missing
# ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT) # ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT)
# Duplicate detection — temp file of tracked paths, grep before delete # Duplicate detection — temp file of tracked paths, grep before delete
# platform_require_cmd — notify script validated before use
# Silent by default — orphans/junk warn(), clean library logs silently # Silent by default — orphans/junk warn(), clean library logs silently
# #
# ============================================================================================== # ==============================================================================================
@@ -52,6 +52,56 @@
# output. Only ambiguous (overlapping-album) pairs produce a notification. # output. Only ambiguous (overlapping-album) pairs produce a notification.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL SAFEGUARDS
# ==============================================================================================
#
# Root Enforcement
# Required by the container interaction and state writes.
#
# Lock Acquisition
# acquire_lock prevents overlapping runs racing on the same artist IDs.
#
# Host Detection
# detect_hosts() aliases LIDARR_URL / LIDARR_API_KEY.
#
# jq Dependency Check
# Fails fast if jq is missing — duplicate detection and every file-count read
# depend on it, and an absent jq would evaluate counts to empty and make every
# artist look like a zero-file phantom.
#
# API Reachability + Version Gate
# check_api then check_arr_version against LIDARR_VERSION_MAJOR before any read.
#
# Empty Library Abort
# A response of 0 artists aborts. An empty list is indistinguishable from
# "no duplicates" and must never be read as a clean result.
#
# Tracked-Count Floor
# check_tracked_count_floor against the baseline shared with lidarr_cleanup.sh.
# A library-wide desync — mid full-rescan, for example — makes trackFileCount read
# far below reality for many artists at once. Confirmed 2026-07-16: without this,
# both sides of a genuinely-real duplicate (ROMES) read as 0-file phantoms and the
# wrong one would have been deleted. Deliberately reuses lidarr_cleanup.sh's own
# baseline so every script depending on tracked counts shares one answer to "is
# Lidarr's data trustworthy right now" rather than forming a separate opinion.
#
# Files Never Deleted
# Removal passes deleteFiles=false. Only the phantom Lidarr entry is dropped;
# nothing on disk is touched, so a wrong call costs a re-add, not media.
#
# Phantom-Only Deletion
# Only the zero-file side of a duplicate pair is ever removed. If both sides hold
# files, or neither does unambiguously, the pair is flagged for manual review
# instead — the script never picks a winner between two real artists.
#
# Manual Review Reporting
# Flagged pairs are named in the summary and notification so an ambiguous
# duplicate surfaces as a decision to make rather than disappearing silently.
#
# Dry Run Support
# --dry-run reports every deletion and exclusion and performs none.
#
# ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
+19 -4
View File
@@ -16,7 +16,7 @@
# Goal: 05 meaningful Lidarr adds per week, not bulk imports. # Goal: 05 meaningful Lidarr adds per week, not bulk imports.
# #
# ============================================================================================== # ==============================================================================================
# FLOW # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# 1. Fetch play completions from Emby activity log (last LOOKBACK_DAYS days) # 1. Fetch play completions from Emby activity log (last LOOKBACK_DAYS days)
@@ -58,11 +58,26 @@
# Max adds: LIDARR_DISCOVERY_MAX_ADDS (default 5) caps each stage # Max adds: LIDARR_DISCOVERY_MAX_ADDS (default 5) caps each stage
# #
# ============================================================================================== # ==============================================================================================
# REQUIREMENTS # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# Last.fm API key — required for both stages # host*.conf (aliased by detect_hosts())
# Configure HOST*_LASTFM_API_KEY in host*.conf #
# HOST*_LASTFM_API_KEY Required for both stages — without it the run exits cleanly
# rather than adding anything unscored.
# LIDARR_URL / LIDARR_API_KEY Target arr
# EMBY_URL / EMBY_API_KEY Play history source for seed selection
#
# master.conf
#
# LIDARR_DISCOVERY_THRESHOLD Score required to accept a candidate (0-100)
# LIDARR_DISCOVERY_LOOKBACK_DAYS Emby play history window
# LIDARR_DISCOVERY_MIN_PLAYS Min plays in the window before an artist is evaluated
# LIDARR_DISCOVERY_USER_CAP_PCT Max % of the play score any one user can contribute,
# so a single heavy listener cannot drive the library
# LIDARR_DISCOVERY_MAX_ADDS Hard cap on artists added per run
# LIDARR_DISCOVERY_REJECT_COOLDOWN Days before a rejected artist is re-evaluated
# LIDARR_DISCOVERY_HISTORY Decision history DB — accepted and rejected
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
+21 -5
View File
@@ -16,7 +16,7 @@
# Goal: 05 meaningful Radarr adds per run, not bulk imports. # Goal: 05 meaningful Radarr adds per run, not bulk imports.
# #
# ============================================================================================== # ==============================================================================================
# FLOW # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# 1. Fetch recently watched movies from Emby (SEED_LIBRARIES, last LOOKBACK_DAYS days) # 1. Fetch recently watched movies from Emby (SEED_LIBRARIES, last LOOKBACK_DAYS days)
@@ -56,12 +56,28 @@
# Max adds: RADARR_DISCOVERY_MAX_ADDS (default 5) # Max adds: RADARR_DISCOVERY_MAX_ADDS (default 5)
# #
# ============================================================================================== # ==============================================================================================
# REQUIREMENTS # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# TMDB API key — required for Stage 2 recommendations # host*.conf (aliased by detect_hosts())
# Configure HOST*_TMDB_API_KEY in host*.conf #
# Free key at: https://www.themoviedb.org/settings/api # HOST*_TMDB_API_KEY Required for Stage 2 recommendations. Free key at
# https://www.themoviedb.org/settings/api — without it the run
# exits cleanly rather than adding anything unscored.
# RADARR_URL / RADARR_API_KEY Target arr
# EMBY_URL / EMBY_API_KEY Play history source for seed selection
#
# master.conf
#
# RADARR_DISCOVERY_THRESHOLD Score required to accept a candidate (0-100)
# RADARR_DISCOVERY_LOOKBACK_DAYS Emby watch history window
# RADARR_DISCOVERY_MAX_SEEDS Max seed movies taken from Stage 1
# RADARR_DISCOVERY_MAX_ADDS Hard cap on movies added per run
# RADARR_DISCOVERY_MIN_VOTE_COUNT Min TMDB votes for a candidate to be considered
# RADARR_DISCOVERY_MIN_RATING Min TMDB vote_average × 10
# RADARR_DISCOVERY_REJECT_COOLDOWN Days before a rejected movie is re-evaluated
# RADARR_DISCOVERY_SEED_LIBRARIES Emby libraries to draw seed movies from
# RADARR_DISCOVERY_HISTORY Decision history DB — accepted and rejected
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
+22 -5
View File
@@ -17,7 +17,7 @@
# Goal: 03 meaningful Sonarr adds per run, not bulk imports. # Goal: 03 meaningful Sonarr adds per run, not bulk imports.
# #
# ============================================================================================== # ==============================================================================================
# FLOW # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# 1. Fetch all Series from Emby SONARR_EMBY_LIBRARIES — build TMDB+TVDB index # 1. Fetch all Series from Emby SONARR_EMBY_LIBRARIES — build TMDB+TVDB index
@@ -63,12 +63,29 @@
# Max adds: SONARR_DISCOVERY_MAX_ADDS (default 3) — TV is a larger commitment than movies # Max adds: SONARR_DISCOVERY_MAX_ADDS (default 3) — TV is a larger commitment than movies
# #
# ============================================================================================== # ==============================================================================================
# REQUIREMENTS # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# TMDB API key — required for Stage 2 recommendations and external_ids lookup # host*.conf (aliased by detect_hosts())
# Configure HOST*_TMDB_API_KEY in host*.conf #
# Free key at: https://www.themoviedb.org/settings/api # HOST*_TMDB_API_KEY Required for Stage 2 recommendations and external_ids lookup.
# Free key at https://www.themoviedb.org/settings/api — without it
# the run exits cleanly rather than adding anything unscored.
# SONARR_URL / SONARR_API_KEY Target arr
# EMBY_URL / EMBY_API_KEY Play history source for seed selection
#
# master.conf
#
# SONARR_DISCOVERY_THRESHOLD Score required to accept a candidate (0-100)
# SONARR_DISCOVERY_LOOKBACK_DAYS Emby episode play history window
# SONARR_DISCOVERY_MAX_SEEDS Max seed series taken from Stage 1
# SONARR_DISCOVERY_MAX_ADDS Hard cap on shows added per run
# SONARR_DISCOVERY_MIN_VOTE_COUNT Min TMDB votes for a candidate to be considered
# SONARR_DISCOVERY_MIN_RATING Min TMDB vote_average × 10
# SONARR_DISCOVERY_REJECT_COOLDOWN Days before a rejected show is re-evaluated
# SONARR_DISCOVERY_USER_EPISODE_CAP Max episodes one user contributes to seed volume
# SONARR_DISCOVERY_MONITOR_MODE Sonarr monitor mode on add
# SONARR_DISCOVERY_HISTORY Decision history DB — accepted and rejected
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
+99 -6
View File
@@ -15,8 +15,11 @@
# FORWARD — a movie classified as anime/kids is sitting outside its dedicated root # FORWARD — a movie classified as anime/kids is sitting outside its dedicated root
# REVERSE — a movie sitting inside the kids/anime root doesn't match that classification # REVERSE — a movie sitting inside the kids/anime root doesn't match that classification
# #
# Report-only. No files are moved and no Radarr API writes happen — this is a detection # Report-only by default — no files are moved and no Radarr API writes happen unless a
# tool. Every rule below was validated against this library's real data before being # mode flag is given. Pass --move to relocate forward misplacements, or --remove-junk to
# delete and import-exclude bad-metadata entries (see OPERATIONAL MODEL below); without
# those flags this is purely a detection tool. Every rule below was validated against
# this library's real data before being
# adopted (see master.conf comments above the curated lists) — this is not a generic # adopted (see master.conf comments above the curated lists) — this is not a generic
# genre-matcher, it's tuned specifically against the false-positive traps that showed up # genre-matcher, it's tuned specifically against the false-positive traps that showed up
# when testing looser rules (documented per-rule below). # when testing looser rules (documented per-rule below).
@@ -52,13 +55,38 @@
# blocklist and likely nothing legitimate to redownload under that exact TMDb match. # blocklist and likely nothing legitimate to redownload under that exact TMDb match.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Report pass (always):
# check_api → check_arr_version → arr_get_tracked_data (cache-first, one call)
# → classify every movie → report FORWARD, REVERSE and JUNK findings
#
# Remove-junk pass (--remove-junk, runs first when combined with --move):
# One entry at a time, halting on the first failure.
# DELETE with deleteFiles=false and addImportExclusion=true — the Radarr entry is
# removed and blocked from re-adding, files on disk are never touched. is_junk
# requires hasFile == false, so there is no file behind these entries anyway.
# Verified by re-fetching and requiring a 404 before counting as removed.
#
# Move pass (--move):
# One movie at a time, verified after each.
# hasFile == true → moveFiles=true, then poll the async MoveMovie command to
# "completed" (bounded by RADARR_MOVE_POLL_TIMEOUT) before the
# DB-field check — the DB flips instantly while the physical
# move is still queued.
# hasFile == false → correct rootFolderPath/path and trigger MoviesSearch instead;
# there is nothing to move.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
# Report, Don't Act # Report by Default, Act Only on Request
# This script never calls Radarr's write API and never touches a file. Every finding is # A bare run never calls Radarr's write API and never touches a file — every finding is
# a candidate for a human decision — moving media and re-pointing Radarr's tracking is a # just a candidate. Acting on them requires an explicit --move or --remove-junk flag, so
# separate, deliberate follow-up action, not something this scan does automatically. # the scan can be scheduled and re-run freely while the curated lists are being tuned
# without any risk of it rearranging the library on its own.
# #
# Curated Lists, Not Bare Genre/Cert Matching # Curated Lists, Not Bare Genre/Cert Matching
# Every signal used here failed at least once as a bare/standalone check during rule # Every signal used here failed at least once as a bare/standalone check during rule
@@ -72,6 +100,71 @@
# if the shared cache is warm) regardless of library size. # if the shared cache is warm) regardless of library size.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL SAFEGUARDS
# ==============================================================================================
#
# Root Enforcement
# Required by the container interaction and state writes.
#
# Lock Acquisition
# acquire_lock, plus acquire_lock "wait" around the write passes with an EXIT trap
# releasing all locks, so an interrupted run never leaves a lock behind.
#
# Host Detection
# detect_hosts() aliases RADARR_URL / RADARR_API_KEY / the root literals.
#
# curl + jq Dependency Check
# Fails fast if either is missing — every classification signal is parsed with jq.
#
# Report-Only Default
# No write happens without --move or --remove-junk.
#
# Required Var Check
# require_var on RADARR_URL and RADARR_API_KEY before any request.
#
# API Reachability + Version Gate
# check_api then check_arr_version against RADARR_VERSION_MAJOR. A major version bump
# can move or rename the fields every rule depends on, so a mismatch aborts rather
# than classifying against an unknown schema.
#
# Empty Library Abort
# A response of 0 movies aborts — an empty list is indistinguishable from a clean
# library and would otherwise report success during an API fault.
#
# Unconfigured Root Skip
# A blank kids/anime root skips that category's checks rather than comparing paths
# against an empty string.
#
# Files Never Deleted
# Junk removal passes deleteFiles=false. Only the Radarr entry is removed, and
# addImportExclusion=true stops it being re-added. is_junk additionally requires
# hasFile == false, so these entries have nothing on disk in the first place.
#
# One At A Time, Stop On First Failure
# Both write passes process one entry at a time and halt on the first failure rather
# than continuing through the library.
#
# Post-Write Verification
# Removal is confirmed by re-fetching and requiring a 404. Moves are confirmed by
# re-fetching and checking rootFolderPath and hasFile. The API response alone is
# never treated as proof.
#
# Async Move Completion Polling
# moveFiles=true flips the DB instantly while the physical move is a separate async
# MoveMovie command. Each move polls its own command to "completed" (bounded by
# RADARR_MOVE_POLL_TIMEOUT) before the DB-field check, so a batch cannot report
# everything moved while files are still queued at the old path.
#
# Junk Vote Threshold
# RADARR_JUNK_MIN_VOTES gates junk detection alongside hasFile == false and a null
# imdbId. All three must hold — a thin-metadata entry that actually has a file, or
# has an IMDb ID, is never treated as junk.
#
# Post-Write Cache Refresh
# The tracked-data cache is refreshed after writes so no other arr script reads a
# stale rootFolderPath or a movie that no longer exists.
#
# ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
-1
View File
@@ -98,7 +98,6 @@
# jq + curl validation — exits if either tool missing # jq + curl validation — exits if either tool missing
# ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT) # ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT)
# notify_emby_scan() — triggers Emby clean after deletion # notify_emby_scan() — triggers Emby clean after deletion
# platform_require_cmd — notify script validated before use
# Silent by default — orphans/junk warn(), clean library logs silently # Silent by default — orphans/junk warn(), clean library logs silently
# #
# ============================================================================================== # ==============================================================================================
+71 -1
View File
@@ -61,9 +61,15 @@
# no Sonarr equivalent for the same reason. # no Sonarr equivalent for the same reason.
# #
# ============================================================================================== # ==============================================================================================
# MOVE MODE (--move) # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# Report pass (always):
# check_api → check_arr_version → arr_get_tracked_data (cache-first, one call)
# → classify every series → report FORWARD and REVERSE disagreements → exit
#
# Move pass (--move only), described in detail below:
#
# Acts on FORWARD misplacements (classified anime/kids, sitting in the wrong root) and on # Acts on FORWARD misplacements (classified anime/kids, sitting in the wrong root) and on
# REVERSE-KIDS leaks (adult certification sitting in the kids root — moved back to # REVERSE-KIDS leaks (adult certification sitting in the kids root — moved back to
# SONARR_GENERAL_ROOT). Does NOT act on REVERSE-ANIME leaks — those are genuine judgment # SONARR_GENERAL_ROOT). Does NOT act on REVERSE-ANIME leaks — those are genuine judgment
@@ -93,6 +99,70 @@
# of library size. Refreshed after --move writes so no other script reads stale data. # of library size. Refreshed after --move writes so no other script reads stale data.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL SAFEGUARDS
# ==============================================================================================
#
# Root Enforcement
# Required by the container interaction and state writes.
#
# Lock Acquisition
# acquire_lock prevents a scheduled run overlapping a manual --move.
#
# Host Detection
# detect_hosts() aliases SONARR_URL / SONARR_API_KEY / the root literals.
#
# curl + jq Dependency Check
# Fails fast if either is missing — every classification signal is parsed with jq,
# and a missing jq would evaluate each signal to empty and classify nothing.
#
# Report-Only Default
# Nothing is written to Sonarr without --move. The scan is safe to schedule and
# safe to run repeatedly while tuning the curated lists.
#
# Required Var Check
# require_var on SONARR_URL and SONARR_API_KEY before any request.
#
# API Reachability + Version Gate
# check_api then check_arr_version against SONARR_VERSION_MAJOR. A major version
# bump can move or rename the fields every rule here depends on, so a mismatch
# aborts rather than classifying against an unknown schema.
#
# Empty Library Abort
# A response of 0 series aborts. An empty list is indistinguishable from "nothing
# is misplaced" and would otherwise report a clean library during an API fault.
#
# Unconfigured Root Skip
# A blank SONARR_GENERAL_ROOT / KIDS_ROOT / ANIME_ROOT skips that category's
# checks rather than erroring — a host with no dedicated root is a valid setup,
# and a blank value must never be compared against as if it were a real path.
#
# One At A Time, Stop On First Failure
# Series are moved individually and the batch halts on the first failure rather
# than continuing. A misclassified root or a failing move is a condition to
# review, not to repeat across the library.
#
# Async Move Completion Polling
# moveFiles=true flips the DB instantly while the physical move is a separate
# async MoveSeries command Sonarr drains one at a time. Each move locates its own
# command and polls it to "completed" (bounded by SONARR_MOVE_POLL_TIMEOUT) before
# anything else is checked. Without this a batch reports every series moved while
# the files are still queued at the old path — and downstream orphan cleanup can
# act on that gap.
#
# Post-Move Re-Verification
# The PUT response is never trusted. The series is re-fetched and both
# rootFolderPath and episodeFileCount are confirmed against expectations before
# the move counts as successful.
#
# Reverse-Anime Leaks Excluded From Moves
# Deliberate style placements (Western/Chinese animation grouped with anime by
# choice) legitimately sit in the anime root. Those are reported, never moved.
#
# Post-Write Cache Refresh
# The tracked-data cache is refreshed after --move writes so no other arr script
# reads a stale rootFolderPath.
#
# ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
-1
View File
@@ -92,7 +92,6 @@
# jq + curl validation — exits if either tool missing # jq + curl validation — exits if either tool missing
# ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT) # ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT)
# notify_emby_scan() — triggers Emby clean after deletion # notify_emby_scan() — triggers Emby clean after deletion
# platform_require_cmd — notify script validated before use
# Silent by default — orphans/junk warn(), clean library logs silently # Silent by default — orphans/junk warn(), clean library logs silently
# #
# ============================================================================================== # ==============================================================================================
+65 -7
View File
@@ -67,13 +67,38 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# WEBHOOK_PORT=0 gate — exits cleanly before any setup if the listener is disabled # Root Enforcement
# acquire_lock "continuous" — exits 0 cleanly if a healthy instance is already running, # Writes /var/log/varaverk and persists the generated secret into master.conf.
# instead of relaunching into a port conflict #
# Secret auto-generate — WEBHOOK_SECRET generated via openssl rand if empty; # WEBHOOK_PORT=0 Gate
# persisted to master.conf immediately so restarts reuse it # Exits cleanly before any setup if the listener is disabled.
# Shared secret gate — webhook URL must include ?key=<WEBHOOK_SECRET>; #
# requests without a valid key are rejected by the Node.js server # node Presence Check
# node is exec'd at the end of this script. Checking up front fails with a clear
# reason at array start rather than an exec error buried in the log after setup.
#
# openssl Presence Check
# Checked before attempting to generate a secret, on the path that needs it.
#
# Lock Acquisition
# acquire_lock "continuous" exits 0 cleanly if a healthy instance is already
# listening, instead of relaunching into an EADDRINUSE port conflict. An array
# stop/start without a full reboot leaves the old node process alive, and without
# this array_started.sh would log a false failure.
#
# Secret Auto-Generation
# WEBHOOK_SECRET generated via openssl rand if empty, and persisted to master.conf
# so restarts reuse it.
#
# Secret Persistence Verification
# The write-back is confirmed by re-reading master.conf. If it did not land, the
# secret would exist only in this process and be regenerated on the next start,
# silently invalidating the key already registered in the arrs — so this aborts
# loudly rather than starting with a secret that will not survive a restart.
#
# Shared Secret Gate
# The webhook URL must include ?key=<WEBHOOK_SECRET>; requests without a valid key
# are rejected by the Node.js server.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -104,11 +129,25 @@ ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
source "$ECOSYSTEM_ROOT/load_config.sh" source "$ECOSYSTEM_ROOT/load_config.sh"
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
[[ "${WEBHOOK_PORT:-0}" -eq 0 ]] && { [[ "${WEBHOOK_PORT:-0}" -eq 0 ]] && {
echo "[webhook] WEBHOOK_PORT=0 — listener disabled" echo "[webhook] WEBHOOK_PORT=0 — listener disabled"
exit 0 exit 0
} }
# node is exec'd at the end of this script — checking here fails with a clear reason at
# array start instead of an exec error buried in the log after all the setup has run.
if ! command -v node >/dev/null 2>&1; then
error "node not found — required to run webhook_listener.js"
notify "Webhook listener failed to start on $(hostname) — node not installed" \
"Webhook Listener" "warning"
exit 1
fi
# Skip gracefully if a healthy instance is already listening — otherwise a # Skip gracefully if a healthy instance is already listening — otherwise a
# restart that doesn't kill the old node process (array stop/start without a # restart that doesn't kill the old node process (array stop/start without a
# full reboot) hits EADDRINUSE and array_started.sh logs a false failure. # full reboot) hits EADDRINUSE and array_started.sh logs a false failure.
@@ -116,10 +155,29 @@ acquire_lock "continuous"
# ── Auto-generate secret if not yet set ───────────────────────────────────── # ── Auto-generate secret if not yet set ─────────────────────────────────────
if [[ -z "${WEBHOOK_SECRET:-}" ]]; then if [[ -z "${WEBHOOK_SECRET:-}" ]]; then
if ! command -v openssl >/dev/null 2>&1; then
error "openssl not found — cannot generate WEBHOOK_SECRET"
notify "Webhook listener failed to start on $(hostname) — openssl not installed" \
"Webhook Listener" "warning"
exit 1
fi
GENERATED=$(openssl rand -hex 32) GENERATED=$(openssl rand -hex 32)
MASTER_CONF="$ECOSYSTEM_ROOT/Configurations/master.conf" MASTER_CONF="$ECOSYSTEM_ROOT/Configurations/master.conf"
sed -i "s/WEBHOOK_SECRET=\"\"/WEBHOOK_SECRET=\"$GENERATED\"/" "$MASTER_CONF" sed -i "s/WEBHOOK_SECRET=\"\"/WEBHOOK_SECRET=\"$GENERATED\"/" "$MASTER_CONF"
WEBHOOK_SECRET="$GENERATED" WEBHOOK_SECRET="$GENERATED"
# If the sed didn't match, the secret only exists in this process. The listener would
# come up, then regenerate a different secret on the next start — silently invalidating
# the key already registered in the arrs. Fail loudly instead.
if ! grep -q "WEBHOOK_SECRET=\"$GENERATED\"" "$MASTER_CONF" 2>/dev/null; then
error "Generated WEBHOOK_SECRET but could not persist it to $MASTER_CONF"
error "Set WEBHOOK_SECRET manually — a non-persisted secret changes on every restart"
notify "Webhook secret not persisted on $(hostname) — set WEBHOOK_SECRET manually" \
"Webhook Listener" "warning"
exit 1
fi
echo "[webhook] Generated WEBHOOK_SECRET — run Tools/webhook_setup.sh to register in arrs" echo "[webhook] Generated WEBHOOK_SECRET — run Tools/webhook_setup.sh to register in arrs"
fi fi
+80 -7
View File
@@ -15,6 +15,30 @@
# begin searching — a search it will never win because we already have it. # begin searching — a search it will never win because we already have it.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Invoked per import by webhook_listener.js with <arr_type> <item_path>:
#
# 1. Validate
# → arr_type must be sonarr|radarr|lidarr; item_path must exist and be a safe
# absolute path
#
# 2. Resolve arr specifics
# → API port, API version and rescan command for that arr type
#
# 3. Discover remote nodes
# → discover_remote_nodes(); no remotes configured means exit cleanly
#
# 4. Per remote node, independently:
# a. Resolve its Tailscale IP — unresolvable skips that node
# b. rsync the single item to the same absolute path (--no-delete)
# c. Skip the rescan if rsync failed — never scan a partial file
# d. Trigger the arr's refresh command, cache-first API key with SSH fallback
#
# One failing node is counted and skipped; the rest still receive the upgrade.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -38,13 +62,47 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Arg validation — exits with usage message if arr_type or item_path missing # Root Enforcement
# Path existence — exits if item_path is not a directory on disk # rsync runs over SSH as root and writes to root@remote at the same absolute path.
# Tailscale resolution — skips a node if its Tailscale IP cannot be resolved #
# rsync exit check — rescan is only triggered if rsync succeeded; a failed # Argument Validation
# transfer does not cause the remote arr to scan a partial file # Exits with a usage message if arr_type or item_path is missing, and rejects an
# SSH fallback — if no cached API key, falls back to SSH to read config.xml # arr_type outside sonarr|radarr|lidarr rather than defaulting to one.
# on the remote rather than failing the rescan step #
# Path Existence
# Exits if item_path is not a directory on disk.
#
# Item Path Depth Guard
# item_path must be an absolute path at least three levels deep. It arrives from the
# arr's webhook payload and is rsynced to the same path on the partner, so a truncated
# or malformed value would push a system directory — or the filesystem root — onto the
# remote. The existence check alone does not catch this, because / is a directory.
#
# No Lock — Deliberate
# This is an event handler invoked per import by webhook_listener.js. Concurrent
# upgrades are normal and expected. A default lock would silently drop overlapping
# events, and a waiting lock would queue them behind a slow transfer, so neither is
# used: each invocation rsyncs a different item path and they do not contend.
#
# Tailscale Resolution
# Skips a node if its Tailscale IP cannot be resolved, rather than attempting the
# transfer against an unresolved or stale address.
#
# rsync Exit Check
# The rescan is only triggered if rsync succeeded. A failed transfer never causes the
# remote arr to scan a partial file into its library.
#
# No Delete on Push
# rsync runs with --no-delete. This pushes one upgraded item; it is not a mirror, and
# must never remove content on the partner that this run does not know about.
#
# SSH Fallback
# If no cached API key is available, falls back to SSH to read config.xml on the
# remote rather than failing the rescan step.
#
# Per-Node Isolation
# One unreachable or failing node is counted and skipped; the remaining nodes still
# receive the upgrade.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -90,6 +148,21 @@ ITEM_PATH="${2:-}"
[[ -d "$ITEM_PATH" ]] || { echo "Path not found: $ITEM_PATH" >&2; exit 1; } [[ -d "$ITEM_PATH" ]] || { echo "Path not found: $ITEM_PATH" >&2; exit 1; }
if [[ "$EUID" -ne 0 ]]; then
echo "Must be run as root" >&2
exit 1
fi
# ITEM_PATH is rsynced to root@remote at the same absolute path. It arrives from the arr's
# webhook payload, so a malformed or truncated value would push a system directory — or the
# filesystem root — onto the partner. -d alone does not catch that: / is a directory.
_depth="${ITEM_PATH//[^\/]/}"
if [[ "$ITEM_PATH" != /* || "${#_depth}" -lt 3 ]]; then
echo "Refusing unsafe item path: '$ITEM_PATH' — expected an absolute path at least 3 levels deep" >&2
exit 1
fi
unset _depth
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../load_config.sh" source "$SCRIPT_DIR/../load_config.sh"
detect_hosts detect_hosts
+24 -5
View File
@@ -68,6 +68,15 @@
# Docker Presence Check # Docker Presence Check
# Verifies docker binary exists before execution. # Verifies docker binary exists before execution.
# #
# Docker Daemon Check
# Verifies the daemon is responsive before enumerating containers. A hung
# daemon returns an empty container list, which would otherwise be read as
# "nothing to stop" and pass a shutdown that never happened.
#
# Host Detection
# detect_hosts() identifies which server is running the script and sets
# MY_ID / LOCAL_SERVER_NAME for logging and notifications.
#
# Lock Acquisition # Lock Acquisition
# acquire_lock prevents overlapping runs (e.g. array_stopping firing twice). # acquire_lock prevents overlapping runs (e.g. array_stopping firing twice).
# #
@@ -140,10 +149,20 @@ fi
detect_hosts detect_hosts
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be stopped"
DOCKER_TIMEOUT=30 DOCKER_TIMEOUT=30
DOCKER_STOP_TIMEOUT=30 # grace period for SIGTERM before docker sends SIGKILL internally DOCKER_STOP_TIMEOUT=30 # grace period for SIGTERM before docker sends SIGKILL internally
# A hung daemon makes docker ps return nothing — indistinguishable from "no containers
# running", which would silently report a clean shutdown that never happened.
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
error "Docker daemon not responding — cannot verify container shutdown"
notify "Container stop aborted on $(hostname) ($MY_ID) — Docker daemon not responding" \
"Docker Container Stop" "warning"
exit 1
fi
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be stopped"
_RETRY_COUNT="${RETRY_COUNT:-3}" _RETRY_COUNT="${RETRY_COUNT:-3}"
log "$ICON_GEAR Config: retries=$_RETRY_COUNT sleep=${SLEEP:-5}s grace=${DOCKER_STOP_TIMEOUT}s cmd-timeout=${DOCKER_TIMEOUT}s" log "$ICON_GEAR Config: retries=$_RETRY_COUNT sleep=${SLEEP:-5}s grace=${DOCKER_STOP_TIMEOUT}s cmd-timeout=${DOCKER_TIMEOUT}s"
@@ -152,7 +171,7 @@ log "$ICON_GEAR Config: retries=$_RETRY_COUNT sleep=${SLEEP:-5}s grace=${DOCKER_
# ━━━ Status ━━━ # ━━━ Status ━━━
# ============================================================================================== # ==============================================================================================
if [[ "$SHOW_STATUS" == true ]]; then if [[ "$SHOW_STATUS" == true ]]; then
mapfile -t RUNNING < <(docker ps --format '{{.Names}}' 2>/dev/null | sort) mapfile -t RUNNING < <(timeout "$DOCKER_TIMEOUT" docker ps --format '{{.Names}}' 2>/dev/null | sort)
echo "" echo ""
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━" echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)" echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
@@ -166,7 +185,7 @@ fi
# ============================================================================================== # ==============================================================================================
# ━━━ Stop Containers ━━━ # ━━━ Stop Containers ━━━
# ============================================================================================== # ==============================================================================================
mapfile -t RUNNING < <(docker ps --format '{{.Names}}' 2>/dev/null | sort) mapfile -t RUNNING < <(timeout "$DOCKER_TIMEOUT" docker ps --format '{{.Names}}' 2>/dev/null | sort)
echo "" echo ""
echo "━━━ $ICON_CONTAINERS Docker Container Stop — $(date '+%Y-%m-%d %H:%M:%S') ━━━" echo "━━━ $ICON_CONTAINERS Docker Container Stop — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
@@ -187,7 +206,7 @@ FAILED=()
for container in "${RUNNING[@]}"; do for container in "${RUNNING[@]}"; do
[[ -z "$container" ]] && continue [[ -z "$container" ]] && continue
c_start=$(date +%s) c_start=$(date +%s)
c_image=$(docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown") c_image=$(timeout "$DOCKER_TIMEOUT" docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown")
log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━" log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━"
if [[ "$DRY_RUN" == true ]]; then if [[ "$DRY_RUN" == true ]]; then
+63 -3
View File
@@ -14,6 +14,33 @@
# so there is no second list to maintain. # so there is no second list to maintain.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Containers are processed one at a time in dependency-safe order:
#
# 1. Build restart order
# → build_restart_order() sorts DAILY_RESTART_CONTAINERS by WATCHDOG_DEPENDENCIES
#
# 2. Skip anything docker_update.sh already rebuilt this run
# → a rebuild onto a new image already restarted it moments ago
#
# 3. Inspect container state
# missing → skip, not an error
# stopped → skip, stopped state is respected
# running → restart
#
# 4. Restart with retry
# → retry_docker wraps each attempt in a timeout, up to RETRY_COUNT
#
# 5. Verify it stayed running
# → verify_running() settles for RESTART_VERIFY_WAIT then checks State.Running
# → a container that crashes immediately is marked failed and notified
#
# 6. Prune dangling images
# → restarts swap onto new images, leaving the old ones dangling
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -37,6 +64,27 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root Enforcement
# Docker operations require root privileges.
#
# Docker Presence Check
# Verifies the docker binary exists before execution. Notifies on absence —
# a missing binary during the maintenance window is worth knowing about.
#
# Docker Daemon Check
# Verifies the daemon is responsive before any restart work. Every container
# would otherwise fail its inspect and be logged as an unknown-status failure,
# burying one daemon fault under a list of bogus per-container errors.
#
# Host Detection
# detect_hosts() identifies which server is running the script and aliases
# HOST*_DAILY_RESTART_CONTAINERS and HOST*_WATCHDOG_DEPENDENCIES to the
# correct host's values.
#
# Empty List Guard
# Exits cleanly with a pointer to the relevant conf key if
# DAILY_RESTART_CONTAINERS is unconfigured for this host.
#
# Dependency Ordering # Dependency Ordering
# Containers restart in dependency-safe order using HOST*_WATCHDOG_DEPENDENCIES. # Containers restart in dependency-safe order using HOST*_WATCHDOG_DEPENDENCIES.
# CONTAINER_DELAY seconds between dependency restart and dependent restart gives # CONTAINER_DELAY seconds between dependency restart and dependent restart gives
@@ -53,6 +101,11 @@
# cannot cause this script to hang indefinitely. Timed-out commands retry # cannot cause this script to hang indefinitely. Timed-out commands retry
# per RETRY_COUNT before marking as failed. # per RETRY_COUNT before marking as failed.
# #
# Stale Rebuild-List Guard
# The rebuilt-container list written by docker_update.sh is discarded if older
# than DOCKER_UPDATE_REBUILT_STALE_HOURS. A stale file would otherwise suppress
# real restarts based on an update run that never happened today.
#
# Lock Acquisition # Lock Acquisition
# acquire_lock() prevents concurrent execution if a previous run is still active. # acquire_lock() prevents concurrent execution if a previous run is still active.
# #
@@ -134,6 +187,14 @@ fi
# detect_hosts() sets MY_ID and aliases HOST*_DAILY_RESTART_CONTAINERS → DAILY_RESTART_CONTAINERS # detect_hosts() sets MY_ID and aliases HOST*_DAILY_RESTART_CONTAINERS → DAILY_RESTART_CONTAINERS
detect_hosts detect_hosts
# Without this, a hung daemon fails every container's inspect individually and the summary
# reports a list of unknown-status failures instead of the one fault that caused them.
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
error "Docker daemon not responding — skipping daily restart"
notify "Daily restart skipped on $(hostname) — Docker daemon not responding" "Docker Daily Restart" "warning"
exit 1
fi
if [[ ${#DAILY_RESTART_CONTAINERS[@]} -eq 0 ]]; then if [[ ${#DAILY_RESTART_CONTAINERS[@]} -eq 0 ]]; then
warn "DAILY_RESTART_CONTAINERS is empty for $MY_ID — nothing to restart" warn "DAILY_RESTART_CONTAINERS is empty for $MY_ID — nothing to restart"
warn "Check HOST${MY_ID#HOST}_DAILY_RESTART_CONTAINERS in host*.conf" warn "Check HOST${MY_ID#HOST}_DAILY_RESTART_CONTAINERS in host*.conf"
@@ -213,7 +274,7 @@ LAST_RESTARTED=""
for container in "${ORDERED_RESTART[@]}"; do for container in "${ORDERED_RESTART[@]}"; do
[[ -z "$container" ]] && continue [[ -z "$container" ]] && continue
c_start=$(date +%s) c_start=$(date +%s)
c_image=$(docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown") c_image=$(timeout "$DOCKER_TIMEOUT" docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown")
log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━" log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━"
if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then
@@ -242,7 +303,6 @@ for container in "${ORDERED_RESTART[@]}"; do
RESTARTED+=("$container") RESTARTED+=("$container")
else else
if retry_docker docker restart "$container"; then if retry_docker docker restart "$container"; then
[[ "${RESTART_VERIFY_WAIT:-3}" -gt 0 ]] && sleep "${RESTART_VERIFY_WAIT:-3}"
if verify_running "$container"; then if verify_running "$container"; then
echo "$ICON_STARTED $container restarted and running in $(format_duration $(( $(date +%s) - c_start )))" echo "$ICON_STARTED $container restarted and running in $(format_duration $(( $(date +%s) - c_start )))"
RESTARTED+=("$container") RESTARTED+=("$container")
@@ -283,7 +343,7 @@ if [[ "$DRY_RUN" == true ]]; then
warn "DRY RUN — would prune dangling images" warn "DRY RUN — would prune dangling images"
PRUNED_SUMMARY="(dry run)" PRUNED_SUMMARY="(dry run)"
else else
PRUNED_OUTPUT=$(docker image prune -f 2>&1) PRUNED_OUTPUT=$(timeout "$DOCKER_TIMEOUT" docker image prune -f 2>&1)
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /' [[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed") PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
fi fi
+14 -6
View File
@@ -51,6 +51,12 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root Enforcement
# Docker network operations require root privileges.
#
# Docker Presence Check
# Verifies the docker binary exists before any network operations.
#
# Lock Acquisition # Lock Acquisition
# Prevents concurrent execution via acquire_lock(). Safe to call from # Prevents concurrent execution via acquire_lock(). Safe to call from
# array start hooks or manually without risk of overlap. # array start hooks or manually without risk of overlap.
@@ -71,8 +77,10 @@
# Warns and exits cleanly if NETWORK_CONNECT_NETWORKS or # Warns and exits cleanly if NETWORK_CONNECT_NETWORKS or
# NETWORK_CONNECT_CONTAINERS are unconfigured. # NETWORK_CONNECT_CONTAINERS are unconfigured.
# #
# Command Validation # Missing Container Tolerance
# Validates unRAID notify script before use. # A configured container that does not exist yet warns and is skipped rather
# than failing the run. This script executes early at array start, before
# every container has necessarily been created.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -130,13 +138,12 @@ fi
# detect_hosts() sets MY_ID and aliases HOST*_NETWORK_CONNECT_* arrays # detect_hosts() sets MY_ID and aliases HOST*_NETWORK_CONNECT_* arrays
detect_hosts detect_hosts
# Validate unRAID notify script — used for network creation alerts
# Docker daemon check — network operations are useless if daemon is hung # Docker daemon check — network operations are useless if daemon is hung
DOCKER_TIMEOUT=15 DOCKER_TIMEOUT=15
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
error "Docker daemon not responding — cannot manage networks" error "Docker daemon not responding — cannot manage networks"
notify "docker_network_connect failed on $(hostname) — Docker daemon not responding" "Network Connect" "warning" notify "docker_network_connect failed on $(hostname) — Docker daemon not responding" \
"Network Connect" "warning"
exit 1 exit 1
fi fi
@@ -291,7 +298,8 @@ if [[ "$DRY_RUN" == true ]]; then
warn "DRY RUN — no changes made" warn "DRY RUN — no changes made"
elif [[ ${#FAILED[@]} -gt 0 ]]; then elif [[ ${#FAILED[@]} -gt 0 ]]; then
echo "$ICON_ERROR Status: SOME OPERATIONS FAILED" echo "$ICON_ERROR Status: SOME OPERATIONS FAILED"
notify "Docker network connect failed on $(hostname)${FAILED[*]}" "Network Connect" "warning" notify "Docker network connect failed on $(hostname)${FAILED[*]}" \
"Network Connect" "warning"
elif [[ ${#NETWORKS_CREATED[@]} -gt 0 ]]; then elif [[ ${#NETWORKS_CREATED[@]} -gt 0 ]]; then
warn "Networks recreated — ${NETWORKS_CREATED[*]} — unRAID update likely wiped them" warn "Networks recreated — ${NETWORKS_CREATED[*]} — unRAID update likely wiped them"
else else
+38 -8
View File
@@ -74,6 +74,28 @@
# Root Enforcement # Root Enforcement
# Docker operations require root privileges. # Docker operations require root privileges.
# #
# Docker Presence Check
# Verifies the docker binary exists before execution.
#
# Docker Daemon Check
# Verifies the daemon is responsive before container discovery. Remainder mode
# derives its entire target list from docker ps — against a hung daemon that
# returns empty and the run silently reports "no containers to update".
#
# Timeout Protection
# Inspect, discovery and image-query commands are wrapped in a timeout so a
# hung daemon cannot stall the maintenance window. docker pull is deliberately
# NOT wrapped — a large image legitimately takes longer than any sane timeout,
# and killing it mid-layer wastes the transfer.
#
# Empty List Guards
# Each mode exits cleanly with a pointer to the relevant conf key when its
# container list is unconfigured for this host.
#
# Rebuild Failure Fallback
# A container that fails to rebuild is excluded from the rebuilt-list handoff
# file, so the follow-up restart script still gives it a normal restart pass.
#
# DAILY_CONTAINER_UPDATES / WEEKLY_CONTAINER_UPDATES Toggles # DAILY_CONTAINER_UPDATES / WEEKLY_CONTAINER_UPDATES Toggles
# Each mode exits cleanly when disabled. Restart scripts run regardless — # Each mode exits cleanly when disabled. Restart scripts run regardless —
# update and restart are independent operations. # update and restart are independent operations.
@@ -192,6 +214,14 @@ fi
detect_hosts detect_hosts
# Remainder mode builds its whole target list from docker ps — a hung daemon returns
# empty and the run would report "no containers to update" instead of failing.
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
error "Docker daemon not responding — skipping image updates"
notify "Docker update skipped on $(hostname) — Docker daemon not responding" "Docker Update" "warning"
exit 1
fi
# ============================================================================================== # ==============================================================================================
# ━━━ Container Discovery ━━━ # ━━━ Container Discovery ━━━
# ============================================================================================== # ==============================================================================================
@@ -234,7 +264,7 @@ if [[ "$REMAINDER_MODE" == true ]]; then
done done
unset _tier _tier_var _tier_arr _c unset _tier _tier_var _tier_arr _c
mapfile -t _all_running < <(docker ps --format '{{.Names}}' | sort) mapfile -t _all_running < <(timeout "$DOCKER_TIMEOUT" docker ps --format '{{.Names}}' | sort)
TARGET_CONTAINERS=() TARGET_CONTAINERS=()
for _c in "${_all_running[@]}"; do for _c in "${_all_running[@]}"; do
[[ -z "${_exclude[$_c]+x}" ]] && TARGET_CONTAINERS+=("$_c") [[ -z "${_exclude[$_c]+x}" ]] && TARGET_CONTAINERS+=("$_c")
@@ -328,13 +358,13 @@ for container in "${TARGET_CONTAINERS[@]}"; do
[[ -z "$container" ]] && continue [[ -z "$container" ]] && continue
log "━━━ $ICON_CONTAINERS $container ━━━" log "━━━ $ICON_CONTAINERS $container ━━━"
if ! docker inspect "$container" &>/dev/null; then if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then
warn "$container — not found, skipping" warn "$container — not found, skipping"
SKIPPED+=("$container") SKIPPED+=("$container")
continue continue
fi fi
IMAGE=$(docker inspect --format='{{.Config.Image}}' "$container" 2>/dev/null) IMAGE=$(timeout "$DOCKER_TIMEOUT" docker inspect --format='{{.Config.Image}}' "$container" 2>/dev/null)
if [[ -z "$IMAGE" ]]; then if [[ -z "$IMAGE" ]]; then
warn "$container — could not determine image, skipping" warn "$container — could not determine image, skipping"
SKIPPED+=("$container") SKIPPED+=("$container")
@@ -352,8 +382,8 @@ for container in "${TARGET_CONTAINERS[@]}"; do
# Capture the image ID the container is currently running on, and the # Capture the image ID the container is currently running on, and the
# image ID :latest points to before the pull. After pulling, we rebuild if # image ID :latest points to before the pull. After pulling, we rebuild if
# either a new digest landed OR the container is behind what :latest is now. # either a new digest landed OR the container is behind what :latest is now.
CONTAINER_IMAGE_ID=$(docker inspect "$container" --format='{{.Image}}' 2>/dev/null || echo "") CONTAINER_IMAGE_ID=$(timeout "$DOCKER_TIMEOUT" docker inspect "$container" --format='{{.Image}}' 2>/dev/null || echo "")
OLD_ID=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "") OLD_ID=$(timeout "$DOCKER_TIMEOUT" docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
log "$ICON_SYNC Pulling $IMAGE..." log "$ICON_SYNC Pulling $IMAGE..."
if [[ "$ENABLE_LOGGING" == "true" ]]; then if [[ "$ENABLE_LOGGING" == "true" ]]; then
@@ -363,7 +393,7 @@ for container in "${TARGET_CONTAINERS[@]}"; do
docker pull "$IMAGE" >/dev/null 2>&1 docker pull "$IMAGE" >/dev/null 2>&1
_pull_rc=$? _pull_rc=$?
fi fi
NEW_ID=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "") NEW_ID=$(timeout "$DOCKER_TIMEOUT" docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
if [[ $_pull_rc -eq 0 ]]; then if [[ $_pull_rc -eq 0 ]]; then
_pull_new=$( [[ -n "$OLD_ID" && "$OLD_ID" != "$NEW_ID" ]] && echo true || echo false) _pull_new=$( [[ -n "$OLD_ID" && "$OLD_ID" != "$NEW_ID" ]] && echo true || echo false)
@@ -436,9 +466,9 @@ if [[ "$DRY_RUN" == true ]]; then
PRUNED_SUMMARY="(dry run)" PRUNED_SUMMARY="(dry run)"
else else
for _old_id in "${OLD_IMAGE_IDS[@]}"; do for _old_id in "${OLD_IMAGE_IDS[@]}"; do
docker rmi "$_old_id" >/dev/null 2>&1 || true timeout "$DOCKER_TIMEOUT" docker rmi "$_old_id" >/dev/null 2>&1 || true
done done
PRUNED_OUTPUT=$(docker image prune -f 2>&1) PRUNED_OUTPUT=$(timeout "$DOCKER_TIMEOUT" docker image prune -f 2>&1)
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /' [[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed") PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
fi fi
+66 -6
View File
@@ -16,6 +16,33 @@
# stopped → leave, missing → skip. Container state is always respected. # stopped → leave, missing → skip. Container state is always respected.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Identical to docker_daily_restart.sh, against WEEKLY_RESTART_CONTAINERS:
#
# 1. Build restart order
# → build_restart_order() sorts WEEKLY_RESTART_CONTAINERS by WATCHDOG_DEPENDENCIES
#
# 2. Skip anything docker_update.sh --weekly already rebuilt this run
# → a rebuild onto a new image already restarted it moments ago
#
# 3. Inspect container state
# missing → skip, not an error
# stopped → skip, stopped state is respected
# running → restart
#
# 4. Restart with retry
# → retry_docker wraps each attempt in a timeout, up to RETRY_COUNT
#
# 5. Verify it stayed running
# → verify_running() settles for RESTART_VERIFY_WAIT then checks State.Running
# → a container that crashes immediately is marked failed and notified
#
# 6. Prune dangling images
# → restarts swap onto new images, leaving the old ones dangling
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -39,6 +66,26 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root Enforcement
# Docker operations require root privileges.
#
# Docker Presence Check
# Verifies the docker binary exists before execution. Notifies on absence.
#
# Docker Daemon Check
# Verifies the daemon is responsive before any restart work. Every container
# would otherwise fail its inspect and be logged as an unknown-status failure,
# burying one daemon fault under a list of bogus per-container errors.
#
# Host Detection
# detect_hosts() identifies which server is running the script and aliases
# HOST*_WEEKLY_RESTART_CONTAINERS and HOST*_WATCHDOG_DEPENDENCIES to the
# correct host's values.
#
# Empty List Guard
# Exits cleanly with a pointer to the relevant conf key if
# WEEKLY_RESTART_CONTAINERS is unconfigured for this host.
#
# Dependency Ordering # Dependency Ordering
# Containers restart in dependency-safe order using HOST*_WATCHDOG_DEPENDENCIES. # Containers restart in dependency-safe order using HOST*_WATCHDOG_DEPENDENCIES.
# CONTAINER_DELAY seconds between dependency restart and dependent restart. # CONTAINER_DELAY seconds between dependency restart and dependent restart.
@@ -51,10 +98,10 @@
# All docker commands wrapped in a 30 second timeout. A hung Docker daemon # All docker commands wrapped in a 30 second timeout. A hung Docker daemon
# cannot cause this script to hang indefinitely. # cannot cause this script to hang indefinitely.
# #
# Host Detection # Stale Rebuild-List Guard
# detect_hosts() identifies which server is running the script and aliases # The rebuilt-container list written by docker_update.sh --weekly is discarded
# HOST*_WEEKLY_RESTART_CONTAINERS and HOST*_WATCHDOG_DEPENDENCIES to the # if older than DOCKER_UPDATE_REBUILT_STALE_HOURS. A stale file would otherwise
# correct host's values. # suppress real restarts based on an update run that never happened this week.
# #
# Lock Acquisition # Lock Acquisition
# acquire_lock() prevents concurrent execution. # acquire_lock() prevents concurrent execution.
@@ -84,6 +131,11 @@
# CONTAINER_DELAY # CONTAINER_DELAY
# Seconds to wait after restarting a dependency before starting its dependents # Seconds to wait after restarting a dependency before starting its dependents
# #
# RESTART_VERIFY_WAIT
# Seconds verify_running() waits after docker restart before checking the
# container is running. Gives the process time to initialise before the
# state is sampled. (default: 3)
#
# DOCKER_UPDATE_REBUILT_WEEKLY_FILE / DOCKER_UPDATE_REBUILT_STALE_HOURS # DOCKER_UPDATE_REBUILT_WEEKLY_FILE / DOCKER_UPDATE_REBUILT_STALE_HOURS
# List of containers docker_update.sh --weekly already rebuilt onto a new image # List of containers docker_update.sh --weekly already rebuilt onto a new image
# this run — read here so they're not restarted a second time. Discarded as # this run — read here so they're not restarted a second time. Discarded as
@@ -132,6 +184,14 @@ fi
# detect_hosts() sets MY_ID and aliases HOST*_WEEKLY_RESTART_CONTAINERS → WEEKLY_RESTART_CONTAINERS # detect_hosts() sets MY_ID and aliases HOST*_WEEKLY_RESTART_CONTAINERS → WEEKLY_RESTART_CONTAINERS
detect_hosts detect_hosts
# Without this, a hung daemon fails every container's inspect individually and the summary
# reports a list of unknown-status failures instead of the one fault that caused them.
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
error "Docker daemon not responding — skipping weekly restart"
notify "Weekly restart skipped on $(hostname) — Docker daemon not responding" "Docker Weekly Restart" "warning"
exit 1
fi
if [[ ${#WEEKLY_RESTART_CONTAINERS[@]} -eq 0 ]]; then if [[ ${#WEEKLY_RESTART_CONTAINERS[@]} -eq 0 ]]; then
warn "WEEKLY_RESTART_CONTAINERS is empty for $MY_ID — nothing to restart" warn "WEEKLY_RESTART_CONTAINERS is empty for $MY_ID — nothing to restart"
warn "Check HOST*_WEEKLY_RESTART_CONTAINERS in host*.conf" warn "Check HOST*_WEEKLY_RESTART_CONTAINERS in host*.conf"
@@ -211,7 +271,7 @@ LAST_RESTARTED=""
for container in "${ORDERED_RESTART[@]}"; do for container in "${ORDERED_RESTART[@]}"; do
[[ -z "$container" ]] && continue [[ -z "$container" ]] && continue
c_start=$(date +%s) c_start=$(date +%s)
c_image=$(docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown") c_image=$(timeout "$DOCKER_TIMEOUT" docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown")
log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━" log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━"
if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then
@@ -281,7 +341,7 @@ if [[ "$DRY_RUN" == true ]]; then
warn "DRY RUN — would prune dangling images" warn "DRY RUN — would prune dangling images"
PRUNED_SUMMARY="(dry run)" PRUNED_SUMMARY="(dry run)"
else else
PRUNED_OUTPUT=$(docker image prune -f 2>&1) PRUNED_OUTPUT=$(timeout "$DOCKER_TIMEOUT" docker image prune -f 2>&1)
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /' [[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed") PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
fi fi
+33
View File
@@ -61,6 +61,15 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root Enforcement
# Failed-import purging deletes directories owned by container users.
#
# Dependency Check
# Verifies curl and jq exist before any API work. jq backs the slskd connection
# probe — without it the probe returns false forever, the script burns its full
# 60 second reconnect wait, then skips every slskd section as "disconnected".
# A missing dependency is reported as itself rather than as a phantom outage.
#
# Active Transfer Protection # Active Transfer Protection
# slskd: skips users with InProgress or Queued transfers before any removal. # slskd: skips users with InProgress or Queued transfers before any removal.
# SABnzbd: age threshold enforced before deletion. # SABnzbd: age threshold enforced before deletion.
@@ -70,6 +79,19 @@
# Each section validates its downloader URL before API calls. Missing or # Each section validates its downloader URL before API calls. Missing or
# unreachable downloaders skip without affecting other sections. # unreachable downloaders skip without affecting other sections.
# #
# No Downloaders Guard
# Exits cleanly when none of SLSKD_URL, SABNZBD_URL or QBIT_URL are set for
# this host — nothing configured is not an error.
#
# Timeout Protection
# Every curl carries --max-time. An unresponsive downloader cannot stall the
# 30 minute maintenance cycle or overlap the next run.
#
# Deletion Scope Limit
# qBittorrent removals pass deleteFiles=false — the torrent record is dropped
# but files on disk are left for the arrs to manage. This script never deletes
# media.
#
# Host Detection # Host Detection
# detect_hosts() identifies which server is running the script and aliases # detect_hosts() identifies which server is running the script and aliases
# all HOST*_SLSKD_*, HOST*_SABNZBD_*, and HOST*_QBIT_* vars to the correct # all HOST*_SLSKD_*, HOST*_SABNZBD_*, and HOST*_QBIT_* vars to the correct
@@ -140,6 +162,17 @@ fi
# Lock first — wait mode since this runs every 30min and previous may still be finishing # Lock first — wait mode since this runs every 30min and previous may still be finishing
acquire_lock "wait" acquire_lock "wait"
# jq backs the slskd connection probe. Missing, the probe never returns true and slskd
# looks permanently disconnected — a 60s wait followed by silently skipped sections.
for _dep in curl jq; do
if ! command -v "$_dep" &>/dev/null; then
error "$_dep not found — required for downloader API calls"
notify "Downloaders reset failed on $(hostname)$_dep not installed" "Downloaders Reset" "warning"
exit 1
fi
done
unset _dep
# detect_hosts() sets MY_ID and aliases all HOST*_SLSKD_*, HOST*_SABNZBD_*, HOST*_QBIT_* vars # detect_hosts() sets MY_ID and aliases all HOST*_SLSKD_*, HOST*_SABNZBD_*, HOST*_QBIT_* vars
detect_hosts detect_hosts
+49 -5
View File
@@ -89,7 +89,50 @@
# #
# FALLBACK_ENABLED Gate # FALLBACK_ENABLED Gate
# Exits cleanly when disabled — safe to run on servers being rebuilt without # Exits cleanly when disabled — safe to run on servers being rebuilt without
# triggering spurious fallback actions. # triggering spurious fallback actions. Fail-closed: anything that is not exactly
# "true" counts as disabled, so a malformed toggle cannot grant this script DDNS
# authority and cross-server container control by accident.
#
# Docker Presence Check
# Verifies the docker binary exists before the state machine starts.
#
# Remote IP Resolution
# resolve_remote_ip() must resolve the partner before any SSH operation, so a
# remote command can never be issued against an unresolved or stale address.
#
# Asymmetric Failover / Handback
# Entering FALLBACK is immediate — a down partner means users are already affected.
# Returning requires FALLBACK_HANDBACK_STRIKES consecutive remote-up checks. The
# asymmetry is deliberate: protecting fast costs a few minutes of redundant
# coverage, handing back fast on a flapping partner costs a second outage.
#
# DDNS Excluded From Tier Loops
# The Tier 1 stop and start loops explicitly skip any container that is also in
# REMOTE_DDNS_CONTAINERS. DDNS is sequenced by the handoff and cutover steps alone,
# so ordinary tier processing can never move DNS at the wrong moment.
#
# Writeback Delay Gate
# Tier writeback rsync only runs once the outage has exceeded that tier's writeback
# delay. A brief blip does not trigger a full data writeback, which would cost more
# than the outage it is compensating for.
#
# FALLBACK_RSYNC_ENABLED Gate
# Writeback is skipped entirely when disabled, and the skip is announced rather than
# silent — containers still hand back, but nobody is left assuming data moved.
#
# Play State Sync Before Cutover
# Handback retries play_state_sync up to PLAY_SYNC_HANDBACK_RETRIES times before DNS
# cuts over, so users land on current watch state. Exhausting retries warns and
# proceeds — stale resume positions are not worth holding DNS on a downed service.
#
# Partnership Suspend Abort
# If partnership goes inactive mid-fallback, _abort_fallback_containers() stops the
# fallback containers and returns to NORMAL rather than leaving this host serving a
# partner it is no longer paired with.
#
# Container Verify Wait
# CONTAINER_VERIFY_WAIT seconds elapse after each start before the running check, so
# a container that starts and immediately crashes is caught rather than counted up.
# #
# Version Parity Check # Version Parity Check
# Refuses handback if remote unRAID version doesn't match. A mismatch may # Refuses handback if remote unRAID version doesn't match. A mismatch may
@@ -271,8 +314,11 @@ if [[ "$EUID" -ne 0 ]]; then
exit 1 exit 1
fi fi
# FALLBACK_ENABLED gate — exits cleanly when disabled (e.g. HOST2 being rebuilt) # FALLBACK_ENABLED gate — exits cleanly when disabled (e.g. HOST2 being rebuilt).
if [[ "${FALLBACK_ENABLED:-false}" == false ]]; then # Fail-closed: anything that isn't exactly "true" disables fallback. Matching only the
# literal "false" would let a typo ("no", "0", "FALSE") hand this script DDNS authority
# and cross-server container control on a toggle nobody meant to set.
if [[ "${FALLBACK_ENABLED:-false}" != "true" ]]; then
warn "FALLBACK_ENABLED=false — fallback monitoring disabled" warn "FALLBACK_ENABLED=false — fallback monitoring disabled"
warn "Set FALLBACK_ENABLED=true in master.conf when both servers are ready" warn "Set FALLBACK_ENABLED=true in master.conf when both servers are ready"
exit 0 exit 0
@@ -289,8 +335,6 @@ detect_hosts
require_partnership require_partnership
resolve_remote_ip resolve_remote_ip
# Validate unRAID notify script — used throughout for state change notifications
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no container or DDNS changes will be made" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no container or DDNS changes will be made"
# Timeout for all docker and SSH docker commands # Timeout for all docker and SSH docker commands
+46 -4
View File
@@ -41,9 +41,23 @@
# ============================================================================================== # ==============================================================================================
# #
# iptables Safety Trap # iptables Safety Trap
# The DROP rule is removed via trap on ANY exit — normal completion, crash, error, # The DROP rule is removed via an EXIT trap that fires on normal completion, error
# ctrl-c. Remote connectivity is always restored regardless of test outcome. # exit, script crash, ctrl-c (SIGINT) and SIGTERM — verified, not assumed. Remote
# You cannot accidentally leave the remote permanently blocked. # connectivity is restored regardless of test outcome.
#
# Stale Rule Sweep
# The trap above cannot cover SIGKILL or a power cut, which are the only ways a DROP
# rule survives the test. One stranded that way makes fallback.sh see the partner as
# permanently down and hold FALLBACK indefinitely, so pre-flight clears any leftover
# rule before doing anything else — including before the reachability check, which
# would otherwise fail and blame the network for the test's own residue.
#
# Root Enforcement
# iptables and container control require root.
#
# iptables Presence Check
# platform_require_cmd confirms iptables exists before the test begins — there is no
# point entering a connectivity simulation that cannot simulate anything.
# #
# FALLBACK_ENABLED Gate # FALLBACK_ENABLED Gate
# Aborts if FALLBACK_ENABLED=false. Testing a disabled fallback system is # Aborts if FALLBACK_ENABLED=false. Testing a disabled fallback system is
@@ -139,6 +153,28 @@ cleanup() {
trap cleanup EXIT trap cleanup EXIT
# ── Stale rule sweep — the one case the trap above cannot cover ───────────────────────────────
# The EXIT trap fires on normal exit, error, ctrl-c and SIGTERM, but not on SIGKILL or a power
# cut. A DROP rule stranded that way makes fallback.sh see the partner as permanently down and
# sit in FALLBACK indefinitely — so clear any leftover from a previous run before starting.
_clear_stale_block() {
[[ -z "${REMOTE_SERVER:-}" ]] && return
local removed=0
while iptables -C OUTPUT -d "$REMOTE_SERVER" -j DROP 2>/dev/null; do
if [[ "$DRY_RUN" == true ]]; then
warn "DRY RUN — would remove stale iptables block on $REMOTE_SERVER"
return
fi
iptables -D OUTPUT -d "$REMOTE_SERVER" -j DROP 2>/dev/null || break
(( removed++ ))
done
if [[ "$removed" -gt 0 ]]; then
warn "$ICON_SHIELD Removed $removed stale iptables block(s) on $REMOTE_SERVER from a previous run"
notify "Fallback test cleared $removed stale iptables block(s) on $(hostname) — a previous test was killed before cleanup" \
"Fallback Test" "warning"
fi
}
# ============================================================================================== # ==============================================================================================
# ━━━ Setup ━━━ # ━━━ Setup ━━━
# ============================================================================================== # ==============================================================================================
@@ -149,7 +185,8 @@ if [[ "$EUID" -ne 0 ]]; then
fi fi
# FALLBACK_ENABLED gate — no point testing if fallback is disabled # FALLBACK_ENABLED gate — no point testing if fallback is disabled
if [[ "${FALLBACK_ENABLED:-false}" == false ]]; then # Fail-closed, matching fallback.sh — anything not exactly "true" counts as disabled.
if [[ "${FALLBACK_ENABLED:-false}" != "true" ]]; then
warn "FALLBACK_ENABLED=false — fallback test aborted" warn "FALLBACK_ENABLED=false — fallback test aborted"
warn "Enable fallback in master.conf before running this test" warn "Enable fallback in master.conf before running this test"
exit 0 exit 0
@@ -238,6 +275,11 @@ echo "━━━━━━━━━━━━━━━━━━━━━━━━
echo "" echo ""
echo "━━━ $ICON_SHIELD Phase 1 — Pre-flight ━━━" echo "━━━ $ICON_SHIELD Phase 1 — Pre-flight ━━━"
# Must run before the reachability check below — a stale DROP rule from a killed run makes
# the partner look unreachable, and the test would abort blaming the network for its own
# leftover.
_clear_stale_block
# Remote reachable # Remote reachable
if ping_remote; then if ping_remote; then
log "$REMOTE_SERVER_NAME is reachable" log "$REMOTE_SERVER_NAME is reachable"
+52 -5
View File
@@ -64,11 +64,45 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# acquire_lock "wait" — wait if previous run still active # Root Enforcement
# detect_hosts() — correct folder lists per host via MY_ID aliases # Media files are owned by container users; deleting them requires root.
# Empty array guards — warns and exits cleanly if no folders or patterns configured #
# Folder existence — skips missing folders with warning, continues others # Profile Required
# platform_require_cmd — notify script validated before use # Exits with usage if no profile is given. There is no default profile — an
# unspecified profile must never fall through to cleaning something.
#
# Lock Acquisition
# acquire_lock "wait" — waits for a previous run to finish rather than
# skipping, so a long anime pass does not cause the media pass to be dropped.
#
# Host Detection
# detect_hosts() aliases HOST*_ANIME_CLEAN_FOLDERS / HOST*_MEDIA_CLEAN_FOLDERS
# to the correct host's values.
#
# Empty Array Guards
# Exits cleanly if the resolved folder list or pattern list is empty. An empty
# pattern list would otherwise build a find with no -iname terms and match
# every file in the tree.
#
# Clean Path Depth Guard
# Every folder must be an absolute path at least three levels deep before it is
# scanned. The patterns include *.sh, *.zip, *.rar and *.exe, so a truncated
# entry like /mnt/user — which passes an existence check — would delete
# matching files across every share on the array.
#
# Folder Existence
# Missing folders are skipped with a warning; remaining folders still process.
#
# Explicit Pattern List
# Only patterns named in ANIME_FILE_PATTERNS / MEDIA_FILE_PATTERNS are removed.
# The script never infers intent from file size, age, or location.
#
# Count Before Delete
# Matching files are counted first; a folder with zero matches short-circuits
# before any rm is constructed.
#
# Dry Run Support
# --dry-run lists every file that would be deleted and removes nothing.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -205,6 +239,19 @@ for FOLDER in "${CLEAN_FOLDERS[@]}"; do
FOLDER_NAME=$(basename "$FOLDER") FOLDER_NAME=$(basename "$FOLDER")
echo "━━━ $ICON_CLEAN $FOLDER_NAME ━━━" echo "━━━ $ICON_CLEAN $FOLDER_NAME ━━━"
# The pattern list includes *.sh, *.zip, *.rar and *.exe. A truncated entry such as
# /mnt/user passes the -d check below and would sweep every share on the array, so
# require an absolute path at least three levels deep before scanning anything.
_depth="${FOLDER//[^\/]/}"
if [[ -z "$FOLDER" || "$FOLDER" != /* || "${#_depth}" -lt 3 ]]; then
error "Refusing to clean unsafe path: '${FOLDER:-empty}' — expected an absolute path at least 3 levels deep"
notify "Media cleaner ($PROFILE) refused unsafe path on $(hostname): '${FOLDER:-empty}'" \
"Media Cleaner" "warning"
FAILED+=("${FOLDER_NAME:-empty}")
echo ""
continue
fi
if [[ ! -d "$FOLDER" ]]; then if [[ ! -d "$FOLDER" ]]; then
warn "$FOLDER_NAME not found — skipping" warn "$FOLDER_NAME not found — skipping"
SKIPPED+=("$FOLDER_NAME") SKIPPED+=("$FOLDER_NAME")
+83 -10
View File
@@ -14,6 +14,30 @@
# or unRAID environment resets after updates. # or unRAID environment resets after updates.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# For each share in MEDIA_PERMISSION_SHARES:
#
# 1. Path safety and existence
# → unsafe or missing paths are refused or skipped, never scanned
#
# 2. Count wrong ownership (diagnostic)
# → find ! -user / ! -group — the number reported as "corrected"
#
# 3. Ownership pass — only if the count is non-zero
# → chown PERMISSIONS_OWNER on non-matching entries only
#
# 4. Directory mode pass
# → chmod PERMISSIONS_DIR_MODE on directories not already at that mode
#
# 5. File mode pass
# → chmod PERMISSIONS_FILE_MODE on files not already at that mode
# → "No such file" errors ignored: volatile dirs (Emby transcodes) race
#
# Every pass is conditional by design — see Conditional Passes below.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -35,16 +59,53 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# acquire_lock "wait" — wait if previous run still active (large share scans) # Root Enforcement
# detect_hosts() — correct share list per host via MY_ID aliases # chown to an arbitrary owner requires root.
# Empty array guard — warns and exits cleanly if no shares configured #
# Folder existence — skips missing shares with warning, continues others # Lock Acquisition
# Separate passes — directories and files chmod'd separately for correctness # acquire_lock "wait" — waits rather than skipping. Share scans are long, and
# Conditional passes — only entries whose owner/mode is actually wrong are touched. # this runs first in the daily window; skipping it would let arr cleanup run
# chown/chmod restamp ctime even when the value doesn't change, # against uncorrected ownership.
# and the arr cleanups gate orphan deletion on ctime #
# platform_require_cmd — notify script validated before use # Host Detection
# Silent by default — only failures produce output, success is silent # detect_hosts() aliases HOST*_MEDIA_PERMISSION_SHARES to this host's shares.
#
# Empty Array Guard
# Exits cleanly if no shares are configured for this host.
#
# Share Path Depth Guard
# Every share must be an absolute path at least three levels deep before it is
# scanned. A truncated entry like /mnt/user passes an existence check and would
# chown and chmod every share on the array — which, because chown/chmod restamp
# ctime, would erase the age signal the arr cleanups depend on across the whole
# library in a single run.
#
# Folder Existence
# Missing shares are skipped with a warning; remaining shares still process.
#
# Separate Passes
# Directories and files are chmod'd in separate passes — directories need the
# execute bit for traversal, media files must not have it.
#
# Conditional Passes
# Only entries whose owner or mode is actually wrong are touched. This is not
# an optimisation: chown/chmod rewrite an inode's ctime even when the value is
# unchanged, so a blanket pass would restamp every file nightly and destroy
# ctime as an age signal. The arr cleanups gate orphan deletion on ctime, and
# mtime cannot substitute — imports preserve the release's original timestamp.
# Making any pass unconditional silently stops orphan collection.
#
# Transcode Race Tolerance
# "No such file or directory" errors from the file pass are ignored. Volatile
# directories such as Emby transcodes delete files mid-scan; that is expected,
# not a permissions failure.
#
# Dry Run Support
# --dry-run counts the dirs, files and ownership entries that would change and
# modifies nothing.
#
# Silent by Default
# Only failures and diagnostics produce output; a clean run is quiet.
# #
# Diagnostic — high corrected count on every run means a container has wrong PUID/PGID: # Diagnostic — high corrected count on every run means a container has wrong PUID/PGID:
# Correct values on unRAID: PUID=99 (nobody) PGID=100 (users) # Correct values on unRAID: PUID=99 (nobody) PGID=100 (users)
@@ -153,6 +214,18 @@ PERMISSIONS_GROUP="${PERMISSIONS_OWNER##*:}"
for SHARE in "${MEDIA_PERMISSION_SHARES[@]}"; do for SHARE in "${MEDIA_PERMISSION_SHARES[@]}"; do
SHARE_NAME=$(basename "$SHARE") SHARE_NAME=$(basename "$SHARE")
# A truncated entry such as /mnt/user passes the -d check below and would chown/chmod
# every share on the array. Because chown/chmod restamp ctime, that would erase the age
# signal the arr cleanups gate orphan deletion on — across the whole library, in one run.
_depth="${SHARE//[^\/]/}"
if [[ -z "$SHARE" || "$SHARE" != /* || "${#_depth}" -lt 3 ]]; then
error "Refusing to touch unsafe path: '${SHARE:-empty}' — expected an absolute path at least 3 levels deep"
notify "Media permissions refused unsafe path on $(hostname): '${SHARE:-empty}'" \
"Media Permissions" "warning"
FAILED+=("${SHARE_NAME:-empty}")
continue
fi
if [[ ! -d "$SHARE" ]]; then if [[ ! -d "$SHARE" ]]; then
warn "$SHARE_NAME not found — skipping" warn "$SHARE_NAME not found — skipping"
SKIPPED+=("$SHARE_NAME") SKIPPED+=("$SHARE_NAME")
+72 -11
View File
@@ -18,7 +18,7 @@
# Audio → MusicBrainz Track ID # Audio → MusicBrainz Track ID
# #
# ============================================================================================== # ==============================================================================================
# SYNC LOGIC # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# For each matched item across ≥2 servers: # For each matched item across ≥2 servers:
@@ -55,21 +55,71 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# PLAY_SYNC_ENABLED gate — exits cleanly when disabled; no partial runs # Root Enforcement
# PARTNERSHIP gate — skips remote sync when partnership is inactive # The probe fingerprint is written under STATE_DIR, which is not user-writable.
# Per-server reachability — unreachable servers are skipped individually; # Without root the fingerprint silently fails to persist and the change probe
# one offline server does not abort the entire sync # never suppresses anything.
# User match required — a user missing from a server is skipped for that #
# server; no cross-account state pollution # jq Dependency Check
# acquire_lock — prevents concurrent runs from racing on the same # Exits if jq is missing. All API response parsing and the epoch comparisons
# items during the 30-minute critical window # depend on it — without jq every comparison would silently evaluate empty.
#
# PLAY_SYNC_ENABLED Gate
# Exits cleanly when disabled; no partial runs.
#
# PLAY_SYNC_REMOTE Gate
# When false, only this host's own servers are synced. Remote hosts are skipped
# before any network call is attempted.
#
# Partnership Gate
# Remote hosts are skipped when PARTNERSHIP_ENABLED=false. Local Emby↔Jellyfin
# sync still runs — a dormant partnership does not disable local work.
#
# Tailscale Resolution Guard
# A remote host whose Tailscale IP cannot be resolved is skipped rather than
# contacted at its literal localhost URL, which would otherwise point the sync
# at this host's own server and cross-contaminate state.
#
# Placeholder Credential Guard
# Servers whose API key is empty or still a placeholder are dropped from the
# list before any request is made.
#
# Per-Server Reachability
# Unreachable servers are skipped individually; one offline server does not
# abort the entire sync.
#
# User Match Required
# A user missing from a server is skipped for that server. State is never
# written to an unrelated account that happens to exist there.
#
# Forward-Only Writes
# The sync only pushes state forward — it never clears a Played flag or resets
# a resume position. The worst outcome of a bad comparison is a no-op, not
# erased watch history.
#
# Lock Acquisition
# acquire_lock prevents concurrent runs racing on the same items during the
# 30 minute critical window. --wait switches from skip to wait for manual runs.
#
# Probe Staleness Ceiling
# PLAY_SYNC_PROBE_MAX_AGE_HOURS forces a full comparison regardless of the
# hash. Fetches happen every run either way, so the probe can only skip
# per-item processing — it can never cause a change to be missed outright.
#
# Dry Run Support
# --dry-run performs all comparisons and writes no state.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION (host*.conf, aliased by detect_hosts) # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# host*.conf
#
# HOST*_TRANSCODE_SERVERS "Name|URL|APIKey|type" entries per host (emby/jellyfin) # HOST*_TRANSCODE_SERVERS "Name|URL|APIKey|type" entries per host (emby/jellyfin)
# All hosts are discovered automatically — no extra config needed. # Read directly for every HOST[0-9]+ defined — this script
# deliberately does NOT call detect_hosts(), because it needs
# every host's servers, not just this one's. Self is identified
# by comparing HOST* values against hostname -s.
# Remote host URLs have localhost rewritten to their Tailscale IP. # Remote host URLs have localhost rewritten to their Tailscale IP.
# #
# master.conf # master.conf
@@ -103,6 +153,10 @@
# play_state_sync.sh --full # play_state_sync.sh --full
# Bypass the change probe — always run the full comparison. # Bypass the change probe — always run the full comparison.
# #
# play_state_sync.sh --wait
# Wait for an in-progress run to finish instead of exiting. For manual runs
# that would otherwise be skipped by the every-30-minute scheduled pass.
#
# play_state_sync.sh --log # play_state_sync.sh --log
# Verbose output — show each item comparison. # Verbose output — show each item comparison.
# #
@@ -129,6 +183,13 @@ parse_args "${_FILTERED[@]}"
# ============================================================================================== # ==============================================================================================
# ━━━ Setup ━━━ # ━━━ Setup ━━━
# ============================================================================================== # ==============================================================================================
# The probe fingerprint lives under STATE_DIR — without root it silently fails to persist
# and the change probe can never suppress a run.
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
[[ "${PLAY_SYNC_ENABLED:-true}" != "true" ]] && echo "Play state sync disabled" && exit 0 [[ "${PLAY_SYNC_ENABLED:-true}" != "true" ]] && echo "Play state sync disabled" && exit 0
SYNC_TYPES="${PLAY_SYNC_TYPES:-Movie,Episode}" SYNC_TYPES="${PLAY_SYNC_TYPES:-Movie,Episode}"
+18 -3
View File
@@ -19,6 +19,24 @@
# to HOST*_DAILY_SYNC_SHARES. Both aliased by detect_hosts(). # to HOST*_DAILY_SYNC_SHARES. Both aliased by detect_hosts().
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Pre-flight — connectivity to the remote, remote array mounted, version parity
# 2. Resolve the share list (BACKUP_VERIFY_SHARES, else DAILY_SYNC_SHARES)
# 3. Per share:
# a. Randomly sample BACKUP_VERIFY_SAMPLE files above BACKUP_VERIFY_MIN_SIZE
# b. Compute each file's MD5 locally
# c. Compute the same file's MD5 on the remote over SSH
# d. Classify: MATCH | MISMATCH | MISSING
# 4. Report per-share and overall counts; notify on any MISMATCH and on
# significant MISSING counts
#
# Sampling rather than full verification is deliberate — a complete checksum of every
# mirrored file would take longer than the interval between runs. Random sampling over
# a weekly cadence surfaces systematic corruption without ever reading the whole library.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -56,9 +74,6 @@
# SSH Timeout # SSH Timeout
# SSH_TIMEOUT caps all SSH calls. One hung connection does not block the run. # SSH_TIMEOUT caps all SSH calls. One hung connection does not block the run.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
+18 -3
View File
@@ -17,6 +17,24 @@
# last 7 days activity timeline, any transfers or days exceeding BANDWIDTH_WARN_GB. # last 7 days activity timeline, any transfers or days exceeding BANDWIDTH_WARN_GB.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Log mode (--log-transfer), called by rsync.sh after every sync:
# 1. Append one line: YYYY-MM-DD|HH:MM|profile|duration|status|bytes
# 2. Trim entries older than BANDWIDTH_LOG_RETENTION days
# One bounded write per rsync run — never grows without limit, never rewrites history.
#
# Report mode (default), scheduled weekly:
# 1. Read the accumulated log
# 2. Aggregate per profile — run count, total bytes, average duration, failures
# 3. Build a 7-day activity timeline
# 4. Flag any single transfer or any single day exceeding BANDWIDTH_WARN_GB
#
# The two modes never run together: logging is a side effect of rsync, reporting is a
# scheduled read. Report mode never writes to the log.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -46,9 +64,6 @@
# Log Directory Guard # Log Directory Guard
# Creates the log directory if it doesn't exist. Exits cleanly if unwritable. # Creates the log directory if it doesn't exist. Exits cleanly if unwritable.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
+21 -2
View File
@@ -15,6 +15,23 @@
# separate message lists all CRITICAL domains. Not one notification per domain. # separate message lists all CRITICAL domains. Not one notification per domain.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Validate openssl is present (platform_require_cmd) — without it nothing can be checked
# 2. Per domain in CERT_MONITOR_DOMAINS:
# a. Open a real TLS connection with openssl s_client
# b. Parse notAfter from the served certificate
# c. Compute days remaining
# d. Classify: HEALTHY (silent) | WARNING (≤ CERT_WARN_DAYS)
# | CRITICAL (≤ CERT_CRIT_DAYS) | FAILED (no connect / no parse)
# 3. Batch by severity — one notification listing all WARNING domains, a separate
# one listing all CRITICAL domains
#
# A domain that fails to connect is reported as FAILED rather than assumed healthy or
# assumed expired — an unreachable host and an expiring cert are different problems.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -44,8 +61,10 @@
# CERT_TIMEOUT caps each openssl connection attempt. One unreachable domain # CERT_TIMEOUT caps each openssl connection attempt. One unreachable domain
# does not block the remaining domains. # does not block the remaining domains.
# #
# Notification Validated # openssl Validated
# platform_require_cmd confirms openssl and notify script are present before use. # platform_require_cmd confirms openssl is present before any domain is checked — every
# check depends on it, so a missing binary is reported as itself rather than as every
# domain failing. The notify script is validated separately by the platform adapter.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
+16 -3
View File
@@ -18,6 +18,22 @@
# configuration issue. Silent on clean runs. # configuration issue. Silent on clean runs.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Reachability — Emby API responding; unreachable exits cleanly rather than
# reporting an empty library as a real result
# 2. Server info and uptime
# 3. Active sessions — count, and the transcode-to-direct-play ratio
# 4. Library counts — movies, episodes, songs
# 5. Activity history over the last EMBY_REPORT_DAYS
# 6. Top EMBY_REPORT_TOP_N items and most active users
# 7. Ramdisk transcode status, read from the shared transcode state
#
# Every figure is queried fresh. The only notification is the transcode-ratio warning;
# everything else is report output.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -54,9 +70,6 @@
# detect_hosts() aliases HOST*_EMBY_URL and HOST*_EMBY_API_KEY → EMBY_URL / EMBY_API_KEY. # detect_hosts() aliases HOST*_EMBY_URL and HOST*_EMBY_API_KEY → EMBY_URL / EMBY_API_KEY.
# Each server reports on its own Emby instance automatically. # Each server reports on its own Emby instance automatically.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
+7
View File
@@ -53,6 +53,13 @@
# collect_hosts() populates ALL_HOST_IDS — if no HOST* vars are defined the # collect_hosts() populates ALL_HOST_IDS — if no HOST* vars are defined the
# output sections iterate over an empty array and exit cleanly. # output sections iterate over an empty array and exit cleanly.
# #
# No Root, No Lock, No detect_hosts — Deliberate
# This is the one script in the ecosystem that intentionally omits all three, and
# they should not be added. It writes nothing, so there is no state for a lock to
# protect and no privileged operation to justify a root gate. It reports on every
# node rather than acting as one, so detect_hosts() would narrow it to this host's
# aliases — the opposite of what it is for. Every HOST* var is read directly instead.
#
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
+23 -2
View File
@@ -19,6 +19,26 @@
# from master.conf if dynamix.cfg is not found. # from master.conf if dynamix.cfg is not found.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Validate smartctl is present (platform_require_cmd)
# 2. Resolve temperature thresholds — dynamix.cfg first, master.conf as fallback
# 3. Enumerate drives, skipping anything in SMART_IGNORE_DRIVES
# 4. Per drive, read live SMART attributes and evaluate:
# overall status FAILED → critical
# Reallocated_Sector_Ct > 0 → concerning
# Current_Pending_Sector > 0 → concerning
# Offline_Uncorrectable > 0 → critical
# Temperature_Celsius vs warn/crit thresholds
# Power_On_Hours → informational only
# NVMe drives expose different attribute names — detected and mapped automatically.
# 5. Report; notify only when something crosses a threshold. Silent when all pass.
#
# Read-only throughout — this queries attributes the drive already maintains and never
# starts a self-test. Running one is smart_long_test.sh's job.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -55,8 +75,9 @@
# Reads hot/max/hotssd/maxssd from dynamix.cfg so smart_health.sh and unRAID's # Reads hot/max/hotssd/maxssd from dynamix.cfg so smart_health.sh and unRAID's
# dashboard use the same thresholds. Falls back to master.conf values if not found. # dashboard use the same thresholds. Falls back to master.conf values if not found.
# #
# Notifications Validated # smartctl Validated
# platform_require_cmd confirms smartctl and notify script are present before use. # platform_require_cmd confirms smartctl is present before any drive is queried. The
# notify script is validated separately by the platform adapter.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
-3
View File
@@ -65,9 +65,6 @@
# Trim uses tmp file + mv — partial writes during log rotation cannot corrupt # Trim uses tmp file + mv — partial writes during log rotation cannot corrupt
# the accumulated history. # the accumulated history.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# STATE FILES # STATE FILES
# ============================================================================================== # ==============================================================================================
+4 -2
View File
@@ -76,8 +76,10 @@
# smart profile produces no output and no notification when nothing worth # smart profile produces no output and no notification when nothing worth
# reporting is found. # reporting is found.
# #
# Notifications Validated # openssl Validated — Non-Fatal
# platform_require_cmd confirms notify and openssl are present before use. # platform_require_cmd checks openssl and, unlike the other monitors, only warns if it
# is missing: the SSL section is skipped and the rest of the digest still runs. The
# notify script is validated separately by the platform adapter.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
-3
View File
@@ -79,9 +79,6 @@
# Docker Stats Timeout # Docker Stats Timeout
# DOCKER_TIMEOUT caps docker stats calls. A hung daemon does not block the report. # DOCKER_TIMEOUT caps docker stats calls. A hung daemon does not block the report.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# STATE FILES # STATE FILES
# ============================================================================================== # ==============================================================================================
+37 -7
View File
@@ -55,13 +55,33 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — all launched scripts require root # Root Enforcement
# acquire_lock — prevents duplicate array start launches # Every script launched here requires root.
# detect_hosts() — MY_ID in notifications #
# platform_require_cmd — notify validated before use # Lock Acquisition
# chmod +x auto-fix — non-executable scripts fixed before launch # acquire_lock prevents duplicate array start launches. Unraid can fire the array
# Full path on failure — shows exact path for debugging # start hook more than once, and a second pass would re-launch continuous scripts
# notify on failures — alert if any script fails to launch # that are already running.
#
# Host Detection
# detect_hosts() sets MY_ID for notifications.
#
# Empty Job List Guard
# Exits with an error and a notification if ARRAY_START_SCRIPTS is empty. An empty
# list would silently bring the array up with no ramdisk, no network setup, no
# watchdogs and no fallback — while reporting a clean start.
#
# Executable Auto-Fix
# Non-executable scripts are chmod +x'd before launch. A permission bit lost to a
# git checkout or a file copy should not silently disable a boot-time component.
#
# Full Path on Failure
# Failures report the exact resolved path, so a missing script is immediately
# distinguishable from a script that ran and failed.
#
# Failure Notification
# Any script that fails to launch raises a notification — array start is unattended,
# so a silent failure here would only surface much later as a missing service.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -109,6 +129,16 @@ acquire_lock
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#ARRAY_START_SCRIPTS[@]} -eq 0 ]]; then
error "ARRAY_START_SCRIPTS is empty — no array start scripts will run"
error "Check ARRAY_START_SCRIPTS in master.conf"
notify "array start scripts skipped on $(hostname) ($MY_ID) — ARRAY_START_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — scripts will not be launched" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — scripts will not be launched"
# ============================================================================================== # ==============================================================================================
+32 -5
View File
@@ -40,11 +40,28 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — all stop scripts require root # Root Enforcement
# acquire_lock — prevents concurrent array stop runs # Every stop script launched here requires root.
# detect_hosts() — MY_ID in notifications and logs #
# platform_require_cmd — notify validated before use # Lock Acquisition
# notify on failures — alert if any stop script fails # acquire_lock prevents concurrent array stop runs. Two overlapping shutdown
# sequences would fight over the same containers.
#
# Host Detection
# detect_hosts() sets MY_ID for notifications and logs.
#
# Empty Job List Guard
# Exits with an error and a notification if ARRAY_STOP_SCRIPTS is empty. An empty
# list means the array stops without saving the conf cache or gracefully stopping
# containers — the failure would only be discovered at the next boot.
#
# Non-Fatal Steps
# A failing stop script is recorded and the remaining ones still run. Abandoning the
# shutdown sequence partway would leave more state unsaved than continuing does.
#
# Failure Notification
# Any failing stop script raises a notification. Shutdown is unattended and its
# failures are invisible until they cause a problem on the way back up.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -97,6 +114,16 @@ fi
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#ARRAY_STOP_SCRIPTS[@]} -eq 0 ]]; then
error "ARRAY_STOP_SCRIPTS is empty — no array stop scripts will run"
error "Check ARRAY_STOP_SCRIPTS in master.conf"
notify "array stop scripts skipped on $(hostname) ($MY_ID) — ARRAY_STOP_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no stop scripts will be executed" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no stop scripts will be executed"
# ============================================================================================== # ==============================================================================================
+35 -5
View File
@@ -44,11 +44,31 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — rsync and container stop/start require root # Root Enforcement
# acquire_lock "strict" — no pile-up; skip cycle if prior run still active # rsync over SSH and container stop/start both require root.
# detect_hosts() — MY_ID and REMOTE_ID for routing and logs #
# resolve_remote_ip — confirms remote reachability before any transfer # Lock Acquisition
# RSYNC_ENABLED gate — global kill switch respected before any rsync call # acquire_lock in strict mode — a cycle is skipped rather than queued if the previous
# one is still running. At a 30-minute cadence, queuing would let a slow sync stack
# windows behind it.
#
# Host Detection
# detect_hosts() sets MY_ID and REMOTE_ID for routing and logs.
#
# Empty Job List Guard
# Exits with an error and a notification if CRITICAL_MAINTENANCE_SCRIPTS is empty —
# a silently empty critical tier would stop downloader resets and play-state sync
# while still reporting success every 30 minutes.
#
# Remote IP Resolution
# resolve_remote_ip confirms the partner is reachable before any transfer is attempted.
#
# RSYNC_ENABLED Gate
# The global kill switch is respected before any rsync call, so disabling rsync
# ecosystem-wide genuinely stops it here too.
#
# Non-Fatal Steps
# A failing job is recorded and the rest of the tier still runs.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -98,6 +118,16 @@ fi
acquire_lock "strict" acquire_lock "strict"
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#CRITICAL_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
error "CRITICAL_MAINTENANCE_SCRIPTS is empty — no critical maintenance scripts will run"
error "Check CRITICAL_MAINTENANCE_SCRIPTS in master.conf"
notify "critical maintenance scripts skipped on $(hostname) ($MY_ID) — CRITICAL_MAINTENANCE_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
resolve_remote_ip resolve_remote_ip
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
+44 -7
View File
@@ -59,13 +59,40 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — rsync and docker operations require root # Root Enforcement
# acquire_lock — prevents concurrent daily windows # rsync and docker operations require root.
# detect_hosts() — aliases correct per-host share lists #
# check_connectivity — verified before any rsync # Lock Acquisition
# check_remote_rootfs — aborts rsync if remote rootfs nearly full # acquire_lock prevents two daily windows overlapping — the window is long and a
# Non-fatal jobs — a failed job logs and continues; remaining jobs still run # second pass would contend for the same shares and containers.
# notify on failure — successful daily run produces no notification #
# Host Detection
# detect_hosts() aliases the correct per-host share and script lists.
#
# Empty Job List Guard
# Exits with an error and a notification if DAILY_MAINTENANCE_SCRIPTS is empty. This
# is the largest tier in the ecosystem — an empty list would silently skip git pull,
# permissions, cleaners, arr cleanup and docker updates while reporting a clean run.
#
# Connectivity Check
# check_connectivity is verified before any rsync is attempted.
#
# Remote Rootfs Check
# check_remote_rootfs aborts rsync if the remote rootfs is nearly full, rather than
# pushing data to a partner that cannot hold it.
#
# Drive Temperature Escalation
# rsync.sh's exit code is honoured per share: exit 1 (temp WARN) skips that share and
# continues; exit 2 (temp CRITICAL) sets ABORT_ALL_SYNCS so every remaining share in
# the window is skipped and a notification is raised. Continuing to hammer drives that
# are already too hot is how a thermal warning becomes a dead disk.
#
# Non-Fatal Jobs
# A failed job is logged and the remaining jobs still run. Partial completion of a
# maintenance window beats abandoning it at the first error.
#
# Quiet on Success
# A successful daily run produces no notification — only failures surface.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -123,6 +150,16 @@ if ! command -v docker &>/dev/null; then
fi fi
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#DAILY_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
error "DAILY_MAINTENANCE_SCRIPTS is empty — no daily maintenance scripts will run"
error "Check DAILY_MAINTENANCE_SCRIPTS in master.conf"
notify "daily maintenance scripts skipped on $(hostname) ($MY_ID) — DAILY_MAINTENANCE_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
resolve_remote_ip resolve_remote_ip
acquire_lock acquire_lock
+40 -7
View File
@@ -40,13 +40,36 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — scripts called here require root # Root Enforcement
# acquire_lock — prevents concurrent intermediate windows # Every script called from here requires root.
# detect_hosts() — aliases correct per-host share lists #
# check_connectivity — verified before any rsync (skipped if no shares) # Lock Acquisition
# check_remote_rootfs — aborts rsync if remote rootfs nearly full # acquire_lock prevents concurrent intermediate windows.
# Non-fatal jobs — a failed arr_sync warns but does not block rsync or artwork fetch #
# Minimal on success — runs 6x/day; full breakdown only on failure or --log # Host Detection
# detect_hosts() aliases the correct per-host share lists.
#
# Empty Job List Guard
# Exits with an error and a notification if INTERMEDIATE_MAINTENANCE_SCRIPTS is empty,
# rather than running six no-op windows a day that all report success.
#
# Connectivity Check
# check_connectivity is verified before any rsync, and skipped entirely when no shares
# are configured — there is nothing to reach a partner for.
#
# Remote Rootfs Check
# check_remote_rootfs aborts rsync if the remote rootfs is nearly full.
#
# Drive Temperature Escalation
# rsync.sh's exit code is honoured per share: exit 1 skips that share, exit 2 aborts
# every remaining sync in the window and notifies.
#
# Non-Fatal Jobs
# A failing arr_sync warns but does not block rsync or the artwork fetch that follow it.
#
# Minimal on Success
# Runs six times a day, so the full breakdown only prints on failure or with --log.
# A quiet run is the normal outcome and should not fill the log.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -101,6 +124,16 @@ fi
acquire_lock acquire_lock
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#INTERMEDIATE_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
error "INTERMEDIATE_MAINTENANCE_SCRIPTS is empty — no intermediate maintenance scripts will run"
error "Check INTERMEDIATE_MAINTENANCE_SCRIPTS in master.conf"
notify "intermediate maintenance scripts skipped on $(hostname) ($MY_ID) — INTERMEDIATE_MAINTENANCE_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
resolve_remote_ip resolve_remote_ip
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
+38 -6
View File
@@ -45,12 +45,34 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — ZFS scrub, SMART tests require root # Root Enforcement
# acquire_lock — prevents concurrent monthly runs # ZFS scrub and SMART tests require root.
# detect_hosts() — MY_ID in notifications and logs #
# Uptime gate — MONTHLY_UPTIME_THRESHOLD_DAYS must be met # Lock Acquisition
# Interval gate — MONTHLY_RUN_INTERVAL_DAYS since last run must be met # acquire_lock prevents concurrent monthly runs. These are long jobs — a scrub can run
# --force flag — bypasses both gates for manual override # for hours — and two at once would double the I/O cost for no benefit.
#
# Host Detection
# detect_hosts() sets MY_ID for notifications and logs.
#
# Empty Job List Guard
# Exits with an error and a notification if MONTHLY_MAINTENANCE_SCRIPTS is empty. A
# monthly job that silently does nothing is the hardest kind to notice missing.
#
# Uptime Gate
# MONTHLY_UPTIME_THRESHOLD_DAYS must be met before the run proceeds. Heavy full-disk
# work immediately after a boot competes with everything else still starting up.
#
# Interval Gate
# MONTHLY_RUN_INTERVAL_DAYS since the last successful run must have elapsed. The
# schedule fires more often than the work should actually happen, so the gate — not
# the cron entry — is what defines the real cadence.
#
# Force Override
# --force bypasses both gates for a deliberate manual run.
#
# Non-Fatal Steps
# A failing job is recorded and the remaining jobs still run.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -113,6 +135,16 @@ acquire_lock
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#MONTHLY_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
error "MONTHLY_MAINTENANCE_SCRIPTS is empty — no monthly maintenance scripts will run"
error "Check MONTHLY_MAINTENANCE_SCRIPTS in master.conf"
notify "monthly maintenance scripts skipped on $(hostname) ($MY_ID) — MONTHLY_MAINTENANCE_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no scripts will be executed" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no scripts will be executed"
[[ "$FORCE_RUN" == true ]] && warn "FORCE — uptime and interval gates bypassed" [[ "$FORCE_RUN" == true ]] && warn "FORCE — uptime and interval gates bypassed"
+37 -6
View File
@@ -39,12 +39,33 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — required for the docker/system reads used in the report # Root Enforcement
# acquire_lock — prevents overlapping weekly runs # Required for the docker and system reads the report is built from.
# detect_hosts() — MY_ID in banner and summary #
# Non-fatal steps — a failed script is logged; remaining scripts still run # Lock Acquisition
# Flag pass-through — --dry-run and --log forwarded to all child scripts # acquire_lock prevents overlapping weekly runs.
# notify() on failure — pushed only outside --dry-run, matching the runtime-mode contract below #
# Host Detection
# detect_hosts() sets MY_ID for the banner and summary.
#
# Empty Job List Guard
# Exits with an error and a notification if COFFEE_REPORT_SCRIPTS is empty. A report
# that silently contains nothing still arrives looking like a report.
#
# Read-Only by Composition
# Every child here is a reporting script. This orchestrator changes nothing itself —
# it only sequences reads and assembles their output.
#
# Non-Fatal Steps
# A failing script is logged and the remaining ones still run, so one unavailable
# subsystem costs a section of the report rather than the whole thing.
#
# Flag Pass-Through
# --dry-run and --log are forwarded to every child script.
#
# Notification Contract
# notify() fires on failure only outside --dry-run, matching the runtime-mode contract
# below — a dry run never sends anything outward.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -91,6 +112,16 @@ acquire_lock
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#COFFEE_REPORT_SCRIPTS[@]} -eq 0 ]]; then
error "COFFEE_REPORT_SCRIPTS is empty — no coffee report scripts will run"
error "Check COFFEE_REPORT_SCRIPTS in master.conf"
notify "coffee report scripts skipped on $(hostname) ($MY_ID) — COFFEE_REPORT_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
# ============================================================================================== # ==============================================================================================
# ━━━ Helpers ━━━ # ━━━ Helpers ━━━
# ============================================================================================== # ==============================================================================================
+41 -6
View File
@@ -50,12 +50,37 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — mount and docker operations require root # Root Enforcement
# acquire_lock — prevents concurrent 7-minute cycles overlapping # Mount and docker operations require root.
# detect_hosts() — aliases RAMDISK_PATH, TRANSCODE_SSD, RAMDISK_WARN_GB per host #
# --dry-run — passed through to every script in TRANSCODE_MANAGEMENT_SCRIPTS # Lock Acquisition
# Exit code — worst exit code across all scripts returned to cron # acquire_lock prevents concurrent 7-minute cycles overlapping. Cleanup and the manager
# notify() — pushed on failure, skipped in --dry-run # both touch the same ramdisk, and two cycles at once could have one deleting files
# while the other is measuring usage to decide whether to flip.
#
# Host Detection
# detect_hosts() aliases RAMDISK_PATH, TRANSCODE_SSD and RAMDISK_WARN_GB per host.
#
# Empty Job List Guard
# Exits with an error and a notification if TRANSCODE_MANAGEMENT_SCRIPTS is empty —
# without it the ramdisk would silently stop being cleaned or flipped, and the first
# symptom would be a full ramdisk stalling playback.
#
# Ordering Is Load-Bearing
# Cleanup runs before the manager so the manager measures real active-session usage
# rather than usage inflated by stale files. Reversing them would trigger flips that
# a cleanup two seconds later would have made unnecessary.
#
# Dry Run Propagation
# --dry-run is passed through to every script in TRANSCODE_MANAGEMENT_SCRIPTS.
#
# Any-Failure Exit Code
# Exits 1 if any child failed, 0 otherwise — the individual exit codes are not
# propagated, only whether anything failed. A failure in an early child is therefore
# never masked by a later success.
#
# Notification Contract
# notify() fires on failure and is skipped in --dry-run.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -112,6 +137,16 @@ fi
# detect_hosts() sets MY_ID and aliases all HOST*_TRANSCODE_* vars # detect_hosts() sets MY_ID and aliases all HOST*_TRANSCODE_* vars
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#TRANSCODE_MANAGEMENT_SCRIPTS[@]} -eq 0 ]]; then
error "TRANSCODE_MANAGEMENT_SCRIPTS is empty — no transcode management scripts will run"
error "Check TRANSCODE_MANAGEMENT_SCRIPTS in master.conf"
notify "transcode management scripts skipped on $(hostname) ($MY_ID) — TRANSCODE_MANAGEMENT_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — passing through to child scripts" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — passing through to child scripts"
# ============================================================================================== # ==============================================================================================
+38 -5
View File
@@ -49,11 +49,34 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — watchdog operations require root # Root Enforcement
# acquire_lock — strict; no pile-up if prior cycle still active # Every watchdog launched from here requires root. Failing once at the top gives one
# detect_hosts() — MY_ID in logs and notifications # clear error instead of the same permission failure repeated per child.
# Array check — exits early if /mnt/user is not shfs-mounted #
# Startup grace — WATCHDOG_STARTUP_GRACE respected before any checks # Lock Acquisition
# acquire_lock in strict mode — if the previous cycle is still running, this one exits
# rather than queuing. At a one-minute cadence a waiting lock would pile up cycles
# behind a slow watchdog and eventually run them all at once.
#
# Host Detection
# detect_hosts() sets MY_ID for logs and notifications.
#
# Empty Job List Guard
# Exits with an error and a notification if WATCHDOG_ORCHESTRATOR_SCRIPTS is empty.
# Without it the cycle reports "0/0 passed" and exits 0 every minute — indistinguishable
# from a healthy run, while nothing at all is being monitored.
#
# Array Check
# Exits early if /mnt/user is not shfs-mounted. Watchdogs that inspect shares would
# otherwise read an unmounted array as missing data and act on it.
#
# Startup Grace
# WATCHDOG_STARTUP_GRACE is respected before any checks run, so containers still
# initialising after boot are not judged as unhealthy.
#
# Non-Fatal Steps
# run_orch_child() records a failing or missing watchdog and continues. One broken
# watchdog never suppresses the rest of the chain.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -104,6 +127,16 @@ acquire_lock
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#WATCHDOG_ORCHESTRATOR_SCRIPTS[@]} -eq 0 ]]; then
error "WATCHDOG_ORCHESTRATOR_SCRIPTS is empty — no watchdogs will run"
error "Check WATCHDOG_ORCHESTRATOR_SCRIPTS in master.conf"
notify "watchdogs skipped on $(hostname) ($MY_ID) — WATCHDOG_ORCHESTRATOR_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
log "$ICON_GEAR Config: grace=${WATCHDOG_STARTUP_GRACE}s heartbeat=${WATCHDOG_ORCHESTRATOR_HEARTBEAT:-true}/${WATCHDOG_ORCHESTRATOR_HEARTBEAT_HOURS:-1}hr scripts=${#WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}" log "$ICON_GEAR Config: grace=${WATCHDOG_STARTUP_GRACE}s heartbeat=${WATCHDOG_ORCHESTRATOR_HEARTBEAT:-true}/${WATCHDOG_ORCHESTRATOR_HEARTBEAT_HOURS:-1}hr scripts=${#WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"
log "$ICON_WATCHDOG Order: $(for s in "${WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"; do printf '%s ' "${s##*/}"; done)" log "$ICON_WATCHDOG Order: $(for s in "${WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"; do printf '%s ' "${s##*/}"; done)"
+39 -9
View File
@@ -40,15 +40,35 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — stop/start containers and rsync require root # Root Enforcement
# acquire_lock — prevents concurrent weekly windows # Container stop/start and rsync both require root.
# detect_hosts() — MY_ID in banner, summary, and notifications #
# check_connectivity — verifies remote before any remote operations # Lock Acquisition
# check_remote_rootfs — aborts rsync if remote rootfs nearly full # acquire_lock prevents concurrent weekly windows. This window stops Emby and the auth
# DOCKER_TIMEOUT — all docker calls protected # stack — two overlapping runs would fight over the same critical containers.
# SSH_TIMEOUT — all SSH calls protected #
# platform_require_cmd — notify validated before use # Host Detection
# Silent on success — runs weekly; only failures warrant notification # detect_hosts() sets MY_ID for the banner, summary and notifications.
#
# Empty Job List Guard
# Exits with an error and a notification if WEEKLY_MAINTENANCE_SCRIPTS is empty, rather
# than taking the weekly outage window and doing nothing with it.
#
# Connectivity Check
# check_connectivity verifies the remote before any remote operation is attempted.
#
# Remote Rootfs Check
# check_remote_rootfs aborts rsync if the remote rootfs is nearly full.
#
# Timeout Protection
# DOCKER_TIMEOUT bounds every docker call and SSH_TIMEOUT every SSH call, so neither a
# hung daemon nor an unresponsive partner can hold the weekly window open indefinitely.
#
# Non-Fatal Steps
# A failing job is recorded and the remaining jobs still run.
#
# Silent on Success
# Runs weekly; only failures warrant a notification.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -109,6 +129,16 @@ if ! command -v docker &>/dev/null; then
fi fi
detect_hosts detect_hosts
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
# from a healthy run. Fail loudly instead of silently doing no work.
if [[ ${#WEEKLY_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
error "WEEKLY_MAINTENANCE_SCRIPTS is empty — no weekly maintenance scripts will run"
error "Check WEEKLY_MAINTENANCE_SCRIPTS in master.conf"
notify "weekly maintenance scripts skipped on $(hostname) ($MY_ID) — WEEKLY_MAINTENANCE_SCRIPTS is empty" \
"$(basename "$0" .sh)" "warning"
exit 1
fi
resolve_remote_ip resolve_remote_ip
WINDOW_START=$(date +%s) WINDOW_START=$(date +%s)
+1
View File
@@ -27,6 +27,7 @@ $monitor = [
'mem' => vv_memory_breakdown(), 'mem' => vv_memory_breakdown(),
'net' => vv_network_stats(), 'net' => vv_network_stats(),
'gpu' => vv_gpu_stats(), 'gpu' => vv_gpu_stats(),
'gpus' => vv_gpu_stats_all(),
'gpu_procs' => vv_gpu_processes(), 'gpu_procs' => vv_gpu_processes(),
'containers' => vv_docker_containers(), 'containers' => vv_docker_containers(),
'stopped' => vv_docker_stopped(), 'stopped' => vv_docker_stopped(),
+1
View File
@@ -25,6 +25,7 @@ echo json_encode([
'mem' => vv_memory_breakdown(), 'mem' => vv_memory_breakdown(),
'net' => vv_network_stats(), 'net' => vv_network_stats(),
'gpu' => vv_gpu_stats(), 'gpu' => vv_gpu_stats(),
'gpus' => vv_gpu_stats_all(),
'gpu_procs' => vv_gpu_processes(), 'gpu_procs' => vv_gpu_processes(),
'containers' => vv_docker_containers(), 'containers' => vv_docker_containers(),
'stopped' => vv_docker_stopped(), 'stopped' => vv_docker_stopped(),
+37 -18
View File
@@ -95,35 +95,54 @@ function vv_docker_stopped(): array {
return $containers; return $containers;
} }
function vv_gpu_stats(): array { // Every installed GPU, one entry per card. Parsed line by line — nvidia-smi emits one row
$out = shell_exec('nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu,temperature.gpu,power.draw,utilization.encoder,utilization.decoder --format=csv,noheader,nounits 2>/dev/null'); // per GPU, so splitting the whole output on commas (as this once did) runs the rows together
if (!$out) return ['available' => false]; // and only ever describes GPU 0.
function vv_gpu_stats_all(): array {
$out = shell_exec('nvidia-smi --query-gpu=index,uuid,name,memory.used,memory.total,utilization.gpu,temperature.gpu,power.draw,utilization.encoder,utilization.decoder --format=csv,noheader,nounits 2>/dev/null');
if (!$out) return [];
$parts = array_map('trim', explode(',', $out)); $gpus = [];
$power = is_numeric($parts[5] ?? '') ? round((float)$parts[5], 1) : null; foreach (explode("\n", trim($out)) as $line) {
return [ if (!$line) continue;
$p = array_map('trim', explode(',', $line));
if (count($p) < 10) continue;
$gpus[] = [
'available' => true, 'available' => true,
'name' => $parts[0] ?? '', 'index' => (int)$p[0],
'memory_used' => (int)($parts[1] ?? 0), 'uuid' => $p[1],
'memory_total' => (int)($parts[2] ?? 0), 'name' => $p[2],
'utilization' => (int)($parts[3] ?? 0), 'memory_used' => (int)$p[3],
'temperature' => (int)($parts[4] ?? 0), 'memory_total' => (int)$p[4],
'power_w' => $power, 'utilization' => (int)$p[5],
'enc_pct' => (int)($parts[6] ?? 0), 'temperature' => (int)$p[6],
'dec_pct' => (int)($parts[7] ?? 0), 'power_w' => is_numeric($p[7]) ? round((float)$p[7], 1) : null,
'enc_pct' => (int)$p[8],
'dec_pct' => (int)$p[9],
]; ];
}
return $gpus;
} }
// Kept for the existing single-GPU consumers — same shape as before, always GPU 0.
function vv_gpu_stats(): array {
$gpus = vv_gpu_stats_all();
return $gpus[0] ?? ['available' => false];
}
// gpu_uuid is included so each process can be attributed to the card it is actually running
// on. Without it a two-GPU box shows every process under every card.
function vv_gpu_processes(): array { function vv_gpu_processes(): array {
$out = shell_exec('nvidia-smi --query-compute-apps=pid,used_gpu_memory,name --format=csv,noheader,nounits 2>/dev/null'); $out = shell_exec('nvidia-smi --query-compute-apps=gpu_uuid,pid,used_gpu_memory,name --format=csv,noheader,nounits 2>/dev/null');
$procs = []; $procs = [];
foreach (explode("\n", trim($out ?? '')) as $line) { foreach (explode("\n", trim($out ?? '')) as $line) {
if (!$line) continue; if (!$line) continue;
$parts = array_map('trim', explode(',', $line)); $parts = array_map('trim', explode(',', $line));
$procs[] = [ $procs[] = [
'pid' => $parts[0] ?? '', 'gpu_uuid' => $parts[0] ?? '',
'memory_mb' => $parts[1] ?? '', 'pid' => $parts[1] ?? '',
'name' => $parts[2] ?? '', 'memory_mb' => $parts[2] ?? '',
'name' => $parts[3] ?? '',
]; ];
} }
return $procs; return $procs;
+55 -11
View File
@@ -104,7 +104,7 @@
<div id="vv-docker-folders-body">Loading...</div> <div id="vv-docker-folders-body">Loading...</div>
</div> </div>
<!-- Row 3: Rsync | GPU | Transcode | Streams --> <!-- Row 3: Rsync | GPU 0 | GPU 1 | Transcode | Streams -->
<div class="vv-card" id="vv-rsync-card" style="grid-column:span 1;"> <div class="vv-card" id="vv-rsync-card" style="grid-column:span 1;">
<h3> <h3>
<span style="display:flex;align-items:center;gap:5px;"> <span style="display:flex;align-items:center;gap:5px;">
@@ -119,14 +119,25 @@
<h3> <h3>
<span style="display:flex;align-items:center;gap:5px;"> <span style="display:flex;align-items:center;gap:5px;">
<span class="vv-ico"><svg width="15" height="10" viewBox="0 0 16 10" fill="none" stroke="#aaa" stroke-width="1.1" stroke-linecap="round" stroke-linejoin="round"><rect x="0.8" y="0.8" width="14.4" height="7" rx="1.2"/><rect x="2.5" y="2.5" width="3.5" height="3.5" rx="0.5"/><line x1="8" y1="3" x2="13" y2="3"/><line x1="8" y1="5" x2="11" y2="5"/><rect x="3" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/><rect x="6.5" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/><rect x="10" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/></svg></span> <span class="vv-ico"><svg width="15" height="10" viewBox="0 0 16 10" fill="none" stroke="#aaa" stroke-width="1.1" stroke-linecap="round" stroke-linejoin="round"><rect x="0.8" y="0.8" width="14.4" height="7" rx="1.2"/><rect x="2.5" y="2.5" width="3.5" height="3.5" rx="0.5"/><line x1="8" y1="3" x2="13" y2="3"/><line x1="8" y1="5" x2="11" y2="5"/><rect x="3" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/><rect x="6.5" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/><rect x="10" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/></svg></span>
GPU <span id="vv-gpu-label">GPU</span>
</span> </span>
<a href="/" target="_blank" class="vv-card-cog" title="Dashboard">⚙</a> <a href="/" target="_blank" class="vv-card-cog" title="Dashboard">⚙</a>
</h3> </h3>
<div id="vv-gpu-body">Loading...</div> <div id="vv-gpu-body">Loading...</div>
</div> </div>
<div class="vv-card" id="vv-transcode" style="grid-column:span 2;"> <div class="vv-card" id="vv-gpu1-card" style="grid-column:span 1;">
<h3>
<span style="display:flex;align-items:center;gap:5px;">
<span class="vv-ico"><svg width="15" height="10" viewBox="0 0 16 10" fill="none" stroke="#aaa" stroke-width="1.1" stroke-linecap="round" stroke-linejoin="round"><rect x="0.8" y="0.8" width="14.4" height="7" rx="1.2"/><rect x="2.5" y="2.5" width="3.5" height="3.5" rx="0.5"/><line x1="8" y1="3" x2="13" y2="3"/><line x1="8" y1="5" x2="11" y2="5"/><rect x="3" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/><rect x="6.5" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/><rect x="10" y="7.8" width="2" height="1.8" rx="0.3" stroke="none" fill="#666"/></svg></span>
<span id="vv-gpu1-label">GPU 1</span>
</span>
<a href="/" target="_blank" class="vv-card-cog" title="Dashboard">⚙</a>
</h3>
<div id="vv-gpu1-body">Loading...</div>
</div>
<div class="vv-card" id="vv-transcode" style="grid-column:span 1;">
<h3> <h3>
<span style="display:flex;align-items:center;gap:5px;"> <span style="display:flex;align-items:center;gap:5px;">
<span class="vv-ico"><svg width="14" height="12" viewBox="0 0 14 12" fill="none" stroke="#aaa" stroke-width="1.1" stroke-linecap="round" stroke-linejoin="round"><rect x="0.5" y="1" width="4" height="10" rx="0.8"/><line x1="0.5" y1="3.5" x2="2.3" y2="3.5"/><line x1="0.5" y1="8.5" x2="2.3" y2="8.5"/><line x1="2.7" y1="3.5" x2="4.5" y2="3.5"/><line x1="2.7" y1="8.5" x2="4.5" y2="8.5"/><line x1="6" y1="6" x2="9.5" y2="6"/><polyline points="8.2,4.3 10.2,6 8.2,7.7"/><rect x="11" y="1" width="2.5" height="10" rx="0.5"/></svg></span> <span class="vv-ico"><svg width="14" height="12" viewBox="0 0 14 12" fill="none" stroke="#aaa" stroke-width="1.1" stroke-linecap="round" stroke-linejoin="round"><rect x="0.5" y="1" width="4" height="10" rx="0.8"/><line x1="0.5" y1="3.5" x2="2.3" y2="3.5"/><line x1="0.5" y1="8.5" x2="2.3" y2="8.5"/><line x1="2.7" y1="3.5" x2="4.5" y2="3.5"/><line x1="2.7" y1="8.5" x2="4.5" y2="8.5"/><line x1="6" y1="6" x2="9.5" y2="6"/><polyline points="8.2,4.3 10.2,6 8.2,7.7"/><rect x="11" y="1" width="2.5" height="10" rx="0.5"/></svg></span>
@@ -1494,9 +1505,30 @@ function vvPollMonitor() {
})(); })();
// ── GPU ───────────────────────────────────────────────────────────────── // ── GPU ─────────────────────────────────────────────────────────────────
const gpu = d.gpu ?? {}; // One card per installed GPU. Falls back to the legacy single-GPU payload so the
const gpuProcs = d.gpu_procs ?? []; // page still renders against a cached response written before 'gpus' existed.
if (gpu.available) { const gpuList = (d.gpus && d.gpus.length) ? d.gpus
: ((d.gpu && d.gpu.available) ? [d.gpu] : []);
const allProcs = d.gpu_procs ?? [];
const renderGpuCard = (gpu, bodyId, labelId, fallbackLabel) => {
const bodyEl = document.getElementById(bodyId);
const labelEl = document.getElementById(labelId);
if (!bodyEl) return;
if (!gpu || !gpu.available) {
if (labelEl) labelEl.textContent = fallbackLabel;
bodyEl.innerHTML = '<p style="color:#555;font-style:italic">No GPU detected</p>';
return;
}
// Attribute processes to this card by UUID. Older payloads have no gpu_uuid on the
// process rows — in that case only the first card claims them, rather than every
// card showing the same list.
const gpuProcs = allProcs.filter(p =>
p.gpu_uuid ? p.gpu_uuid === gpu.uuid : (gpu.index ?? 0) === 0
);
const vramPct = gpu.memory_total > 0 ? Math.round(gpu.memory_used / gpu.memory_total * 100) : 0; const vramPct = gpu.memory_total > 0 ? Math.round(gpu.memory_used / gpu.memory_total * 100) : 0;
const utilPct = gpu.utilization ?? 0; const utilPct = gpu.utilization ?? 0;
const temp = gpu.temperature ?? 0; const temp = gpu.temperature ?? 0;
@@ -1505,7 +1537,9 @@ function vvPollMonitor() {
const procCount = gpuProcs.length; const procCount = gpuProcs.length;
const procColor = procCount > 0 ? '#4caf50' : '#555'; const procColor = procCount > 0 ? '#4caf50' : '#555';
// Process list — which apps are actually using the GPU (name + VRAM) if (labelEl) labelEl.textContent = 'GPU ' + (gpu.index ?? 0);
// Process list — which apps are actually using this GPU (name + VRAM)
let gpuProcHtml = ''; let gpuProcHtml = '';
if (procCount > 0) { if (procCount > 0) {
const rows = gpuProcs.map(p => { const rows = gpuProcs.map(p => {
@@ -1521,7 +1555,7 @@ function vvPollMonitor() {
</div>`; </div>`;
} }
document.getElementById('vv-gpu-body').innerHTML = bodyEl.innerHTML =
// header row: name + process count pill // header row: name + process count pill
`<div style="display:flex;align-items:center;justify-content:space-between;margin-bottom:8px;"> `<div style="display:flex;align-items:center;justify-content:space-between;margin-bottom:8px;">
<div style="font-size:11px;color:#888;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1;">${gpu.name}</div> <div style="font-size:11px;color:#888;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1;">${gpu.name}</div>
@@ -1538,9 +1572,19 @@ function vvPollMonitor() {
<span style="color:#666;margin-left:12px;">Power</span> <span style="color:#666;margin-left:12px;">Power</span>
<span style="color:#aaa;font-weight:bold;">${powerStr}</span> <span style="color:#aaa;font-weight:bold;">${powerStr}</span>
</div>` + gpuProcHtml; </div>` + gpuProcHtml;
} else { };
document.getElementById('vv-gpu-body').innerHTML = '<p style="color:#555;font-style:italic">No GPU detected</p>';
} renderGpuCard(gpuList[0], 'vv-gpu-body', 'vv-gpu-label', 'GPU');
renderGpuCard(gpuList[1], 'vv-gpu1-body', 'vv-gpu1-label', 'GPU 1');
// Row 3 is 8 columns wide: Rsync(1) + GPU0(1) + GPU1(1) + Transcode(1) + Streams(4).
// On a single-GPU host the second card is hidden and Transcode widens back to 2 so the
// row still fills exactly — otherwise it would leave a one-column hole.
const gpu1Card = document.getElementById('vv-gpu1-card');
const transcodeCard = document.getElementById('vv-transcode');
const twoGpus = gpuList.length > 1;
if (gpu1Card) gpu1Card.style.display = twoGpus ? '' : 'none';
if (transcodeCard) transcodeCard.style.gridColumn = twoGpus ? 'span 1' : 'span 2';
// ── Scripts ───────────────────────────────────────────────────────────── // ── Scripts ─────────────────────────────────────────────────────────────
vvLastScripts = d.scripts ?? {}; vvLastScripts = d.scripts ?? {};
+106 -6
View File
@@ -41,9 +41,59 @@
# using version-stable field names. # using version-stable field names.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Deletion Is Opt-In, Never Inherited
# DEFAULT_RSYNC_OPTS deliberately omits --delete. Media shares only ever gain files
# here; the arr cleanup scripts own deletion and are the only things that understand
# whether a file is genuinely orphaned. --delete appears in exactly two places: a
# profile that sets it explicitly, and the push pass of --merge-run.
#
# Profiles Replace, Not Extend
# PROFILE_RSYNC_OPTS does not inherit DEFAULT_RSYNC_OPTS. A profile states its full
# flag set, so reading one profile tells you exactly what will run — no tracing
# through a base list to discover an inherited --delete.
#
# Local Is Authoritative
# In --merge-run the remote may contribute content the local lacks, but never a
# competing version. Pass 1 pulls only with --ignore-existing, so the push in pass 2
# can safely be authoritative. Reversing that order would let the remote overwrite
# local files before the delete pass.
#
# Pre-flight Before Payload
# Connectivity, version parity, remote rootfs, remote share, remote disks and drive
# temperature are all checked before a byte moves. A transfer aborted halfway is more
# expensive to reason about than one that never started.
#
# Bounded, Resumable Transfers
# No single attempt may exceed RSYNC_MAX_RUNTIME_HOURS. This is safe only because
# --partial is in the default opts: a terminated transfer resumes rather than
# restarting, so bounding it costs nothing and prevents one huge or stuck transfer
# from holding its lock indefinitely and starving every other profile of a turn.
#
# Silent on Success
# Only failures produce visible output. A quiet run is a successful one.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root Enforcement
# rsync over SSH as root and container stop/start both require it.
#
# Docker Presence Check
# Verified before any profile container operations.
#
# Source Path Depth Guard
# The source must be an absolute path at least three levels deep. It is pushed to
# root@remote at the same absolute path and --merge-run adds --delete, so a truncated
# argument is a remote-side hazard: /mnt/user would sync every share at once.
#
# Source Existence Check
# The local directory must exist. Without this, --merge-run's pull pass would create
# a mistyped directory, populate it from the remote, then push it back with --delete.
#
# Global Rsync Gate # Global Rsync Gate
# check_rsync_enabled() — RSYNC_ENABLED=false exits cleanly before any operation. # check_rsync_enabled() — RSYNC_ENABLED=false exits cleanly before any operation.
# #
@@ -72,8 +122,23 @@
# acquire_rsync_lock() — per-profile lock prevents parallel runs of the same profile. # acquire_rsync_lock() — per-profile lock prevents parallel runs of the same profile.
# Global concurrent limit prevents too many simultaneous rsync processes. # Global concurrent limit prevents too many simultaneous rsync processes.
# #
# Notification Validated # Runtime Ceiling
# platform_require_cmd confirms the notify script is present before use. # RSYNC_MAX_RUNTIME_HOURS terminates a single attempt that overruns, releasing the
# per-profile lock for the next scheduled run. --partial in the default opts means the
# paused transfer resumes rather than restarting from scratch.
#
# Merge-Run Ordering
# Pass 1 pulls remote-unique content with --ignore-existing before pass 2 pushes with
# --delete. Content the remote had and the local did not is preserved locally before
# anything is deleted remotely.
#
# Merge-Run Delete Interlock
# --delete is applied only when pass 1 completed. Its entire justification is that the
# local is now the authoritative superset, and a capped or failed pull means it is not:
# the remote still holds content the local never received. On an incomplete pull the
# push proceeds without --delete and notifies, so local content still propagates while
# nothing remote-unique is destroyed. The delete happens on a later run whose pull
# succeeded.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -176,6 +241,25 @@ parse_args "${RAW_ARGS[@]}"
exit 1 exit 1
} }
# ── Source path guards ────────────────────────────────────────────────────────────────────────
# This script pushes to root@remote at the same absolute path, and --merge-run adds --delete on
# the push pass. A truncated or mistyped source is therefore a remote-side data hazard, not just
# a local no-op: /mnt/user would sync every share at once, and / would target the filesystem
# root. Every real job path (see HOST*_*_SYNC_SHARES) is at least three levels deep.
_rsync_depth="${DIRECTORY//[^\/]/}"
if [[ "$DIRECTORY" != /* || "${#_rsync_depth}" -lt 3 ]]; then
error "Refusing unsafe source path: '$DIRECTORY' — expected an absolute path at least 3 levels deep"
exit 1
fi
unset _rsync_depth
# Without this, --merge-run's pull pass would create a mistyped local directory, fill it with
# remote content, then push it back with --delete.
if [[ ! -d "$DIRECTORY" ]]; then
error "Source directory does not exist locally: $DIRECTORY"
exit 1
fi
# ============================================================================================== # ==============================================================================================
# ━━━ Setup ━━━ # ━━━ Setup ━━━
# ============================================================================================== # ==============================================================================================
@@ -380,23 +464,39 @@ if [[ "$MERGE_RUN" == true ]]; then
_pull_exit=$? _pull_exit=$?
_pull_bytes=$(echo "$_pull_output" | \ _pull_bytes=$(echo "$_pull_output" | \
awk '/Total transferred file size:/{gsub(/,/,"",$NF); gsub(/[^0-9]/,"",$NF); print $NF+0}') awk '/Total transferred file size:/{gsub(/,/,"",$NF); gsub(/[^0-9]/,"",$NF); print $NF+0}')
_pull_complete=false
if [[ "$_pull_exit" -eq 0 ]]; then if [[ "$_pull_exit" -eq 0 ]]; then
echo "$ICON_DONE Merge pass 1 complete — ${_pull_bytes:-0} bytes pulled" echo "$ICON_DONE Merge pass 1 complete — ${_pull_bytes:-0} bytes pulled"
_pull_complete=true
elif [[ "$_pull_exit" -eq 124 ]]; then elif [[ "$_pull_exit" -eq 124 ]]; then
# Same cap as the main push — never let one direction of the merge run unbounded. # Same cap as the main push — never let one direction of the merge run unbounded.
# Files already completed before the timeout stay pulled; next scheduled run # Files already completed before the timeout stay pulled; next scheduled run
# picks up whatever's still remote-unique (no --partial here, so an in-flight # picks up whatever's still remote-unique (no --partial here, so an in-flight
# file at the moment of the kill is discarded, not corrupted). # file at the moment of the kill is discarded, not corrupted).
warn "Merge pass 1 exceeded ${RSYNC_MAX_RUNTIME_HOURS}h cap — continuing with push, will resume pull next scheduled run" warn "Merge pass 1 exceeded ${RSYNC_MAX_RUNTIME_HOURS}h cap — pull incomplete"
else else
warn "Merge pass 1 failed (exit $_pull_exit) — continuing with push" warn "Merge pass 1 failed (exit $_pull_exit)"
echo "$_pull_output" | grep -iE "error|rsync:|permission denied" | while read -r _line; do echo "$_pull_output" | grep -iE "error|rsync:|permission denied" | while read -r _line; do
warn " $_line" warn " $_line"
done done
fi fi
unset _pull_opts _pull_output _pull_exit _pull_bytes
# Pass 2: push with --delete — local is now the authoritative superset # Pass 2: push with --delete — but ONLY if pass 1 actually finished.
# --delete is justified solely by "local is now the authoritative superset", and that
# premise holds only on a complete pull. If pass 1 was capped or failed, the remote still
# holds content the local never received; deleting it here would destroy it permanently
# and make the "resume the pull next run" promise impossible to keep. Push without
# --delete instead — local content still propagates, nothing remote-unique is lost, and
# the delete happens on a later run whose pull completed.
if [[ "$_pull_complete" == true ]]; then
RSYNC_OPTS+=(--delete) RSYNC_OPTS+=(--delete)
else
warn "Skipping --delete this run — pass 1 did not complete, remote may hold content the local lacks"
warn "Local content still pushes; the delete pass runs once a full pull succeeds"
notify "Merge run on $(hostname) pushed without --delete — pass 1 incomplete for $PROFILE_NAME" \
"Rsync" "warning"
fi
unset _pull_opts _pull_output _pull_exit _pull_bytes _pull_complete
fi fi
START=$(date +%s) START=$(date +%s)
+24
View File
@@ -16,9 +16,33 @@
# SABnzbd) grow fastest — inactive containers typically remain small. # SABnzbd) grow fastest — inactive containers typically remain small.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Two independent passes, each with its own threshold:
#
# System logs — every path in LOG_FILES
# size < LOG_MIN_SIZE_MB → skip, recent diagnostic history is worth keeping
# size >= LOG_MIN_SIZE_MB → truncate in place
#
# Docker logs — /var/lib/docker/containers/**/*-json.log
# container name resolved for reporting via docker inspect
# size < LOG_DOCKER_MAX_MB → skip
# size >= LOG_DOCKER_MAX_MB → truncate in place
# containers directory missing → whole pass skipped, not an error
#
# Truncation is always `: > file`, never rm — see Truncate, Never Delete below.
# Freed bytes are totalled per pass and reported in the summary.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
# Truncate, Never Delete
# Logs are emptied in place, never removed. The writing process keeps its open file
# handle and keeps logging; deleting the inode would leave a running daemon writing
# to a file nothing can read, and would consume more tmpfs, not less.
#
# Size Thresholds, Not Blind Truncation # Size Thresholds, Not Blind Truncation
# A 2MB syslog contains useful recent diagnostic history — not worth clearing. # A 2MB syslog contains useful recent diagnostic history — not worth clearing.
# A 500MB Docker log is consuming rootfs and contains mostly noise — clear it. # A 500MB Docker log is consuming rootfs and contains mostly noise — clear it.
+107 -6
View File
@@ -22,6 +22,21 @@
# Own conf is never in the backup — it's always on disk. # Own conf is never in the backup — it's always on disk.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Gates — PARTNERSHIP_ENABLED, and a validated PERSISTENT_CONF_CACHE path
# 2. No backup directory → exit 0, nothing to restore
# 3. For each host*.conf in the backup:
# own conf → skip (always on disk)
# already in RAM cache → skip — conf_sync.sh reached the partner, its copy
# is fresher than this one
# otherwise → copy into the RAM cache, mode 600
# 4. Clear the backup unconditionally — see Remove After Use below
#
# Counterpart to conf_cache_save.sh, which writes this backup at array stop.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -40,9 +55,73 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# require_partnership — exits early if PARTNERSHIP_ENABLED=false # Root Enforcement
# detect_hosts() — determines which conf files belong to partners vs self # Reads the plugin-directory backup and writes the RAM cache.
# No-backup guard — exits cleanly if PERSISTENT_CONF_CACHE doesn't exist #
# Lock Acquisition
# acquire_lock prevents this racing conf_cache_save.sh or the conf cache watchdog
# over the same backup directory — this script deletes it at the end.
#
# Partnership Gate
# require_partnership exits early if PARTNERSHIP_ENABLED=false.
#
# Host Detection
# detect_hosts() determines which conf files are partner confs and which is our own.
#
# Cache Path Sanity Guard
# PERSISTENT_CONF_CACHE is validated as an absolute path at least three levels deep
# before anything is read or removed. This script ends with rm -rf on that path, and
# the directory-exists check alone would not catch a collapsed value — / is a
# directory.
#
# No-Backup Guard
# Exits cleanly if the backup directory does not exist — the normal case when the
# partner was reachable at boot.
#
# Fresh-Copy Precedence
# A conf already present in the RAM cache is never overwritten from the backup.
# conf_sync.sh reaching the partner means its copy is current; the backup is by
# definition older.
#
# Own-Conf Exclusion
# Our own conf is never restored from the backup over the live on-disk copy.
#
# Credential File Permissions
# The RAM cache directory is created 700 and each restored conf written 600 — these
# carry partner NPM/lldap passwords and API keys and live under a world-readable /tmp.
#
# Dry Run Support
# --dry-run reports what would be restored and removed, and changes nothing.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# PERSISTENT_CONF_CACHE
# Reboot-surviving backup written by conf_cache_save.sh. Consumed and cleared here.
#
# CONF_RAM_CACHE_DIR
# Destination RAM cache (tmpfs, /tmp/.cache/vv/d) that load_config.sh reads
# partner vars from.
#
# PARTNERSHIP_ENABLED
# Checked via require_partnership().
#
# ==============================================================================================
# RUNTIME MODES
# ==============================================================================================
#
# conf_cache_restore.sh
# Restore partner confs into the RAM cache, then clear the backup.
# Runs at array start, after conf_sync.sh has had its chance.
#
# conf_cache_restore.sh --dry-run
# Report what would be restored and removed without changing anything
#
# conf_cache_restore.sh --log
# Verbose per-file output
# #
# ============================================================================================== # ==============================================================================================
@@ -50,11 +129,31 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../load_config.sh" source "$SCRIPT_DIR/../load_config.sh"
parse_args "$@" parse_args "$@"
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock
detect_hosts detect_hosts
require_partnership require_partnership
RAM_CACHE="$CONF_RAM_CACHE_DIR" RAM_CACHE="$CONF_RAM_CACHE_DIR"
SAVE_DIR="$PERSISTENT_CONF_CACHE" SAVE_DIR="${PERSISTENT_CONF_CACHE:-}"
# SAVE_DIR is rm -rf'd at the end of this script and is built from ${SCRIPTS_DIR}. If that is
# ever unset the path collapses toward / — and the -d check below would pass, since / is a
# directory. Require an absolute path at least three levels deep before touching it.
_slashes="${SAVE_DIR//[^\/]/}"
if [[ -z "$SAVE_DIR" || "$SAVE_DIR" != /* || "${#_slashes}" -lt 3 ]]; then
error "PERSISTENT_CONF_CACHE is unset or unsafe ('${SAVE_DIR:-unset}') — refusing to restore or clear"
notify "conf_cache_restore aborted on $(hostname) — PERSISTENT_CONF_CACHE is '${SAVE_DIR:-unset}'" \
"Conf Cache Restore" "warning"
exit 1
fi
unset _slashes
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
@@ -80,8 +179,10 @@ for conf in "$SAVE_DIR"/host*.conf; do
continue continue
fi fi
mkdir -p "$RAM_CACHE" # Partner confs carry credentials (NPM/lldap passwords, API keys). Default umask would
if cp "$conf" "$RAM_CACHE/$base"; then # leave them 644 in a world-readable /tmp path — restrict on the way in, not afterwards.
mkdir -p "$RAM_CACHE" && chmod 700 "$RAM_CACHE"
if cp "$conf" "$RAM_CACHE/$base" && chmod 600 "$RAM_CACHE/$base"; then
echo "Restored $base from persistent backup → RAM cache ✅" echo "Restored $base from persistent backup → RAM cache ✅"
(( restored++ )) (( restored++ ))
else else
+101 -6
View File
@@ -18,6 +18,19 @@
# Path adapts to storage mode: $SCRIPTS_DIR/.cache/vv/d (internal or appdata). # Path adapts to storage mode: $SCRIPTS_DIR/.cache/vv/d (internal or appdata).
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Gates — PARTNERSHIP_ENABLED, and a validated PERSISTENT_CONF_CACHE path
# 2. No RAM cache present → exit 0, nothing to snapshot
# 3. For each host*.conf in the RAM cache:
# own conf → skip (always on disk, never needs saving)
# partner conf → copy to $PERSISTENT_CONF_CACHE, mode 600
#
# Counterpart to conf_cache_restore.sh, which consumes and then clears this backup
# at the next array start.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -34,9 +47,70 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# require_partnership — exits early if PARTNERSHIP_ENABLED=false # Root Enforcement
# detect_hosts() — determines which confs to save (partner confs only) # Writes into $PERSISTENT_CONF_CACHE under the plugin directory.
# No-cache guard — exits cleanly if RAM cache is empty or missing #
# Lock Acquisition
# acquire_lock prevents this racing conf_cache_restore.sh or the conf cache
# watchdog over the same backup directory.
#
# Partnership Gate
# require_partnership exits early if PARTNERSHIP_ENABLED=false.
#
# Host Detection
# detect_hosts() determines which confs are partner confs and which is our own.
#
# Cache Path Sanity Guard
# PERSISTENT_CONF_CACHE is validated as an absolute path at least three levels
# deep before anything is written. It is built from ${SCRIPTS_DIR}; if that were
# unset the copy target would collapse to "/host2.conf", dropping partner
# passwords and API keys at the filesystem root.
#
# No-Cache Guard
# Exits cleanly if the RAM cache is missing or empty — nothing to snapshot is a
# normal state, not an error.
#
# Own-Conf Exclusion
# Our own conf is never written into the partner backup. Restoring it later
# would overwrite live local config with a stale copy.
#
# Credential File Permissions
# The backup directory is created 700 and each conf written 600. These files
# carry partner NPM/lldap passwords and API keys and must not inherit the
# default umask on a path that survives reboot.
#
# Dry Run Support
# --dry-run reports every file it would write and writes none.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# PERSISTENT_CONF_CACHE
# Reboot-surviving destination for the partner conf backup. Built from
# ${SCRIPTS_DIR}, so it follows the active storage mode.
#
# CONF_RAM_CACHE_DIR
# Source RAM cache (tmpfs, /tmp/.cache/vv/d) populated by conf_sync.sh.
#
# PARTNERSHIP_ENABLED
# Checked via require_partnership().
#
# ==============================================================================================
# RUNTIME MODES
# ==============================================================================================
#
# conf_cache_save.sh
# Snapshot partner confs from the RAM cache to the persistent backup.
# Runs first in ARRAY_STOP_SCRIPTS.
#
# conf_cache_save.sh --dry-run
# Report what would be saved without writing anything
#
# conf_cache_save.sh --log
# Verbose per-file output
# #
# ============================================================================================== # ==============================================================================================
@@ -44,11 +118,30 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../load_config.sh" source "$SCRIPT_DIR/../load_config.sh"
parse_args "$@" parse_args "$@"
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock
detect_hosts detect_hosts
require_partnership require_partnership
RAM_CACHE="$CONF_RAM_CACHE_DIR" RAM_CACHE="$CONF_RAM_CACHE_DIR"
SAVE_DIR="$PERSISTENT_CONF_CACHE" SAVE_DIR="${PERSISTENT_CONF_CACHE:-}"
# Credentials get written here. An empty SAVE_DIR would make the cp target "/host2.conf",
# dropping partner passwords and API keys at the filesystem root.
_slashes="${SAVE_DIR//[^\/]/}"
if [[ -z "$SAVE_DIR" || "$SAVE_DIR" != /* || "${#_slashes}" -lt 3 ]]; then
error "PERSISTENT_CONF_CACHE is unset or unsafe ('${SAVE_DIR:-unset}') — refusing to save partner confs"
notify "conf_cache_save aborted on $(hostname) — PERSISTENT_CONF_CACHE is '${SAVE_DIR:-unset}'" \
"Conf Cache Save" "warning"
exit 1
fi
unset _slashes
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be written" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be written"
@@ -69,8 +162,10 @@ for conf in "$RAM_CACHE"/host*.conf; do
continue continue
fi fi
mkdir -p "$SAVE_DIR" # Partner confs carry credentials — restrict on write rather than leaving them at the
if cp "$conf" "$SAVE_DIR/$base"; then # default umask on a path that survives reboot.
mkdir -p "$SAVE_DIR" && chmod 700 "$SAVE_DIR"
if cp "$conf" "$SAVE_DIR/$base" && chmod 600 "$SAVE_DIR/$base"; then
echo "Saved $base$SAVE_DIR" echo "Saved $base$SAVE_DIR"
(( saved++ )) (( saved++ ))
else else
+98 -8
View File
@@ -25,6 +25,21 @@
# works whether the remote is in internal or appdata storage mode. # works whether the remote is in internal or appdata storage mode.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Gates — PARTNERSHIP_ENABLED, CONF_SYNC_ENABLED
# 2. Cache own conf into the local RAM cache (skipped in --push-only / --pull-only)
# 3. Per partner:
# a. Resolve the partner's own SCRIPTS_DIR by reading their varaverk.cfg over SSH,
# so a partner in appdata storage mode is still found
# b. Pull — scp their host*.conf from their disk into our RAM cache
# c. Push — scp our host*.conf into their RAM cache
# A partner that fails SSH is counted and skipped; the others still sync.
#
# Every file written locally or remotely is restricted to 600, in a 700 directory.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -42,10 +57,65 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# require_partnership — exits early if PARTNERSHIP_ENABLED=false # Root Enforcement
# detect_hosts() — partner list for push/pull routing # Reads the on-disk conf and writes the RAM cache; SSH/scp run as root.
# acquire_lock — prevents concurrent sync runs #
# SSH reachability — partners that fail SSH are skipped, not fatal # Lock Acquisition
# acquire_lock prevents concurrent sync runs writing the same cache files.
#
# Partnership Gate
# require_partnership exits early if PARTNERSHIP_ENABLED=false.
#
# CONF_SYNC_ENABLED Gate
# Exits cleanly when disabled, without removing it from the schedule.
#
# Host Detection
# detect_hosts() builds the partner list used for push/pull routing.
#
# SSH Reachability
# Partners that fail SSH are counted and skipped, never fatal — one unreachable
# partner does not prevent the others from syncing.
#
# SSH Timeouts
# Every ssh and scp call is wrapped in timeout with ConnectTimeout and BatchMode,
# so an unresponsive or password-prompting partner cannot stall the run.
#
# Remote Path Discovery
# The partner's SCRIPTS_DIR is read from their own varaverk.cfg rather than assumed,
# so a partner in appdata storage mode is still found. Falls back to the default
# plugin path if the file cannot be read.
#
# Credential File Permissions
# Cache directories are created 700 and every conf written 600 — on both ends. These
# files carry NPM/lldap passwords and API keys, and the cache lives under a
# world-readable /tmp path. The pushed copy is chmod'd on the partner too, since our
# own credentials land on their disk.
#
# Dry Run Support
# --dry-run reports every pull and push without transferring anything.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# CONF_SYNC_ENABLED
# Master toggle for conf syncing (default: true)
#
# CONF_RAM_CACHE_DIR
# tmpfs cache both ends read partner vars from (/tmp/.cache/vv/d). Cleared every
# reboot, which is why conf_cache_save.sh / conf_cache_restore.sh exist.
#
# SSH_KEY
# Key used for all partner ssh/scp operations
#
# PARTNERSHIP_ENABLED
# Checked via require_partnership()
#
# host*.conf
#
# HOST* — hostnames used to build the partner list via detect_hosts()
# #
# ============================================================================================== # ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
@@ -73,6 +143,14 @@ for arg in "$@"; do
esac esac
done done
parse_args "${FILTERED_ARGS[@]}" parse_args "${FILTERED_ARGS[@]}"
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock
detect_hosts detect_hosts
require_partnership require_partnership
@@ -105,7 +183,10 @@ _remote_scripts_dir() {
# ── Ensure cache dir exists ─────────────────────────────────────────────────── # ── Ensure cache dir exists ───────────────────────────────────────────────────
if [[ "$DRY_RUN" == false ]]; then if [[ "$DRY_RUN" == false ]]; then
mkdir -p "$CACHE_DIR" # These conf files carry credentials (NPM/lldap passwords, API keys). The cache lives in
# a world-readable /tmp path, so the directory and every file written into it below are
# restricted explicitly rather than left at the default umask.
mkdir -p "$CACHE_DIR" && chmod 700 "$CACHE_DIR"
fi fi
# ── Copy own conf into local cache ─────────────────────────────────────────── # ── Copy own conf into local cache ───────────────────────────────────────────
@@ -114,8 +195,12 @@ if [[ "$PUSH_ONLY" == false ]] && [[ "$PULL_ONLY" == false ]]; then
if [[ "$DRY_RUN" == true ]]; then if [[ "$DRY_RUN" == true ]]; then
warn "DRY RUN — would copy $(basename "$MY_CONF")$CACHE_DIR/" warn "DRY RUN — would copy $(basename "$MY_CONF")$CACHE_DIR/"
else else
cp "$MY_CONF" "$CACHE_DIR/${MY_ID,,}.conf" && \ if cp "$MY_CONF" "$CACHE_DIR/${MY_ID,,}.conf" && \
echo "Own conf cached ✅" || warn "Failed to cache own conf" chmod 600 "$CACHE_DIR/${MY_ID,,}.conf"; then
echo "Own conf cached ✅"
else
warn "Failed to cache own conf"
fi
fi fi
else else
warn "Own conf not found: $MY_CONF" warn "Own conf not found: $MY_CONF"
@@ -151,6 +236,7 @@ for host_var in $(compgen -v | grep -E '^HOST[0-9]+$' | sort); do
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \ -o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
"root@${partner_ip}:${remote_conf}" \ "root@${partner_ip}:${remote_conf}" \
"$CACHE_DIR/${partner_slot}.conf" 2>/dev/null; then "$CACHE_DIR/${partner_slot}.conf" 2>/dev/null; then
chmod 600 "$CACHE_DIR/${partner_slot}.conf" 2>/dev/null
echo "Pulled ${partner_slot}.conf from $partner_host" echo "Pulled ${partner_slot}.conf from $partner_host"
(( PULLED++ )) (( PULLED++ ))
else else
@@ -172,12 +258,16 @@ for host_var in $(compgen -v | grep -E '^HOST[0-9]+$' | sort); do
# Ensure partner's cache dir exists, then SCP own conf into it # Ensure partner's cache dir exists, then SCP own conf into it
timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \ timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \ -o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
"root@${partner_ip}" "mkdir -p '$CACHE_DIR'" 2>/dev/null "root@${partner_ip}" "mkdir -p '$CACHE_DIR' && chmod 700 '$CACHE_DIR'" 2>/dev/null
if timeout "$SSH_TIMEOUT" scp -i "$SSH_KEY" \ if timeout "$SSH_TIMEOUT" scp -i "$SSH_KEY" \
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \ -o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
"$MY_CONF" \ "$MY_CONF" \
"root@${partner_ip}:${CACHE_DIR}/${MY_ID,,}.conf" 2>/dev/null; then "root@${partner_ip}:${CACHE_DIR}/${MY_ID,,}.conf" 2>/dev/null; then
# Our own conf lands on the partner carrying our credentials — restrict it there too.
timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
"root@${partner_ip}" "chmod 600 '${CACHE_DIR}/${MY_ID,,}.conf'" 2>/dev/null
echo "Pushed ${MY_ID,,}.conf to $partner_host" echo "Pushed ${MY_ID,,}.conf to $partner_host"
(( PUSHED++ )) (( PUSHED++ ))
else else
+13
View File
@@ -16,6 +16,19 @@
# completely masking real events. # completely masking real events.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Build the expected rsyslog filter content
# 2. Compare against the file already on disk
# identical → exit silently, no write, no rsyslog restart
# missing or different → write the filter, then restart rsyslog
# 3. Verify rsyslog came back up after the restart
#
# Runs before any container starts, so the first wave of veth messages at array
# start is already being filtered.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
+20
View File
@@ -14,6 +14,26 @@
# starts requires a docker restart to pick up the new values. # starts requires a docker restart to pick up the new values.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# For each of the three limits (instances, watches, queued events):
#
# 1. Read the current kernel value via sysctl -n
# 2. Compare against the configured target
# exactly equal → skip, nothing to do
# anything else → apply via sysctl -w
#
# Note this enforces the configured value exactly, in both directions: a limit currently
# set HIGHER than the target is lowered back to it. That is deliberate — the conf is the
# single declared source of truth for these limits — but it means raising a limit by hand
# will be silently undone at the next array start. Raise the target in master.conf instead.
#
# Applied at every array start because these are runtime kernel settings that do not
# survive a reboot, and must land before containers launch — a container inherits the
# limits in force at its start, not dynamically.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
+54
View File
@@ -40,6 +40,35 @@
# If remote unreachable → skips remote cleanly, logs warning. # If remote unreachable → skips remote cleanly, logs warning.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Stop the Transfer, Not the Schedule
# The default mode kills only the rsync subprocess and lets the orchestrator notice
# the exit and wind down on its own. Killing the orchestrator too would abandon the
# remaining shares silently; letting it finish its own loop keeps the schedule honest
# about what ran and what did not.
#
# Exact-Name Process Matching
# Targets are found with pgrep -x rsync — exact process name, never a pattern match
# against a command line. A loose pattern on a box running arbitrary containers could
# match something that merely mentions rsync in its arguments.
#
# Liveness Checked Before Every Signal
# kill -0 confirms a PID is still alive immediately before signalling it. PIDs are
# reused, and a transfer that exited on its own between discovery and signalling must
# not have its number sent a kill.
#
# Interrupting Is Safe by Construction
# rsync runs with --partial, so a killed transfer resumes rather than restarting.
# That is what makes stopping mid-sync a routine operation rather than a costly one.
#
# Clean Up What the Interruption Left
# A killed rsync leaves its lock file behind and may leave profile containers stopped.
# Both are cleared afterwards, so the next scheduled run is not blocked by a lock
# whose owner no longer exists.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -61,6 +90,31 @@
# Remote containers deferred to docker_watchdog.sh. # Remote containers deferred to docker_watchdog.sh.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# LOCK_DIR
# Directory holding rsync and orchestrator lock files. Scanned after a kill to
# clear locks whose owning PID is gone.
#
# DOCKER_TIMEOUT
# Timeout applied to the docker calls used when recovering containers a killed
# rsync left stopped.
#
# PROFILE_CRITICAL_CONTAINER_NAMES
# Per-profile container lists — used to work out which containers an interrupted
# profile sync had stopped and therefore needs restarting.
#
# SSH_KEY / SSH timeouts
# Used to reach the partner when stopping its rsync as well.
#
# host*.conf
#
# HOST* — resolved via detect_hosts() for MY_ID and remote routing
#
# ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
# ============================================================================================== # ==============================================================================================
# #
+25
View File
@@ -37,6 +37,31 @@
# After REBOOT_VM_WAIT seconds, libvirt stops anyway — reboot takes priority. # After REBOOT_VM_WAIT seconds, libvirt stops anyway — reboot takes priority.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Warn, Do Not Block
# Pre-flight reports what is in flight — rsync running, mover running — and proceeds
# anyway. This is an operator-invoked tool: the person running it has already decided
# to reboot, and refusing would just push them to /sbin/reboot with no warning, no
# wall message and no clean array stop. The warnings name the specific script to run
# first (rsync_stop.sh, mover_stop.sh) so the safer path is the obvious one.
#
# Announce Before Acting
# A wall message and a notification go out REBOOT_SLEEP seconds ahead, both naming
# which host is rebooting and why. On a two-server setup "the server is rebooting" is
# ambiguous and therefore useless.
#
# Clean Array Stop First
# The reboot routes through the normal array stop sequence rather than calling
# /sbin/reboot directly. Array stop failures are reported and the reboot continues —
# an already-committed reboot should not be abandoned halfway, leaving services down
# and the machine still up.
#
# Flush Before Cutting Power
# sync runs immediately before /sbin/reboot so buffered writes reach disk.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+50
View File
@@ -18,6 +18,41 @@
# do not need to be hardcoded and work across hosts. # do not need to be hardcoded and work across hosts.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Sonarr then Radarr, each independently:
#
# 1. Fetch quality profiles and resolve the configured names to live IDs
# 2. Fetch the library
# 3. Per item, derive the expected profile from its rootFolderPath
# contains "kids" or "anime" → kids profile
# anything else → default profile
# 4. Compare against the item's current profile
# already correct → skip, no call made
# wrong → queue for the bulk editor PUT
# 5. Issue the change and report counts
#
# ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Resolve Profile IDs at Runtime
# Profiles are looked up by name from the API on every run rather than hardcoding IDs.
# Profile IDs differ between hosts and change when profiles are recreated; a stale
# hardcoded ID would silently assign the wrong profile rather than failing.
#
# Change Only What Is Wrong
# Items already on the correct profile are never touched. That makes the tool safe to
# re-run at any time and keeps the reported count meaningful — a non-zero result means
# real drift, not just "it ran".
#
# Root Folder Is the Source of Truth
# Classification comes from where the item actually lives, not from its metadata. The
# classification scans own the harder question of whether an item is in the right root;
# this tool just makes the profile agree with the answer already on disk.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -69,6 +104,21 @@ done
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../load_config.sh" source "$SCRIPT_DIR/../load_config.sh"
parse_args "$@" parse_args "$@"
# ==============================================================================================
# ━━━ Setup ━━━
# ==============================================================================================
# Consistent with the rest of Arrs_Stack/ and Tools/ — this issues bulk profile PUTs against
# the arrs and should not be runnable by an unprivileged account that happens to hold a key.
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
# Bulk PUTs against the same items — two concurrent runs could interleave and leave a subset
# of series/movies on the wrong profile.
acquire_lock
detect_hosts detect_hosts
KIDS_PROFILE_NAME="${ARR_KIDS_PROFILE_NAME:-Kids shows}" KIDS_PROFILE_NAME="${ARR_KIDS_PROFILE_NAME:-Kids shows}"
+56
View File
@@ -19,6 +19,58 @@
# this for scans it triggers itself; this tool covers everything else. # this for scans it triggers itself; this tool covers everything else.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Validate the arr type and resolve its URL / API key / API version
# 2. Poll /command for any active rescan-type command (ARR_RESCAN_COMMANDS)
# none active → nothing to watch, exit cleanly
# 3. Wait for it to leave the active state, bounded by the timeout argument
# (default 7200s — a full-library scan is slow by nature)
# 4. Once finished, fetch the library and arr_cache_write() the real numbers
#
# Triggers nothing. It only watches a scan someone or something else started.
#
# ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Watch, Never Trigger
# This tool deliberately starts nothing. Its whole job is to be the thing that notices a
# scan finished, so a rescan begun by hand, by the arr's own UI, or by another script
# still gets its result written through to the shared cache.
#
# Complement to the Write Guard
# arr_cache_write() refuses to write while a rescan is active, because a mid-scan snapshot
# reads as real data loss to consumers like check_tracked_count_floor. That guard protects
# the cache but leaves nobody to write the true number afterwards. arr_full_rescan.sh
# covers the scans it starts itself; this covers every other origin.
#
# Bounded Wait
# The wait is capped by the timeout argument, so a scan that stalls or never reports
# completion cannot leave this running indefinitely holding its lock.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# host*.conf (aliased by detect_hosts())
#
# SONARR_URL / SONARR_API_KEY
# RADARR_URL / RADARR_API_KEY
# LIDARR_URL / LIDARR_API_KEY
# Connection details for whichever arr is passed as the first argument.
#
# common.sh
#
# ARR_RESCAN_COMMANDS
# Command names treated as a rescan for this purpose — shared with
# arr_cache_write()'s own active-scan guard so both agree on what counts.
#
# ARR_API_VERSION
# Per-arr API version map (v3 for Sonarr/Radarr, v1 for Lidarr).
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -59,6 +111,10 @@ if [[ "$EUID" -ne 0 ]]; then
exit 1 exit 1
fi fi
# This writes the shared tracked-data cache on completion. Two monitors watching the same
# arr would both write it, and the loser's stale snapshot could land last.
acquire_lock
detect_hosts detect_hosts
ARR_TYPE="${PARSED_ARGS[0]:-}" ARR_TYPE="${PARSED_ARGS[0]:-}"
+35 -3
View File
@@ -23,9 +23,32 @@
# repairs this tool is meant for — worth remembering before pointing it at an entire share. # repairs this tool is meant for — worth remembering before pointing it at an entire share.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# For each path given on the command line:
#
# 1. Path guard — refuse anything shallower than two components
# 2. Existence check — a missing path is a failure for that entry, not the run
# 3. Report scale — file count, dir count, total size, so the operator sees the job size
# 4. Count entries with wrong ownership (the diagnostic number in the summary)
# 5. chown -R PERMISSIONS_OWNER across the path
# 6. find -type d → chmod PERMISSIONS_DIR_MODE
# 7. find -type f → chmod PERMISSIONS_FILE_MODE
#
# Unlike the nightly media_shares_permissions.sh, steps 57 are unconditional — see the
# ctime note in PURPOSE above for why that distinction matters.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
# Unconditional by Design, Unlike the Nightly Job
# media_shares_permissions.sh applies its passes conditionally to protect ctime as an
# age signal for arr orphan collection. This tool deliberately does not: it exists to
# repair paths that are known-wrong, where correctness matters more than preserving a
# clock. That is exactly why it is a targeted manual tool and not a scheduled one.
#
# Permissions Model # Permissions Model
# Directories (PERMISSIONS_DIR_MODE, default 755): # Directories (PERMISSIONS_DIR_MODE, default 755):
# Owner (nobody) — rwx enter, list, create files # Owner (nobody) — rwx enter, list, create files
@@ -53,9 +76,6 @@
# Each path is verified before processing — missing paths log an error and # Each path is verified before processing — missing paths log an error and
# are skipped rather than silently passing. # are skipped rather than silently passing.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# Silent on Success # Silent on Success
# Only failures and the wrong-owner diagnostic produce visible output. # Only failures and the wrong-owner diagnostic produce visible output.
# #
@@ -137,6 +157,18 @@ for share_path in "${PARSED_ARGS[@]}"; do
echo "" echo ""
echo "━━━ $ICON_PERMS $(basename "$share_path") ━━━" echo "━━━ $ICON_PERMS $(basename "$share_path") ━━━"
# chown -R below. Paths come straight from the command line, so a spacing typo
# ("/mnt/user /Movies" instead of "/mnt/user/Movies") would hand this a bare top-level
# directory — and chown -R nobody:users on / or /etc breaks the system outright.
# Require at least two path components; that still allows a deliberate whole-share
# repair like /mnt/user while refusing /, /mnt, /etc, /boot and friends.
_bpr_slashes="${share_path//[^\/]/}"
if [[ "$share_path" != /* || "${#_bpr_slashes}" -lt 2 ]]; then
error "$share_path — refusing: expected an absolute path at least 2 levels deep"
FAIL+=("$share_path")
continue
fi
if [[ ! -d "$share_path" ]]; then if [[ ! -d "$share_path" ]]; then
error "$share_path — not found" error "$share_path — not found"
FAIL+=("$share_path") FAIL+=("$share_path")
+31 -3
View File
@@ -13,6 +13,37 @@
# Output: ContainerName_YYYY-MM-DD_HH-MM.tar.gz — timestamped, no overwrite. # Output: ContainerName_YYYY-MM-DD_HH-MM.tar.gz — timestamped, no overwrite.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Validate arguments — container name, appdata path, output directory all required
# 2. Verify the output directory exists and holds enough free space for the archive
# 3. Record whether the container is currently running
# 4. Stop the container if it was running
# 5. tar czf the appdata directory to a timestamped archive
# 6. Verify the archive with tar --test-file
# 7. Restart the container only if it was running before — a container found stopped
# stays stopped
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# This tool takes everything as arguments rather than from conf:
#
# <container> Container to stop for the duration of the export
# <appdata_path> Directory to archive
# <output_dir> Destination for the archive — must already exist
#
# That is deliberate. It is used for one-off exports of arbitrary containers, including
# ones being removed from the stack entirely, so there is no meaningful configured list
# to draw from and nothing host-specific to alias.
#
# Note the archive is written as root and is not chowned afterwards. That is fine for the
# operator-invoked use this tool is for, but worth knowing if the output directory is a
# user share reached over SMB.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -46,9 +77,6 @@
# DOCKER_TIMEOUT (default: 30s) caps all docker calls. Guards against a hung # DOCKER_TIMEOUT (default: 30s) caps all docker calls. Guards against a hung
# daemon blocking the script indefinitely. # daemon blocking the script indefinitely.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
# ============================================================================================== # ==============================================================================================
+51 -4
View File
@@ -30,13 +30,60 @@
# that may be intentionally paused. # that may be intentionally paused.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Default (dangling only):
# docker image prune -f — removes untagged, unreferenced images only. Every tagged
# image survives, running containers are untouched, and nothing stopped is removed.
#
# --all (full orphan cleanup):
# 1. Remove containers in exited/created state
# 2. Remove every image not used by a RUNNING container
# Step 1 is what makes step 2 reach further: with the stopped containers gone, their
# images are no longer referenced and become eligible. That is also precisely why
# --all is destructive to anything deliberately kept stopped.
#
# Reclaimed space is reported for both modes.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# This tool is deliberately unconfigured — it reads live Docker state rather than any
# configured list, so there is nothing host-specific to alias and no thresholds to tune.
# Scope is controlled entirely by the mode flag (default vs --all).
#
# Note it does not call detect_hosts(): nothing here is host-specific, and it acts only on
# the local daemon.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — Docker prune operations require root # Root Enforcement
# acquire_lock prevents concurrent prune runs # Docker prune operations require root.
# --dry-run — shows what would be removed without taking any action #
# --status — lists current dangling images and stopped containers; no changes # Lock Acquisition
# acquire_lock prevents concurrent prune runs racing on the same image store.
#
# Dangling-Only Default
# The default mode removes untagged, unreferenced images only. Reaching anything
# tagged, running, or deliberately stopped requires --all explicitly — the safe
# behaviour is what you get by not thinking about it.
#
# Running Containers Never Touched
# Neither mode removes a running container or an image a running container uses.
# --all widens the blast radius to STOPPED containers and their images, never to
# anything currently up.
#
# Dry Run Support
# --dry-run lists everything that would be removed and removes nothing. Worth using
# before --all specifically, since that mode deletes intentionally-stopped containers.
#
# Status Mode
# --status lists current dangling images and stopped containers without changing
# anything, so the scope of a prospective --all is visible up front.
# #
# ============================================================================================== # ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
+22
View File
@@ -29,6 +29,28 @@
# Emby's config path is detected from the Docker mount — no hardcoded paths. # Emby's config path is detected from the Docker mount — no hardcoded paths.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Diagnose, Never Repair
# The script reports which databases fail integrity_check and stops there. SQLite
# "repair" means dumping and rebuilding, which silently discards whatever rows were
# corrupt — an outcome nobody should get without deciding to. The remediation steps
# are printed instead, for a human to run against a backup.
#
# Stop Emby Before Reading
# integrity_check against a live database gives unreliable answers and can itself
# contend with Emby's writes. Emby is stopped for the duration and restarted after,
# so the check runs against a quiescent file.
#
# Always Restart, Even on Failure
# Emby is brought back up regardless of what the check found, and an EXIT trap armed
# before it is stopped restores it even if this script dies partway through. It is
# disarmed only once the normal restart has run. Leaving the media server down because
# a diagnostic reported a problem turns an investigation into an outage — and the
# original running state is honoured, so an Emby that was already stopped stays stopped.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+25 -2
View File
@@ -17,7 +17,7 @@
# fetch as fallback. # fetch as fallback.
# #
# ============================================================================================== # ==============================================================================================
# FLOW # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# 1. Pull play completions from Emby activity log (all time, or --days N) # 1. Pull play completions from Emby activity log (all time, or --days N)
@@ -25,6 +25,29 @@
# 3. For each played artist not in Lidarr → add to Lidarr # 3. For each played artist not in Lidarr → add to Lidarr
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Add-Only, Never Remove
# The tool only adds artists the arr is missing. It never deletes or unmonitors anything,
# so the worst outcome of a bad match is an extra tracked entry — trivially reversible —
# rather than lost tracking on something already curated.
#
# Library Membership Is the Whole Criterion
# No scoring, no thresholds, no metadata quality gates. If it is in Emby, the arr should
# know about it. This is deliberately not a discovery tool; the playback_aware_* scripts
# own that job and its judgement calls.
#
# Provider ID Over Title
# Matching prefers the MusicBrainz ID and falls back to case-insensitive title only when the ID
# is absent. Titles differ across sources by punctuation, year suffixes and articles;
# matching on them alone would re-add things the arr already tracks.
#
# Cache-First Read
# The arr library comes from the shared tracked-data cache, kept warm by
# arr_cache_prefill.sh, with a live fetch as fallback. One read regardless of library size.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -40,7 +63,7 @@
# modification. Safe to run multiple times — the second run finds nothing to add. # modification. Safe to run multiple times — the second run finds nothing to add.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION (host*.conf) # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# HOST*_LIDARR_URL / HOST*_LIDARR_API_KEY — Lidarr connection (aliased by detect_hosts) # HOST*_LIDARR_URL / HOST*_LIDARR_API_KEY — Lidarr connection (aliased by detect_hosts)
+25 -2
View File
@@ -16,7 +16,7 @@
# fetch as fallback. # fetch as fallback.
# #
# ============================================================================================== # ==============================================================================================
# FLOW # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# 1. Fetch all Movie items from Emby (with TMDB provider IDs) # 1. Fetch all Movie items from Emby (with TMDB provider IDs)
@@ -27,6 +27,29 @@
# falling back to case-insensitive title comparison. # falling back to case-insensitive title comparison.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Add-Only, Never Remove
# The tool only adds movies the arr is missing. It never deletes or unmonitors anything,
# so the worst outcome of a bad match is an extra tracked entry — trivially reversible —
# rather than lost tracking on something already curated.
#
# Library Membership Is the Whole Criterion
# No scoring, no thresholds, no metadata quality gates. If it is in Emby, the arr should
# know about it. This is deliberately not a discovery tool; the playback_aware_* scripts
# own that job and its judgement calls.
#
# Provider ID Over Title
# Matching prefers the TMDB ID and falls back to case-insensitive title only when the ID
# is absent. Titles differ across sources by punctuation, year suffixes and articles;
# matching on them alone would re-add things the arr already tracks.
#
# Cache-First Read
# The arr library comes from the shared tracked-data cache, kept warm by
# arr_cache_prefill.sh, with a live fetch as fallback. One read regardless of library size.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -42,7 +65,7 @@
# without modification. Safe to run multiple times — the second run finds nothing to add. # without modification. Safe to run multiple times — the second run finds nothing to add.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION (master.conf / host*.conf) # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# RADARR_EMBY_LIBRARIES — Emby library names to scan (master.conf); empty = all libraries # RADARR_EMBY_LIBRARIES — Emby library names to scan (master.conf); empty = all libraries
+25 -2
View File
@@ -16,7 +16,7 @@
# fetch as fallback. # fetch as fallback.
# #
# ============================================================================================== # ==============================================================================================
# FLOW # OPERATIONAL MODEL
# ============================================================================================== # ==============================================================================================
# #
# 1. Fetch all Series items from Emby (with TVDB provider IDs) # 1. Fetch all Series items from Emby (with TVDB provider IDs)
@@ -27,6 +27,29 @@
# falling back to case-insensitive title comparison. # falling back to case-insensitive title comparison.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Add-Only, Never Remove
# The tool only adds series the arr is missing. It never deletes or unmonitors anything,
# so the worst outcome of a bad match is an extra tracked entry — trivially reversible —
# rather than lost tracking on something already curated.
#
# Library Membership Is the Whole Criterion
# No scoring, no thresholds, no metadata quality gates. If it is in Emby, the arr should
# know about it. This is deliberately not a discovery tool; the playback_aware_* scripts
# own that job and its judgement calls.
#
# Provider ID Over Title
# Matching prefers the TVDB ID and falls back to case-insensitive title only when the ID
# is absent. Titles differ across sources by punctuation, year suffixes and articles;
# matching on them alone would re-add things the arr already tracks.
#
# Cache-First Read
# The arr library comes from the shared tracked-data cache, kept warm by
# arr_cache_prefill.sh, with a live fetch as fallback. One read regardless of library size.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -42,7 +65,7 @@
# without modification. Safe to run multiple times — the second run finds nothing to add. # without modification. Safe to run multiple times — the second run finds nothing to add.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION (master.conf / host*.conf) # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
# SONARR_EMBY_LIBRARIES — Emby library names to scan (master.conf); empty = all libraries # SONARR_EMBY_LIBRARIES — Emby library names to scan (master.conf); empty = all libraries
+43 -3
View File
@@ -22,6 +22,49 @@
# covering the remote server until the next detection cycle. # covering the remote server until the next detection cycle.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# 1. Read and display the current state file so the operator sees what is being discarded
# 2. Write a fresh state file:
# state=NORMAL, fallback_start=0, handback_strikes=0
# tier2_started=false, tier3_started=false, tier4_started=false
# 3. Report the new state
#
# Nothing else is touched. No container is started or stopped, no DDNS is moved, no rsync
# is run. fallback.sh picks the new state up on its next cycle and proceeds from NORMAL.
#
# ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# State File Only, Never Containers
# This tool changes what fallback.sh believes, not what is actually running. That
# separation is the point: reconciling the real stack is an operator judgement call,
# and a tool that tried to do both could act on a belief that was already wrong.
#
# The Operator Asserts Reality
# Resetting to NORMAL is a claim that the stack really is normal — right containers on
# the right host, DDNS pointing the right way. The script cannot verify that, so it
# shows the current state before overwriting it and leaves the check to the human. If
# the assertion is wrong, fallback.sh will act on a false NORMAL.
#
# Full Reset, Not Partial Edit
# Every field is rewritten rather than patching individual keys. A partially-reset file
# — NORMAL state with tier flags still true — is a state fallback.sh has no handling
# for and would be worse than either extreme.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# FALLBACK_STATE_FILE
# Path to the fallback state file this tool rewrites. Shared with fallback.sh —
# both must agree or the reset writes somewhere fallback.sh never reads.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -40,9 +83,6 @@
# Interactive mode prompts for YES before writing. Use --force to bypass in # Interactive mode prompts for YES before writing. Use --force to bypass in
# non-interactive contexts (cron, scripts). # non-interactive contexts (cron, scripts).
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# ============================================================================================== # ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
# ============================================================================================== # ==============================================================================================
+25
View File
@@ -30,6 +30,31 @@
# lost on unmount — warn the user but proceed (this is expected for maintenance). # lost on unmount — warn the user but proceed (this is expected for maintenance).
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Redirect Before Unmounting
# The symlink is pointed at the SSD fallback first, and only then is the tmpfs
# unmounted. Reversing that order would leave in-flight transcodes writing into a
# path that is being pulled out from under them. New sessions land on SSD from the
# moment of the flip; existing ones drain.
#
# Unmount Is the Data Loss
# Everything on a tmpfs disappears when it is unmounted — there is nothing to migrate
# and no way to preserve it. That is acceptable only because the contents are
# regenerable transcode segments, which is exactly why this operation is safe for
# transcodes and would not be for any other kind of ramdisk.
#
# Not Mounted Is Success
# An already-unmounted ramdisk is reported and treated as done, not as an error. The
# goal state is "ramdisk not mounted", and the script is idempotent toward it.
#
# State File Reflects Reality
# The transcode state file is updated to record the SSD target, so transcode_manager.sh
# agrees with what actually happened rather than trying to flip back to a ramdisk that
# no longer exists.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+23
View File
@@ -31,6 +31,29 @@
# NVMe any : 510 min # NVMe any : 510 min
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Read-Only Diagnostic
# A SMART self-test is executed by the drive's own firmware and writes nothing to the
# filesystem. The array stays fully online and in use throughout; the only cost is
# background I/O contention while the drive read-scans itself.
#
# Full Scan, Not the Short Test
# The extended test reads every sector. The short test samples, and a bad sector in
# the unsampled region is exactly the kind of latent fault that only surfaces when
# something tries to read it — often during a rebuild, when redundancy is already gone.
#
# Per-Drive Isolation
# A drive that fails to start a test, does not support one, or reports an error never
# prevents the remaining drives from being tested. Partial coverage beats none.
#
# Report, Do Not Act
# Results are reported and notified; nothing is disabled, replaced or rebuilt on the
# basis of a failed test. Acting on a drive is a hardware decision with cost and
# downtime attached, and belongs to the operator.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+70
View File
@@ -41,6 +41,64 @@
# safely after a partial run or a manually resolved conflict. # safely after a partial run or a manually resolved conflict.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Move, Never Copy or Delete
# Files are relocated with mv and nothing is ever removed. A misjudged move is
# reversible by hand; a delete is not. The script has no cleanup pass by design.
#
# Idempotent by Collision Skip
# Re-running is safe because an existing destination is skipped rather than
# overwritten. A partial run, an interrupted run, or a manually resolved conflict
# can all be followed by a plain re-run.
#
# Deterministic Ordering
# Multiple trailers per show are sorted before numbering, so the same input always
# produces the same trailer.ext / trailer-2.ext assignment. Without the sort, the
# numbering would depend on filesystem iteration order and a re-run could rename
# files differently.
#
# Filesystem Only
# Neither Trailarr nor Emby is touched. Emby's periodic library scan picks the moved
# files up on its own — poking either service would add failure modes to what is
# otherwise a pure file move.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS
# ==============================================================================================
#
# Root Enforcement
# Moves files owned by container users and chowns the created trailers/ directory.
#
# Lock Acquisition
# acquire_lock prevents two runs numbering the same show's trailers concurrently.
#
# Host Detection
# detect_hosts() aliases HOST*_SONARR_TV_ROOT → SONARR_TV_ROOT.
#
# TV Root Validation
# Exits if SONARR_TV_ROOT is unset or is not a directory — the scan below is rooted
# there, and an empty value would walk from the current directory.
#
# Bounded Scan Depth
# find runs with -mindepth 2 -maxdepth 2, so only files sitting directly in a series
# folder are considered. Trailers already correctly placed inside trailers/ are out
# of range and cannot be picked up and re-moved.
#
# Collision Skip
# An existing destination is never overwritten — the source is left in place and
# counted as a conflict for review.
#
# Permission Matching on Created Directories
# mkdir as root would leave trailers/ as root-owned. It is chowned to
# PERMISSIONS_OWNER and chmodded to PERMISSIONS_DIR_MODE so it matches the rest of
# the library and does not become an exception the permissions job has to correct.
#
# Dry Run Support
# --dry-run reports every move and touches nothing.
#
# ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
@@ -50,6 +108,12 @@
# Host filesystem path to the TV library. Aliased by detect_hosts() → # Host filesystem path to the TV library. Aliased by detect_hosts() →
# SONARR_TV_ROOT. # SONARR_TV_ROOT.
# #
# master.conf
#
# PERMISSIONS_OWNER / PERMISSIONS_DIR_MODE
# Applied to each created trailers/ directory so it matches library convention.
# Shared with media_shares_permissions.sh. (defaults: nobody:users, 755)
#
# ============================================================================================== # ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
# ============================================================================================== # ==============================================================================================
@@ -74,6 +138,12 @@ parse_args "$@"
# ============================================================================================== # ==============================================================================================
# ━━━ Setup ━━━ # ━━━ Setup ━━━
# ============================================================================================== # ==============================================================================================
# Moves container-owned media files and chowns the trailers/ dirs it creates.
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock acquire_lock
# detect_hosts() sets MY_ID and aliases HOST*_SONARR_TV_ROOT → SONARR_TV_ROOT # detect_hosts() sets MY_ID and aliases HOST*_SONARR_TV_ROOT → SONARR_TV_ROOT
+39
View File
@@ -28,6 +28,45 @@
# 4. Watchdog monitors normally on next cycle. If it crashes again → re-added. # 4. Watchdog monitors normally on next cycle. If it crashes again → re-added.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Clearing Is an Assertion, Not a Fix
# Removing a container from the skip list tells docker_watchdog.sh to start restarting
# it again. It does nothing about why the container was crash-looping. Clearing before
# the underlying fault is fixed just restarts the loop that put it there.
#
# The Skip List Is Protection, Not Punishment
# A container lands here because restarting it repeatedly was making things worse, not
# better. The entry exists so the watchdog stops burning cycles and notifications on
# something only a human can fix.
#
# Persistent by Design
# The list lives on /boot/config and survives reboots deliberately. A crash loop that
# a reboot would clear is exactly the case where the watchdog would resume looping
# after the reboot — persistence is what stops that.
#
# Auto-Clear Is the Normal Path
# docker_watchdog.sh removes a container from the list on its own once it sees it
# running healthily. This tool is for the case where you have fixed the problem and
# do not want to wait for that, not the routine way entries leave the list.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# DOCKER_WATCHDOG_FAILED_FILE
# Persistent skip list on /boot/config, shared with docker_watchdog.sh. Both must
# agree on the path or the watchdog will not see what this tool changes.
#
# WATCHDOG_CONTAINER_RESTART_LIMIT / WATCHDOG_CONTAINER_RESTART_WINDOW
# The thresholds docker_watchdog.sh uses to decide a container belongs on the list.
# Shown here for context — this tool does not apply them, it only views and edits
# the resulting list.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+57 -2
View File
@@ -43,6 +43,34 @@
# acquire_lock — prevents concurrent registration runs # acquire_lock — prevents concurrent registration runs
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# WEBHOOK_PORT
# Port the listener binds and the registered webhook URL points at. 0 disables the
# listener entirely (start_webhook_listener.sh exits early), so registering against
# a port of 0 would produce URLs nothing is serving.
#
# WEBHOOK_SECRET
# Shared secret embedded in the registered URL as ?key=. Generated here on first run
# and written back into master.conf — the write-back is verified, because the arrs
# are registered with this value and a failed persist would leave them holding a
# secret this host does not have.
#
# host*.conf (aliased by detect_hosts())
#
# SONARR_URL / SONARR_API_KEY
# RADARR_URL / RADARR_API_KEY
# LIDARR_URL / LIDARR_API_KEY
# Each arr the webhook is registered in. An arr with no URL or key configured on
# this host is skipped rather than failing the run.
#
# SSH_KEY
# Used when propagating the same secret to the partner via OVERRIDE_SECRET.
#
# ==============================================================================================
# RUNTIME MODES # RUNTIME MODES
# ============================================================================================== # ==============================================================================================
# #
@@ -76,6 +104,13 @@ MASTER_CONF="$ECOSYSTEM_ROOT/Configurations/master.conf"
source "$ECOSYSTEM_ROOT/load_config.sh" source "$ECOSYSTEM_ROOT/load_config.sh"
parse_args "$@" parse_args "$@"
# Writes the generated secret into master.conf via sed -i and SSHes to the partner.
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock acquire_lock
detect_hosts detect_hosts
@@ -84,9 +119,25 @@ WEBHOOK_NAME="Varaverk Upgrade"
# ── Resolve or generate the secret ────────────────────────────────────────── # ── Resolve or generate the secret ──────────────────────────────────────────
# OVERRIDE_SECRET env var is set when called recursively via SSH from the # OVERRIDE_SECRET env var is set when called recursively via SSH from the
# primary host, so both ends use the same secret. # primary host, so both ends use the same secret.
#
# Every write-back below is verified by re-reading master.conf. The secret gets baked into
# the webhook URL registered in each arr — if the sed silently fails to match, the arrs end
# up holding a secret this host does not have, and the listener rejects every delivery.
_persist_secret() {
local secret="$1"
sed -i "s/WEBHOOK_SECRET=\"\"/WEBHOOK_SECRET=\"$secret\"/" "$MASTER_CONF"
if ! grep -q "WEBHOOK_SECRET=\"$secret\"" "$MASTER_CONF" 2>/dev/null; then
error "Could not persist WEBHOOK_SECRET to $MASTER_CONF"
error "Registering the arrs now would leave them with a secret this host does not have"
notify "Webhook setup aborted on $(hostname) — could not persist WEBHOOK_SECRET" \
"Webhook Setup" "warning"
exit 1
fi
}
if [[ -n "${OVERRIDE_SECRET:-}" ]]; then if [[ -n "${OVERRIDE_SECRET:-}" ]]; then
if [[ -z "${WEBHOOK_SECRET:-}" ]]; then if [[ -z "${WEBHOOK_SECRET:-}" ]]; then
sed -i "s/WEBHOOK_SECRET=\"\"/WEBHOOK_SECRET=\"$OVERRIDE_SECRET\"/" "$MASTER_CONF" _persist_secret "$OVERRIDE_SECRET"
fi fi
WEBHOOK_SECRET="$OVERRIDE_SECRET" WEBHOOK_SECRET="$OVERRIDE_SECRET"
fi fi
@@ -95,8 +146,12 @@ if [[ -z "${WEBHOOK_SECRET:-}" ]]; then
if [[ "$DRY_RUN" == true ]]; then if [[ "$DRY_RUN" == true ]]; then
WEBHOOK_SECRET="<would-generate>" WEBHOOK_SECRET="<would-generate>"
else else
if ! command -v openssl >/dev/null 2>&1; then
error "openssl not found — cannot generate WEBHOOK_SECRET"
exit 1
fi
GENERATED=$(openssl rand -hex 32) GENERATED=$(openssl rand -hex 32)
sed -i "s/WEBHOOK_SECRET=\"\"/WEBHOOK_SECRET=\"$GENERATED\"/" "$MASTER_CONF" _persist_secret "$GENERATED"
WEBHOOK_SECRET="$GENERATED" WEBHOOK_SECRET="$GENERATED"
echo "Generated WEBHOOK_SECRET — saved to master.conf" echo "Generated WEBHOOK_SECRET — saved to master.conf"
fi fi
+27
View File
@@ -27,6 +27,33 @@
# VM pools, temp pools, etc.). Specifying a pool by name bypasses the ignore list. # VM pools, temp pools, etc.). Specifying a pool by name bypasses the ignore list.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Scrub While Online
# A scrub runs against a live, in-use pool by design. Nothing is unmounted and no
# service is stopped — the cost is background I/O, not downtime, which is what makes
# monthly cadence practical at all.
#
# Detect Before Redundancy Is Gone
# Silent corruption otherwise surfaces only when the bad block is finally read, which
# is often during a resilver — precisely when the redundancy needed to repair it is
# already spent. Scrubbing is what moves that discovery to a moment when ZFS can still
# fix it from a good copy.
#
# Never Start a Second Scrub
# A pool already scrubbing is left alone rather than restarted. Restarting discards the
# progress of the run in flight and begins the whole read again.
#
# Per-Pool Isolation
# One pool failing, being unavailable, or already scrubbing never blocks the others.
#
# Report Errors, Repair Nothing by Hand
# ZFS self-heals from redundancy during the scrub. What this reports is what ZFS could
# not fix — those are operator decisions about hardware, not something a script should
# attempt to resolve.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
+48 -3
View File
@@ -30,6 +30,37 @@
# /tmp resets on reboot — correct, transcode state should not persist across boots. # /tmp resets on reboot — correct, transcode state should not persist across boots.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Size Is a Ceiling, Not a Reservation
# tmpfs allocates on write. RAMDISK_SIZE caps how large the ramdisk may grow; it does
# not take that RAM away from the system up front. Sizing it generously costs nothing
# until transcodes actually fill it, which is why the ceiling can sit well above
# normal usage without starving anything.
#
# Clean Symlink State Every Boot
# TRANSCODE_LINK is reset to the ramdisk at every array start rather than left wherever
# the last flip put it. transcode_manager.sh flips it to SSD under pressure, and that
# flip is a runtime response to a full ramdisk — carrying it across a reboot would mean
# starting on the fallback with an empty ramdisk sitting unused.
#
# Pre-Create Before Emby Starts
# transcoding-temp/ is created on the ramdisk before any container launches. Emby
# searches accessible paths for an existing transcoding-temp at startup and binds to
# the first it finds — if only the SSD copy exists, every session lands there until
# Emby is restarted. Ordering here is not cosmetic; it decides where transcodes go.
#
# Idempotent Re-Runs
# An already-mounted ramdisk is left mounted and only the symlink and permissions are
# verified. Re-running never tears down a mount that active sessions are writing into.
#
# State Belongs in /tmp
# The transcode state DB lives in /tmp and resets on reboot. Flip counters and the
# current target describe a running system; carrying them across a boot would make the
# manager act on pressure that no longer exists.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -43,9 +74,6 @@
# If RAMDISK_PATH is already a mountpoint, reports status and exits cleanly # If RAMDISK_PATH is already a mountpoint, reports status and exits cleanly
# without attempting to remount or changing anything. # without attempting to remount or changing anything.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# Silent on Success # Silent on Success
# Startup script runs on every boot — no output when healthy. # Startup script runs on every boot — no output when healthy.
# #
@@ -127,6 +155,23 @@ acquire_lock
# detect_hosts() sets MY_ID and aliases RAMDISK_SIZE, TRANSCODE_SSD etc. # detect_hosts() sets MY_ID and aliases RAMDISK_SIZE, TRANSCODE_SSD etc.
detect_hosts detect_hosts
# This script mounts a tmpfs over RAMDISK_PATH and, when TRANSCODE_LINK exists but is not a
# symlink, rm -rf's it before replacing it. Neither of those checks catches a collapsed path:
# / and /mnt both satisfy -e, and mounting a tmpfs over a system directory hides its contents
# for the life of the mount. Require at least two path components before either is touched.
for _tc_pair in "RAMDISK_PATH:$RAMDISK_PATH" "TRANSCODE_LINK:$TRANSCODE_LINK"; do
_tc_name="${_tc_pair%%:*}"
_tc_path="${_tc_pair#*:}"
_tc_slashes="${_tc_path//[^\/]/}"
if [[ -z "$_tc_path" || "$_tc_path" != /* || "${#_tc_slashes}" -lt 2 ]]; then
error "$_tc_name is unset or unsafe ('${_tc_path:-unset}') — refusing to mount or relink"
notify "Ramdisk setup aborted on $(hostname) ($MY_ID) — $_tc_name is '${_tc_path:-unset}'" \
"Ramdisk Setup" "warning"
exit 1
fi
done
unset _tc_pair _tc_name _tc_path _tc_slashes
log "Identity: $MY_ID ($LOCAL_SERVER_NAME)" log "Identity: $MY_ID ($LOCAL_SERVER_NAME)"
log "Ramdisk: $RAMDISK_PATH ($RAMDISK_SIZE)" log "Ramdisk: $RAMDISK_PATH ($RAMDISK_SIZE)"
log "Fallback: $TRANSCODE_SSD" log "Fallback: $TRANSCODE_SSD"
+52
View File
@@ -15,9 +15,41 @@
# 2. Not currently open by any process (checked via lsof pre-built map) # 2. Not currently open by any process (checked via lsof pre-built map)
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# Two locations cleaned in sequence, each through the same routine:
#
# Ramdisk — only if mountpoint -q confirms it is actually mounted. An unmounted
# ramdisk means the underlying directory is the real filesystem, and
# cleaning it would delete from disk rather than from tmpfs.
# SSD fallback — only if the directory exists.
#
# Per location:
# 1. Path sanity check — refuse anything shallower than two components
# 2. Count total and age-eligible files (-mmin +TRANSCODE_MAX_AGE)
# 3. One lsof +D call → in-memory open-file map for the whole location
# 4. Per eligible file: open → skip and count as active; otherwise rm -f
# 5. Remove empty directories older than TRANSCODE_ORPHAN_AGE, never transcoding-temp
#
# Afterwards, if the ramdisk recovered enough headroom, the symlink is flipped back to
# it so new sessions return to RAM.
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
# Two Independent Conditions, Both Required
# Age and open-file state are checked separately and a file must pass both. Age alone
# would delete a long-running session's segments; lsof alone cannot see HLS segments,
# which are written and closed atomically. Neither signal is sufficient on its own,
# which is exactly why both are applied rather than picking the better one.
#
# Mounted-Only Ramdisk Cleaning
# The ramdisk is only cleaned when it is genuinely mounted. If the tmpfs failed to
# mount, that same path is an ordinary directory on the array — cleaning it then would
# delete real files from disk while believing it was clearing RAM.
#
# lsof Called Once, Not Per File # lsof Called Once, Not Per File
# On a busy Live TV system the ramdisk contains thousands of HLS segment files. # On a busy Live TV system the ramdisk contains thousands of HLS segment files.
# Calling lsof once per file creates thousands of subprocess calls every 7 minutes. # Calling lsof once per file creates thousands of subprocess calls every 7 minutes.
@@ -55,6 +87,13 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Location Path Guard
# cleanup_location() refuses any path that is not absolute with at least two
# components. Neither caller's own check catches a collapsed value — mountpoint -q
# returns true for /, and -d is true for both / and /mnt — so the guard lives inside
# the function that does the deleting, covering both call sites.
#
#
# Wait Lock # Wait Lock
# acquire_lock "wait" — waits if a previous cleanup run is still active rather # acquire_lock "wait" — waits if a previous cleanup run is still active rather
# than exiting. The caller's 7-minute interval can overlap on a slow system. # than exiting. The caller's 7-minute interval can overlap on a slow system.
@@ -188,6 +227,19 @@ cleanup_location() {
return return
fi fi
# This function deletes every file under $location past the age gate. Neither caller's
# own check catches a collapsed path: mountpoint -q returns true for /, and -d is true
# for / and /mnt alike. Require at least two path components so a blank or truncated
# RAMDISK_PATH / TRANSCODE_SSD can never point this at a system directory.
local _loc_slashes="${location//[^\/]/}"
if [[ "$location" != /* || "${#_loc_slashes}" -lt 2 ]]; then
error "Refusing to clean unsafe location: '$location' ($label)"
notify "Transcode cleanup refused unsafe path on $(hostname): '$location'" \
"Transcode Cleanup" "warning"
LOCATION_REMOVED=0 LOCATION_FREED="0B" LOCATION_SKIPPED=0 LOCATION_ACTIVE=0 LOCATION_STREAMING=0 LOCATION_TOO_YOUNG=0 LOCATION_FAILED=0
return
fi
local file_count eligible_count local file_count eligible_count
file_count=$(find "$location" -type f 2>/dev/null | wc -l) file_count=$(find "$location" -type f 2>/dev/null | wc -l)
eligible_count=$(find "$location" -type f -mmin +"$max_age" 2>/dev/null | wc -l) eligible_count=$(find "$location" -type f -mmin +"$max_age" 2>/dev/null | wc -l)
+30 -3
View File
@@ -45,6 +45,36 @@
# Format: DATE|RAMDISK_USED_GB|FLIP_COUNT|RAM_SESSION_COUNT|SSD_SESSION_COUNT|FILES_CLEANED # Format: DATE|RAMDISK_USED_GB|FLIP_COUNT|RAM_SESSION_COUNT|SSD_SESSION_COUNT|FILES_CLEANED
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Flip the Destination, Never Move the Sessions
# Switching targets only repoints the symlink so NEW sessions land elsewhere. Existing
# transcodes keep writing to the path they opened and drain naturally. Moving files
# mid-transcode would break every stream currently playing, which is the opposite of
# what pressure relief is for.
#
# Asymmetric Thresholds
# Flipping away happens at RAMDISK_WARN_GB, flipping back at RAMDISK_LOW_GB — two
# separate values, not one. A single threshold would flip on every fluctuation around
# it; the gap between them is what makes the decision stable under load.
#
# Safety Checks Are Unconditional
# Symlink, ramdisk presence, SSD presence, transcoding-temp and permissions are all
# verified on every run in every mode, including the fixed ramdisk/ssd modes. Mode
# controls where transcodes go, not whether the plumbing gets checked.
#
# Degrade Toward the Fallback
# Every failure path resolves toward SSD, never toward an unusable target. A vanished
# ramdisk flips to SSD immediately rather than leaving sessions pointed at nothing —
# transcoding slower is recoverable, transcoding nowhere is not.
#
# Observe Without Emby
# With Emby down, threshold logic is skipped but the symlink is still verified. There
# is no session pressure to react to, and acting on stale usage would flip the target
# for sessions that no longer exist.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -55,9 +85,6 @@
# Docker Timeout # Docker Timeout
# DOCKER_TIMEOUT caps all docker calls against a hung daemon. # DOCKER_TIMEOUT caps all docker calls against a hung daemon.
# #
# Notification Validated
# platform_require_cmd confirms the notify script is present before use.
#
# Silent by Default # Silent by Default
# Runs every 7 minutes — only speaks when something changes or needs attention. # Runs every 7 minutes — only speaks when something changes or needs attention.
# #
+107 -6
View File
@@ -21,6 +21,25 @@
# No-op when FALLBACK_ENABLED=false or CONF_SYNC_ENABLED=false. # No-op when FALLBACK_ENABLED=false or CONF_SYNC_ENABLED=false.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL MODEL
# ==============================================================================================
#
# One decision, driven entirely by remote reachability:
#
# 1. Gates
# → PARTNERSHIP_ENABLED, FALLBACK_ENABLED, CONF_SYNC_ENABLED, REMOTE_ID
# → any gate closed means exit 0, no work, no output
#
# 2. ping_remote
# REACHABLE → remove $PERSISTENT_CONF_CACHE if it exists, exit
# UNREACHABLE → refresh the backup from the RAM cache
#
# 3. Refresh (remote offline only)
# → copy every host*.conf from the RAM cache except this host's own
# → own conf is excluded: it is already on disk, the backup exists
# solely to survive a reboot without the partner's vars
#
# ==============================================================================================
# DESIGN PRINCIPLES # DESIGN PRINCIPLES
# ============================================================================================== # ==============================================================================================
# #
@@ -38,11 +57,70 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# require_partnership — exits early if PARTNERSHIP_ENABLED=false # Root Enforcement
# FALLBACK_ENABLED gate — exits if fallback is disabled # Writes to and removes $PERSISTENT_CONF_CACHE, which lives under the plugin
# CONF_SYNC_ENABLED gate — exits if conf sync is disabled # directory and is not user-writable.
# REMOTE_ID presence check — exits if partner identity is unset #
# --dry-run mode — shows what would happen without touching the backup # Lock Acquisition
# acquire_lock() prevents concurrent execution. Without it a slow run can still
# be copying confs into the backup while the next run, seeing the remote back
# online, rm -rf's the directory out from under it.
#
# Cache Path Sanity Guard
# PERSISTENT_CONF_CACHE is validated as an absolute path at least three levels
# deep before any rm -rf. It is built from ${SCRIPTS_DIR} — if that is ever
# unset the path collapses toward the filesystem root, and this script would
# otherwise recursively delete whatever it collapsed to.
#
# Partnership Gate
# require_partnership() exits early if PARTNERSHIP_ENABLED=false.
#
# FALLBACK_ENABLED / CONF_SYNC_ENABLED Gates
# Exits cleanly if either is disabled — the backup only has meaning when
# fallback can actually consume it.
#
# REMOTE_ID Presence Check
# Exits if partner identity is unset. Without a partner there is nothing to
# back up and the own-conf exclusion below could not be applied correctly.
#
# Own-Conf Exclusion
# This host's own conf is never written into the partner backup. Restoring it
# later would overwrite live local config with a stale copy.
#
# Dry Run Support
# --dry-run reports every removal and copy without touching the backup.
#
# ==============================================================================================
# CONFIGURATION
# ==============================================================================================
#
# master.conf
#
# PERSISTENT_CONF_CACHE
# Destination for the partner conf backup. Must survive a reboot, so it
# lives under ${SCRIPTS_DIR}, not in /tmp.
#
# FALLBACK_ENABLED
# Master fallback toggle. Backup is pointless when fallback cannot run.
#
# CONF_SYNC_ENABLED
# Conf sync toggle. When off, no RAM cache is being maintained to back up.
#
# PARTNERSHIP_ENABLED
# Checked via require_partnership().
#
# ==============================================================================================
# RUNTIME MODES
# ==============================================================================================
#
# conf_cache_watchdog.sh
# Refresh or remove the persistent partner conf backup based on remote state
#
# conf_cache_watchdog.sh --dry-run
# Report what would be written or removed without changing the backup
#
# conf_cache_watchdog.sh --log
# Verbose per-file output
# #
# ============================================================================================== # ==============================================================================================
@@ -50,6 +128,17 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../../load_config.sh" source "$SCRIPT_DIR/../../load_config.sh"
parse_args "$@" parse_args "$@"
# ==============================================================================================
# ━━━ Setup ━━━
# ==============================================================================================
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock
detect_hosts detect_hosts
require_partnership require_partnership
@@ -58,7 +147,19 @@ require_partnership
[[ -z "${REMOTE_ID:-}" ]] && exit 0 [[ -z "${REMOTE_ID:-}" ]] && exit 0
RAM_CACHE="/tmp/.cache/vv/d" RAM_CACHE="/tmp/.cache/vv/d"
SAVE_DIR="$PERSISTENT_CONF_CACHE" SAVE_DIR="${PERSISTENT_CONF_CACHE:-}"
# SAVE_DIR is rm -rf'd below and is built from ${SCRIPTS_DIR}. If that is ever unset the
# path collapses toward / — require an absolute path at least three levels deep so a
# collapsed or empty value can never name a system directory.
_slashes="${SAVE_DIR//[^\/]/}"
if [[ -z "$SAVE_DIR" || "$SAVE_DIR" != /* || "${#_slashes}" -lt 3 ]]; then
error "PERSISTENT_CONF_CACHE is unset or unsafe ('${SAVE_DIR:-unset}') — refusing to manage conf backup"
notify "conf_cache_watchdog aborted on $(hostname) — PERSISTENT_CONF_CACHE is '${SAVE_DIR:-unset}'" \
"Conf Cache Watchdog" "warning"
exit 1
fi
unset _slashes
if ping_remote; then if ping_remote; then
if [[ -d "$SAVE_DIR" ]]; then if [[ -d "$SAVE_DIR" ]]; then
+71 -11
View File
@@ -33,6 +33,33 @@
# External check — verifies the full stack (DNS → NPM → backend), not just NPM running. # External check — verifies the full stack (DNS → NPM → backend), not just NPM running.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Gate on Root Cause
# The internet check runs first and short-circuits everything after it. DDNS
# and NPM checks both depend on outbound connectivity — running them during an
# outage produces three alarms for one fault and can trigger container restarts
# that fix nothing.
#
# Verify the Path, Not the Process
# NPM is checked by fetching an external URL rather than asking whether the
# container is running. A running container behind broken DNS or a broken
# upstream still serves nothing. Checking the whole path is the only result
# that means anything to a user.
#
# Restart Only What Restarting Fixes
# DDNS and NPM are restarted because a restart forces a record update or
# reloads proxy config — the restart is the remedy. Tailscale is notify-only:
# its failures are auth, key expiry or ACL problems that a restart cannot
# resolve and may obscure.
#
# Strike Before Restarting NPM
# NPM sits in front of every externally reachable service, so restarting it is
# itself disruptive. A single failed fetch can be a transient upstream blip;
# two consecutive failures justify the interruption.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -42,6 +69,19 @@
# Single Instance Lock # Single Instance Lock
# acquire_lock prevents concurrent runs. # acquire_lock prevents concurrent runs.
# #
# Host Detection Before Config Resolution
# detect_hosts() runs before the HOST*_NETWORK_WATCHDOG_* names are built. MY_ID
# is not exported, so an orchestrated run starts with it empty — resolving these
# any earlier silently produces empty DDNS and NPM config and skips both checks.
#
# Unresolved Config Warning
# Warns when neither DDNS nor NPM config resolves for this host. "Not configured"
# is a legitimate state, but it looks identical to a broken lookup, so it is
# stated out loud rather than passed over in silence.
#
# NETWORK_WATCHDOG_ENABLED Toggle
# Exits cleanly when disabled, without removing it from the orchestrator list.
#
# Internet Gates All Checks # Internet Gates All Checks
# If internet is down, DDNS and NPM checks are skipped — no cascade of false positives. # If internet is down, DDNS and NPM checks are skipped — no cascade of false positives.
# #
@@ -49,6 +89,17 @@
# Single curl failure could be transient DNS hiccup or CDN blip. # Single curl failure could be transient DNS hiccup or CDN blip.
# Two consecutive failures confirms NPM is the problem. # Two consecutive failures confirms NPM is the problem.
# #
# Notify-Only for Tailscale
# Tailscale is never restarted. Its failures are auth, key expiry or ACL issues
# that a restart cannot fix and would only obscure.
#
# Timeout Protection
# Every connectivity probe carries an explicit timeout, so a black-holed route
# cannot stall the every-minute watchdog chain.
#
# Dry Run Support
# --dry-run performs all checks and reports restarts without issuing them.
#
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
@@ -93,9 +144,24 @@ source "$SCRIPT_DIR/../../load_config.sh"
parse_args "$@" parse_args "$@"
# ==============================================================================================
# ━━━ Setup ━━━
# ==============================================================================================
if [[ "$EUID" -ne 0 ]]; then
error "Must be run as root"
exit 1
fi
acquire_lock
detect_hosts
# ============================================================================================== # ==============================================================================================
# ━━━ Resolve host-specific config ━━━ # ━━━ Resolve host-specific config ━━━
# ============================================================================================== # ==============================================================================================
# MUST run after detect_hosts() — MY_ID is set there and is not exported, so orchestrated
# runs start with it empty. Building these names any earlier yields "_NETWORK_WATCHDOG_*",
# which is always unset, and the DDNS and NPM checks silently skip as "not configured".
_ddns_domain_var="${MY_ID}_NETWORK_WATCHDOG_DDNS_DOMAIN" _ddns_domain_var="${MY_ID}_NETWORK_WATCHDOG_DDNS_DOMAIN"
_ddns_container_var="${MY_ID}_NETWORK_WATCHDOG_DDNS_CONTAINER" _ddns_container_var="${MY_ID}_NETWORK_WATCHDOG_DDNS_CONTAINER"
_npm_url_var="${MY_ID}_NETWORK_WATCHDOG_NPM_URL" _npm_url_var="${MY_ID}_NETWORK_WATCHDOG_NPM_URL"
@@ -104,19 +170,13 @@ DDNS_DOMAIN="${!_ddns_domain_var:-}"
DDNS_CONTAINER="${!_ddns_container_var:-}" DDNS_CONTAINER="${!_ddns_container_var:-}"
NPM_URL="${!_npm_url_var:-}" NPM_URL="${!_npm_url_var:-}"
# ============================================================================================== # A check that is configured but resolves empty means the lookup broke, not that the
# ━━━ Setup ━━━ # operator opted out. Say so — silent skipping is what hid this for so long.
# ============================================================================================== if [[ -z "$DDNS_DOMAIN" && -z "$NPM_URL" ]]; then
if [[ "$EUID" -ne 0 ]]; then warn "No DDNS or NPM config resolved for ${MY_ID:-unknown host}"
error "Must be run as root" warn "Expected ${MY_ID}_NETWORK_WATCHDOG_DDNS_DOMAIN / ${MY_ID}_NETWORK_WATCHDOG_NPM_URL in host*.conf"
exit 1
fi fi
acquire_lock
detect_hosts
[[ "${NETWORK_WATCHDOG_ENABLED:-true}" != "true" ]] && echo "Network watchdog disabled" && exit 0 [[ "${NETWORK_WATCHDOG_ENABLED:-true}" != "true" ]] && echo "Network watchdog disabled" && exit 0
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be restarted" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be restarted"
+57
View File
@@ -73,6 +73,63 @@
# attention. # attention.
# #
# ============================================================================================== # ==============================================================================================
# OPERATIONAL SAFEGUARDS
# ==============================================================================================
#
# Root Enforcement
# Truncating container-owned log files requires root.
#
# Lock Acquisition
# acquire_lock() prevents concurrent runs. Two instances would race on the
# strike state file and the growth baseline, double-counting strikes and
# potentially truncating a file one cycle early.
#
# Host Detection
# detect_hosts() aliases HOST*_WATCHDOG_APPDATA_SIZES to the correct host's
# suppress ceilings.
#
# WATCHDOG_CHECK_APPDATA Toggle
# Exits cleanly before any scanning when the master toggle is off.
#
# Path Existence Guard
# Every entry in WATCHDOG_APPDATA_PATHS is skipped unless it is a non-empty
# string naming a real directory. An unconfigured array cannot cause a scan
# from an unintended location.
#
# Truncate-Never-Delete
# Action is always truncate -s 0, never rm. The container keeps its open file
# handle and space is reclaimed immediately, so a still-running service does
# not lose its log destination mid-write.
#
# Filename Restriction
# Only *.log and *.log.* files are ever truncation candidates. Databases,
# caches, game saves and every other growing file are alert-only — detected
# and reported, never modified.
#
# Truncation Opt-In
# WATCHDOG_APPDATA_TRUNCATE_LOGS defaults to false. Without it explicitly
# enabled the action cycle escalates to a critical notification and holds
# strikes rather than touching any file.
#
# Strike Threshold
# Nothing acts on first detection. WATCHDOG_APPDATA_STRIKE_LIMIT consecutive
# cycles are required, separating a legitimate library scan or save burst
# from a genuine runaway. Strikes auto-clear when the condition resolves.
#
# Suppress Ceiling
# Containers listed in WATCHDOG_APPDATA_SIZES are exempt from growth alerts
# while under their configured ceiling — prevents known-large stable data
# from generating recurring false alarms.
#
# Dry Run Support
# --dry-run reports every truncation that would occur and performs none.
#
# Atomic Baseline Update
# The growth baseline is written to a temp file and moved into place, so an
# interrupted run cannot leave a half-written baseline that would read as
# false growth on the next cycle.
#
# ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
# #
+16 -2
View File
@@ -89,6 +89,17 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root Enforcement
# Container restarts and daemon service control require root.
#
# Docker Enabled Check
# Exits cleanly when Docker is disabled in the platform's own settings. A
# deliberately disabled Docker service is not a fault and must not be
# "healed" by restarting the daemon.
#
# Docker Presence Check
# Verifies the docker binary exists before the cycle begins.
#
# Lock Acquisition # Lock Acquisition
# Prevents concurrent execution via acquire_lock(). Safe at array start — # Prevents concurrent execution via acquire_lock(). Safe at array start —
# only one watchdog instance runs at a time. # only one watchdog instance runs at a time.
@@ -110,8 +121,11 @@
# #
# Docker Daemon Health Check # Docker Daemon Health Check
# First operation every cycle. Daemon not responding within DOCKER_TIMEOUT → # First operation every cycle. Daemon not responding within DOCKER_TIMEOUT →
# restart via /etc/rc.d/rc.docker → verify recovery. If still hung: log # restart via platform_restart_service docker → verify recovery. If still hung:
# critical, skip cycle. stability_watchdog.sh handles further escalation. # log critical, skip cycle, and set daemon_confirmed_down so
# stability_watchdog.sh owns any further escalation. The restart itself is
# bounded by a 180 second timeout — a daemon stop can block for 30+ minutes on
# a busy host, and the watchdog must not be held hostage to it.
# #
# RAM Emergency Deferral # RAM Emergency Deferral
# Reads RW_STATE_FILE each cycle. If resource_watchdog.sh has set # Reads RW_STATE_FILE each cycle. If resource_watchdog.sh has set
+59 -1
View File
@@ -42,6 +42,32 @@
# Cleared when pressure resolves and containers are restarted. # Cleared when pressure resolves and containers are restarted.
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Graduated Response
# Pressure is answered with the smallest effective action first — throttle,
# then pause, then stop. Each level is only reached because the level below it
# failed to relieve pressure. Nothing jumps straight to stopping containers.
#
# Reversibility First
# docker pause suspends a container without losing its state and is instantly
# reversible, so it is preferred at level 2. docker stop, which discards
# in-memory state, is held back to level 3 and applied only to services
# explicitly listed as expendable in RW_STOP_CONTAINERS.
#
# Hysteresis on Recovery
# Restoring requires RW_RECOVER_CYCLES consecutive clear cycles and
# de-escalates one level at a time. Recovering instantly on a single good
# reading would flap — restore, re-trigger, restore — under sustained load.
#
# Cross-Watchdog Coordination
# Level 3 publishes mem_shutdown_active=true so docker_watchdog.sh defers its
# restart logic. Two watchdogs acting on the same containers with opposite
# intent would otherwise fight: one stopping to free RAM, the other restarting
# to restore health.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -51,13 +77,45 @@
# Single Instance Lock # Single Instance Lock
# acquire_lock prevents concurrent runs from racing on state file writes. # acquire_lock prevents concurrent runs from racing on state file writes.
# #
# Docker Presence Check
# Verifies the docker binary exists before any pressure response — every
# level-2 and level-3 action depends on it.
#
# Host Detection
# detect_hosts() aliases HOST*_RW_PAUSE_CONTAINERS, HOST*_RW_STOP_CONTAINERS
# and the downloader credentials to the correct host's values.
#
# State File Verification
# Exits if RW_STATE_FILE cannot be created. Without durable state the script
# cannot track recovery cycles or know which containers it paused, and would
# never restore them.
#
# RW_CRITICAL_CONTAINERS # RW_CRITICAL_CONTAINERS
# Containers listed here are never paused or stopped regardless of pressure level. # Containers listed here are never paused or stopped regardless of pressure
# level. Enforced by is_critical(), which gates both the pause and the stop
# path — not just the configuration lists.
# #
# RW_ENABLED Flag # RW_ENABLED Flag
# Set RW_ENABLED=false to disable the entire script without removing it from # Set RW_ENABLED=false to disable the entire script without removing it from
# the orchestrator schedule. # the orchestrator schedule.
# #
# Timeout Protection
# All docker commands wrapped in a 15 second timeout. Pressure response runs
# during a degraded system, which is exactly when the daemon is most likely
# to be slow — a hang here would stall the whole watchdog chain every minute.
#
# Downloader Availability Guards
# SABnzbd and qBittorrent throttling no-ops when the service is disabled or
# its URL/credentials are unset. A missing downloader never blocks the
# container-level pressure response.
#
# Recovery Hysteresis
# Restoration requires RW_RECOVER_CYCLES consecutive clear cycles and
# de-escalates one level per cycle, preventing flapping under sustained load.
#
# Dry Run Support
# --dry-run reports every throttle, pause and stop without performing any.
#
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
# ============================================================================================== # ==============================================================================================
+86 -2
View File
@@ -54,6 +54,39 @@
# (Container health is owned by docker_watchdog — not checked here.) # (Container health is owned by docker_watchdog — not checked here.)
# #
# ============================================================================================== # ==============================================================================================
# DESIGN PRINCIPLES
# ==============================================================================================
#
# Last Line of Defense
# Every other watchdog tries to heal a specific subsystem. This one assumes
# those attempts have already failed and holds the only irreversible remedy in
# the ecosystem — a reboot. That authority is why nearly every check here is
# gated behind strikes, tiers and abort conditions.
#
# Evidence Before Reboot
# Strikes are the default; bypassing them requires corroboration, not just a
# worse number. Tier 2 needs low RAM AND active OOM kills before it acts —
# low RAM alone is a reading, low RAM plus processes being killed is a crisis.
#
# Unrecoverable Conditions Skip the Queue
# Tier 1 conditions share one property: the system cannot heal from them and
# waiting makes recovery less likely. A full rootfs or a kernel oops degrades
# further every cycle, and strike-counting through it only guarantees the
# reboot happens from a worse state.
#
# Data Safety Outranks Uptime
# Reboots abort while a ZFS pool is unhealthy, parity is running, or the mover
# is active. Interrupting those risks the data itself, which no amount of
# uptime justifies. Tier 1 is the sole exception — an imminent crash will
# interrupt them anyway, less gracefully.
#
# Clear Ownership Boundaries
# Container health belongs to docker_watchdog.sh and is deliberately not
# checked here. The Docker daemon check writes a flag for docker_watchdog
# rather than acting on it. Two watchdogs remediating the same subsystem
# would race.
#
# ==============================================================================================
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
@@ -61,10 +94,61 @@
# Reboot and container stop require root. # Reboot and container stop require root.
# #
# Single Instance Lock # Single Instance Lock
# acquire_lock prevents a second watchdog instance from starting. # acquire_lock prevents a second watchdog instance from starting. Two instances
# could each count strikes against the same condition and reach the reboot
# threshold in half the intended time.
# #
# State File Verification # State File Verification
# All state files verified writable at startup — errors if any cannot be created. # All state files verified writable at startup — errors if any cannot be created.
# Strike counts and the reboot log live in these files; if they silently failed
# to persist, every cycle would look like strike 1 and the reboot rate limit
# would never accumulate.
#
# Reboot Rate Limiting
# No more than SYS_WATCHDOG_REBOOT_LIMIT reboots within
# SYS_WATCHDOG_REBOOT_WINDOW_HRS. On hitting the limit the host powers off
# instead of rebooting — a fault that survives repeated reboots will not be
# fixed by more of them, and a box cycling endlessly is worse than one that
# is cleanly down and obviously needs attention.
#
# Abort Conditions
# Reboots are aborted while a ZFS pool is unhealthy, parity is running, or the
# mover is active — each individually toggleable. Interrupting any of these
# risks the data itself.
#
# Critical Tier Override
# Tier 1 conditions bypass both strikes and abort conditions. These are states
# the system cannot recover from and which degrade every cycle; waiting only
# guarantees the eventual reboot happens from a worse position.
#
# Strike Threshold
# Tier 3 requires SYS_WATCHDOG_STRIKE_LIMIT consecutive failing cycles. A
# single bad sample — a momentary load spike, a transient RAM dip — never
# reboots the system.
#
# OOM Corroboration
# Tier 2 escalation requires low RAM AND active OOM kills in the same cycle.
# Low RAM alone stays in the strike system.
#
# Ownership Boundary
# Container health is not checked here — docker_watchdog.sh owns it. The Docker
# daemon check writes daemon_confirmed_down for docker_watchdog rather than
# remediating, so the two never act on the same subsystem.
#
# Aborted-Reboot Recovery
# An EXIT trap is armed the moment containers start being stopped for a reboot
# and disarmed only once the reboot is committed. If the script dies anywhere
# in between, the trap restarts everything it stopped — the failure mode is a
# running system, never a host left with all containers down and no reboot.
#
# Sync Before Reboot
# sync is issued before both /sbin/poweroff and /sbin/reboot so pending writes
# are flushed. Container stop is additionally bounded by a 60 second timeout so
# one unresponsive container cannot hold the shutdown sequence open forever.
#
# Dry Run Support
# --dry-run runs the full detection path and reports the reboot or shutdown
# that would occur without issuing either.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -74,7 +158,7 @@
# Full variable listing in master.conf. Key variables: # Full variable listing in master.conf. Key variables:
# #
# SYS_WATCHDOG_REBOOT_WINDOW_HRS — reboot rate limit window (default: 12) # SYS_WATCHDOG_REBOOT_WINDOW_HRS — reboot rate limit window (default: 12)
# SYS_WATCHDOG_MAX_REBOOTS — max reboots in window before giving up (default: 3) # SYS_WATCHDOG_REBOOT_LIMIT — max reboots in window before giving up (default: 3)
# SYS_WATCHDOG_STRIKE_LIMIT — consecutive failures before reboot (default: 2) # SYS_WATCHDOG_STRIKE_LIMIT — consecutive failures before reboot (default: 2)
# SYS_WATCHDOG_OOM_LIMIT — OOM kills/cycle to trigger URGENT bypass (default: 3) # SYS_WATCHDOG_OOM_LIMIT — OOM kills/cycle to trigger URGENT bypass (default: 3)
# SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED — containers exempt from memory shutdown # SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED — containers exempt from memory shutdown
+34 -4
View File
@@ -35,10 +35,33 @@
# OPERATIONAL SAFEGUARDS # OPERATIONAL SAFEGUARDS
# ============================================================================================== # ==============================================================================================
# #
# Root check — child scripts require root # Root Enforcement
# acquire_lock — prevents concurrent system watchdog runs # Every child script requires root. Failing here gives one clear error instead
# detect_hosts() — MY_ID in notifications and logs # of the same permission failure repeated once per child.
# Non-fatal steps — a failed step is logged; remaining steps still run #
# Lock Acquisition
# acquire_lock() prevents concurrent system watchdog runs. This is called every
# cycle by watchdog_orchestrator.sh — a slow child must not cause two chains to
# overlap and run the same watchdog twice.
#
# Host Detection
# detect_hosts() sets MY_ID for notifications and logs.
#
# Empty List Guard
# Warns and exits if SYSTEM_WATCHDOG_SCRIPTS is unconfigured. An empty list
# would otherwise report "0/0 passed" every cycle — indistinguishable from
# healthy, while no system monitoring is actually running.
#
# Missing Script Tolerance
# run_orch_child() records a missing or failing child as a failed step and
# continues. One broken watchdog never suppresses the rest of the chain.
#
# Non-Fatal Steps
# A failed step is logged and surfaces in the summary and notification, but
# remaining steps still execute. Partial coverage beats a halted chain.
#
# Dry Run Propagation
# --dry-run and --log are passed through to every child script.
# #
# ============================================================================================== # ==============================================================================================
# CONFIGURATION # CONFIGURATION
@@ -86,6 +109,13 @@ acquire_lock
detect_hosts detect_hosts
# An empty list reports "0/0 passed" every cycle — reads as healthy while nothing is monitored.
if [[ ${#SYSTEM_WATCHDOG_SCRIPTS[@]} -eq 0 ]]; then
warn "SYSTEM_WATCHDOG_SCRIPTS is empty — no system watchdogs will run"
warn "Check SYSTEM_WATCHDOG_SCRIPTS in master.conf"
exit 0
fi
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — passing --dry-run to all sub-scripts" [[ "$DRY_RUN" == true ]] && warn "DRY RUN — passing --dry-run to all sub-scripts"
# ============================================================================================== # ==============================================================================================
+3 -1
View File
@@ -116,7 +116,9 @@
unset _VV_CONF_CACHE _conf_base _disk_conf_basenames unset _VV_CONF_CACHE _conf_base _disk_conf_basenames
# ━━━ Source shared functions ━━━ # ━━━ Source shared functions ━━━
# common.sh sourced last — it calls detect_hosts() which needs HOST* vars to be set. # common.sh sourced last — it defines detect_hosts(), which needs HOST* vars to be set.
# It does NOT call it. Each script calls detect_hosts() itself, so MY_ID and REMOTE_ID
# are unset until it does — anything building HOST*-prefixed variable names must come after.
if [[ ! -f "$LOAD_CONFIG_DIR/common.sh" ]]; then if [[ ! -f "$LOAD_CONFIG_DIR/common.sh" ]]; then
echo "[FATAL] common.sh not found at $LOAD_CONFIG_DIR/common.sh" >&2 echo "[FATAL] common.sh not found at $LOAD_CONFIG_DIR/common.sh" >&2
exit 1 exit 1