did way to much,,,,,, mostly added monitors, but almost every file was edited in some way
This commit is contained in:
+413
-222
@@ -5,6 +5,13 @@
|
||||
# All user-facing variables for the unRAID script ecosystem.
|
||||
# Scripts source this file — edit here, changes apply everywhere on next git pull.
|
||||
#
|
||||
# ── HOW THIS FILE WORKS ───────────────────────────────────────────────────────────────────────
|
||||
# Every script sources Master.conf and common.sh at startup.
|
||||
# Change a value here and it affects all scripts that use it — no hunting through files.
|
||||
# To disable something: comment it out with # rather than deleting it.
|
||||
# To add a new rsync profile: add a key to each PROFILE_* array.
|
||||
# To add a new media maintenance job: add a line to MEDIA_MAINTENANCE_JOBS.
|
||||
#
|
||||
# ── INDEX ─────────────────────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# Section Description
|
||||
@@ -36,47 +43,70 @@
|
||||
# PHP-FPM PHP-FPM max children config
|
||||
# CLEAR LOGS System log file paths
|
||||
# WEBGUI WATCHDOG WebGUI nginx + emhttp monitoring and restart
|
||||
# ZFS MEMORY SNAPSHOT Weekly ZFS health and memory diagnostic report
|
||||
#
|
||||
# ── MEDIA ──────────────────────────────────────────────────────────────────────────────────
|
||||
# MEDIA PERMISSIONS Share list, mode and owner for permissions script
|
||||
# MEDIA CLEANER Anime and media folder lists and file patterns
|
||||
# MEDIA MANAGEMENT Orchestrator job list for media_management.sh
|
||||
# ARR CLEANUP Lidarr, Sonarr, Radarr orphan file cleanup
|
||||
#
|
||||
# ── TRANSCODES ─────────────────────────────────────────────────────────────────────────────
|
||||
# TRANSCODE MANAGER Ramdisk and SSD fallback transcode management
|
||||
#
|
||||
# ── MONITOR ────────────────────────────────────────────────────────────────────────────────
|
||||
# CERTIFICATE MONITOR SSL certificate expiry monitoring
|
||||
# BACKUP VERIFY Random sample checksum verification against remote
|
||||
# SMART HEALTH Drive SMART attribute monitoring
|
||||
# BANDWIDTH MONITOR Daily rsync transfer logging and weekly summary
|
||||
# HEALTH DIGEST Aggregated system health digest — always/smart/weekly
|
||||
# EMBY SESSION REPORT Weekly Emby usage statistics via API
|
||||
#
|
||||
# ── SYSTEM WATCHDOG ────────────────────────────────────────────────────────────────────────
|
||||
# SYSTEM WATCHDOG System health monitoring — last line of defense
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Host Configuration ━━━
|
||||
# Hostnames must match Tailscale machine names exactly — case sensitive
|
||||
# Hostnames must match Tailscale machine names exactly — case sensitive.
|
||||
# Used by detect_hosts() in common.sh to determine which server is local and which is remote.
|
||||
# Both servers run identical scripts — host detection makes them bidirectional.
|
||||
HOST1="unRAID-Gmer4Lfe"
|
||||
HOST2="unRAID-Jayred365"
|
||||
# SSH keys for server-to-server rsync and failover operations
|
||||
# Each server authenticates with its own key — both must be authorised on the remote
|
||||
|
||||
# SSH keys for server-to-server rsync and failover container operations.
|
||||
# HOST1_SSH_KEY is used when HOST1 SSHes to HOST2 and vice versa.
|
||||
# Both keys must be in /root/.ssh/ and authorised in the remote server's authorized_keys.
|
||||
HOST1_SSH_KEY="/root/.ssh/Gmer4Lfe-rsync-key"
|
||||
HOST2_SSH_KEY="/root/.ssh/Jayred365-rsync-key"
|
||||
|
||||
# ━━━ Logging ━━━
|
||||
# true = verbose [LOG] lines in all script output / false = user-facing output only
|
||||
# Controls verbose [LOG] output across all scripts.
|
||||
# true = show detailed [LOG] lines — useful for debugging or first-time setup
|
||||
# false = show only user-facing output — cleaner for scheduled runs
|
||||
ENABLE_LOGGING=true
|
||||
|
||||
# ━━━ Notifications ━━━
|
||||
# unRAID native — fires through Settings → Notification Settings
|
||||
# Recommended: configure unRAID to send errors/warnings only so normal completions stay quiet
|
||||
# Two independent notification channels — either or both can be active simultaneously.
|
||||
|
||||
# unRAID native notification system — integrates with the bell icon in the WebGUI.
|
||||
# Recommended: set Settings → Notification Settings to errors/warnings only so
|
||||
# normal completions don't create noise. The ecosystem sends:
|
||||
# normal — job completed successfully (informational)
|
||||
# warning — something failed or needs attention
|
||||
NOTIFY_UNRAID=true
|
||||
# Discord webhook — paste full webhook URL to enable, leave blank to disable
|
||||
|
||||
# Discord webhook URL — paste the full webhook URL from your Discord server settings.
|
||||
# Leave blank to disable Discord notifications entirely.
|
||||
DISCORD_WEBHOOK=""
|
||||
|
||||
# ━━━ Git / Repo ━━━
|
||||
# Gitea self-hosted repository for the script ecosystem
|
||||
# git_pull_execute.sh uses these to clone or pull the latest version on both servers
|
||||
# Gitea self-hosted repository settings used by git_pull_execute.sh.
|
||||
# Running git_pull_execute.sh on either server pulls the latest scripts and sets
|
||||
# executable permissions automatically — keeps both servers in sync with one command.
|
||||
REPO_SSH="git@192.168.50.2:FailedProxy/Unraid_Scripts.git"
|
||||
TARGET_DIR="/mnt/user/appdata/unraid_scripts"
|
||||
GITEA_SSH_KEY="/root/.ssh/unraid_gitea"
|
||||
SSH_PORT=221
|
||||
GITEA_SSH_KEY="/root/.ssh/unraid_gitea" # SSH key for authenticating to Gitea
|
||||
SSH_PORT=221 # Gitea SSH port — default Gitea uses 22
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RSYNC ─────────────────────────────────────────────────────────────────────────────────────
|
||||
@@ -86,39 +116,35 @@
|
||||
# Global fallback values used when no profile match is found for a directory.
|
||||
# Shares in DAILY_SYNC_SHARES always use these globals — no profile is defined for them.
|
||||
# Appdata shares (Arrs_Stack, Critical-Data etc.) match profiles by directory basename.
|
||||
# If a profile key exists in a PROFILE_* array that value overrides the global.
|
||||
# If a profile key is missing the global below is used as the fallback.
|
||||
|
||||
# Network transfer speed cap in KB/s — 12500 ≈ 100Mbit
|
||||
BW_LIMIT=12500
|
||||
# Number of retry attempts if rsync fails before giving up
|
||||
RETRY_COUNT=3
|
||||
# Seconds to wait between retry attempts
|
||||
SLEEP=300
|
||||
# Containers to stop on remote before rsync and restart after — empty by default
|
||||
# Profiles below override this per appdata share
|
||||
CRITICAL_CONTAINER_NAMES=()
|
||||
# Containers that need a delay before starting — e.g. Authelia needs its DB ready first
|
||||
DELAYED_CONTAINERS=()
|
||||
# Seconds to wait before starting delayed containers
|
||||
CONTAINER_DELAY=5
|
||||
# Directories to exclude from transfer — empty by default, profiles override per share
|
||||
EXCLUDE_DIRS=()
|
||||
# Default rsync options used when no profile match is found
|
||||
# --delete removes files on remote that no longer exist on source
|
||||
# --inplace writes directly to destination file rather than temp file — better for large files
|
||||
# --no-whole-file forces delta transfer even on local-like connections
|
||||
BW_LIMIT=12500 # network transfer speed cap in KB/s — 12500 ≈ 100Mbit
|
||||
RETRY_COUNT=3 # number of retry attempts if rsync fails before giving up
|
||||
SLEEP=300 # seconds to wait between retry attempts
|
||||
CRITICAL_CONTAINER_NAMES=() # containers to stop on REMOTE before rsync — profiles override
|
||||
DELAYED_CONTAINERS=() # containers needing delay before starting — profiles override
|
||||
CONTAINER_DELAY=5 # seconds to wait before starting delayed containers
|
||||
EXCLUDE_DIRS=() # directories to exclude from transfer — profiles override
|
||||
|
||||
# Default rsync options used when no profile match is found.
|
||||
# --delete removes files on remote that no longer exist on source (mirror behaviour)
|
||||
# --inplace writes directly to destination file — better for large files, avoids temp copies
|
||||
# --no-whole-file forces delta transfer even on fast local-like connections
|
||||
DEFAULT_RSYNC_OPTS=(-av --info=progress2 --human-readable --bwlimit="$BW_LIMIT" --delete --inplace --no-whole-file)
|
||||
|
||||
# ━━━ Remote Health Checks ━━━
|
||||
# Abort rsync if remote rootfs (/) usage is at or above this percentage.
|
||||
# When the remote array is down or drives are missing, rsync writes land on rootfs instead
|
||||
# of /mnt/user — this fills the filesystem rapidly and can crash the remote server.
|
||||
# 75% gives enough headroom to detect the problem before it becomes critical.
|
||||
# Pre-flight check run before every rsync — aborts if remote rootfs (/) usage is at or
|
||||
# above this percentage. When the remote array is down or drives are missing, rsync
|
||||
# writes land on rootfs instead of /mnt/user — this fills the filesystem rapidly and
|
||||
# can crash the remote server. 75% gives headroom to detect the problem early.
|
||||
ROOTFS_WARN=75
|
||||
|
||||
# ━━━ Daily Sync Shares ━━━
|
||||
# Media shares synced once daily by Orchestrators/daily_sync.sh
|
||||
# These shares have no profile — all use DEFAULT_RSYNC_OPTS above
|
||||
# Add or remove paths to control what gets synced each night
|
||||
# Media shares synced once daily by Orchestrators/daily_sync.sh.
|
||||
# These shares have no profile entry — all use DEFAULT_RSYNC_OPTS above.
|
||||
# Add or remove paths here to control what syncs each night.
|
||||
# For shares needing custom bandwidth or container stops — create a profile below instead.
|
||||
DAILY_SYNC_SHARES=(
|
||||
/mnt/user/Anime_Movies-Old
|
||||
/mnt/user/Anime_Shows-Old
|
||||
@@ -136,36 +162,38 @@ DAILY_SYNC_SHARES=(
|
||||
|
||||
# ━━━ Rsync Profile System ━━━
|
||||
# Profiles allow per-share rsync behaviour without touching script logic.
|
||||
# Each profile is matched automatically by the basename of the directory passed to rsync.sh
|
||||
# (lowercased). Example: /mnt/user/appdata-Failover/Arrs_Stack → profile key = arrs_stack
|
||||
# The profile key is matched automatically by the basename of the directory
|
||||
# passed to rsync.sh (lowercased).
|
||||
#
|
||||
# How fallthrough works:
|
||||
# If a key exists in a profile array → that value is used for this run
|
||||
# If a key is missing → the global default above is used instead
|
||||
# Shares in DAILY_SYNC_SHARES → always use globals, no profile defined
|
||||
# Example:
|
||||
# rsync.sh /mnt/user/appdata-Failover/Arrs_Stack
|
||||
# basename = Arrs_Stack → lowercased = arrs_stack → matches [arrs_stack] profile
|
||||
#
|
||||
# To add a new profile:
|
||||
# 1. Add a key to each array below with your chosen profile name
|
||||
# 1. Add a key to each PROFILE_* array below using your chosen name
|
||||
# 2. Call rsync.sh with a directory whose basename matches that key
|
||||
# 3. Any array you omit falls back to its global default automatically
|
||||
#
|
||||
# IMPORTANT: PROFILE_RSYNC_OPTS does NOT inherit from DEFAULT_RSYNC_OPTS —
|
||||
# if you define a profile entry you must list all desired options explicitly
|
||||
# IMPORTANT: PROFILE_RSYNC_OPTS does NOT inherit from DEFAULT_RSYNC_OPTS.
|
||||
# If you define a profile entry you must list ALL desired options explicitly.
|
||||
#
|
||||
# Profile descriptions:
|
||||
# arrs_stack — Sonarr, Radarr, Lidarr, Readarr, Prowlarr, Bazarr, Pinchflat appdata
|
||||
# Lower bandwidth — runs alongside media syncs, containers stopped during sync
|
||||
# critical-data — Auth stack: NPM, Authelia, Mariadb, Redis, LLDAP
|
||||
# High bandwidth — small data, synced frequently, Authelia needs delayed start
|
||||
# Current profiles:
|
||||
# arrs_stack — Sonarr, Radarr, Lidarr, Readarr, Prowlarr, Bazarr, Pinchflat
|
||||
# Lower bandwidth — runs alongside media syncs
|
||||
# Containers stopped during sync for data consistency
|
||||
# critical-data — Auth stack: NPM, Authelia, Mariadb-Authelia, Redis-Authelia, LLDAP
|
||||
# High bandwidth — small data, synced frequently
|
||||
# Authelia needs delayed start — database containers must be ready first
|
||||
# gmer4lfe — Server-specific appdata: Organizr, UptimeKuma, VaultWarden
|
||||
# Medium bandwidth — personal services, no container stop needed
|
||||
# important-data — NextCloud + Postgres database
|
||||
# High bandwidth — NextCloud needs graceful stop before sync
|
||||
# emby — Emby media server appdata and metadata
|
||||
# Medium bandwidth — large appdata, no containers stopped (metadata only)
|
||||
# NextCloud needs delayed start — Postgres must be accepting connections
|
||||
# emby — Emby media server appdata and metadata only
|
||||
# Medium bandwidth — large appdata directory, no containers stopped
|
||||
|
||||
# SPACE-SEPARATED STRINGS — rsync options per profile
|
||||
# If defined for a profile, these replace DEFAULT_RSYNC_OPTS entirely for that run
|
||||
# Rsync options per profile — replaces DEFAULT_RSYNC_OPTS entirely for that profile run
|
||||
# SPACE-SEPARATED STRINGS — converted to array at runtime by rsync.sh
|
||||
declare -A PROFILE_RSYNC_OPTS=(
|
||||
[arrs_stack]="-av --info=progress2 --human-readable --bwlimit=$BW_LIMIT --delete --inplace"
|
||||
[critical-data]="-av --human-readable --bwlimit=$BW_LIMIT --delete"
|
||||
@@ -174,16 +202,16 @@ declare -A PROFILE_RSYNC_OPTS=(
|
||||
[emby]="-av --human-readable --bwlimit=$BW_LIMIT"
|
||||
)
|
||||
|
||||
# Bandwidth limit in KB/s per profile — overrides global BW_LIMIT for this profile
|
||||
# Bandwidth limit in KB/s per profile — overrides global BW_LIMIT for this profile only
|
||||
declare -A PROFILE_BW_LIMIT=(
|
||||
[arrs_stack]=5000 # lower — runs alongside other jobs
|
||||
[critical-data]=9500 # high — small data, sync fast
|
||||
[arrs_stack]=5000 # lower — runs alongside other jobs, avoids saturating link
|
||||
[critical-data]=9500 # high — small data, get it synced fast
|
||||
[gmer4lfe]=8000 # medium
|
||||
[important-data]=9500 # high — database sync, prioritise speed
|
||||
[emby]=8000 # medium — large files, steady transfer
|
||||
[important-data]=9500 # high — database sync needs to be fast and clean
|
||||
[emby]=8000 # medium — large files, steady sustained transfer
|
||||
)
|
||||
|
||||
# Retry attempts per profile — overrides global RETRY_COUNT
|
||||
# Retry attempts per profile before giving up — overrides global RETRY_COUNT
|
||||
declare -A PROFILE_RETRY_COUNT=(
|
||||
[arrs_stack]=3
|
||||
[critical-data]=3
|
||||
@@ -192,7 +220,7 @@ declare -A PROFILE_RETRY_COUNT=(
|
||||
[emby]=3
|
||||
)
|
||||
|
||||
# Sleep between retries in seconds — overrides global SLEEP
|
||||
# Seconds between retry attempts — overrides global SLEEP
|
||||
declare -A PROFILE_SLEEP=(
|
||||
[arrs_stack]=300
|
||||
[critical-data]=300
|
||||
@@ -201,9 +229,10 @@ declare -A PROFILE_SLEEP=(
|
||||
[emby]=300
|
||||
)
|
||||
|
||||
# Containers stopped on REMOTE before rsync and restarted after — SPACE-SEPARATED STRINGS
|
||||
# Only containers that need to be stopped for data consistency — running containers are fine
|
||||
# for most media but databases and auth stacks need clean state during sync
|
||||
# Containers stopped on REMOTE before rsync and restarted after completion.
|
||||
# Only include containers that need to be stopped for data consistency.
|
||||
# Databases and auth stacks need clean state — media servers generally do not.
|
||||
# SPACE-SEPARATED STRINGS — converted to array at runtime
|
||||
declare -A PROFILE_CRITICAL_CONTAINER_NAMES=(
|
||||
[arrs_stack]="Sonarr Lidarr Readarr Radarr Prowlarr Bazarr Pinchflat"
|
||||
[critical-data]="Mariadb-Authelia Redis-Authelia Lldap-Gmer4Lfe NginxProxyManager Authelia"
|
||||
@@ -212,28 +241,30 @@ declare -A PROFILE_CRITICAL_CONTAINER_NAMES=(
|
||||
[emby]=""
|
||||
)
|
||||
|
||||
# Containers that need a delay before starting after rsync — SPACE-SEPARATED STRINGS
|
||||
# Authelia needs its database containers (Mariadb, Redis) to be ready before it starts
|
||||
# NextCloud needs Postgres to be accepting connections before it starts
|
||||
# Containers that need a delay before starting after rsync completes.
|
||||
# Used when a container depends on another that was also stopped — it needs its
|
||||
# dependency to be ready before it can start successfully.
|
||||
# SPACE-SEPARATED STRINGS — converted to array at runtime
|
||||
declare -A PROFILE_DELAYED_CONTAINERS=(
|
||||
[arrs_stack]=""
|
||||
[critical-data]="Authelia"
|
||||
[critical-data]="Authelia" # Authelia needs Mariadb + Redis ready before starting
|
||||
[gmer4lfe]=""
|
||||
[important-data]="NextCloud"
|
||||
[important-data]="NextCloud" # NextCloud needs Postgres accepting connections first
|
||||
[emby]=""
|
||||
)
|
||||
|
||||
# Seconds to wait before starting delayed containers — per profile
|
||||
# Seconds to wait before starting delayed containers — gives dependencies time to initialise
|
||||
declare -A PROFILE_CONTAINER_DELAY=(
|
||||
[arrs_stack]=5
|
||||
[critical-data]=10 # Authelia needs DB ready — 10s gives Mariadb/Redis time to start
|
||||
[critical-data]=10 # 10s gives Mariadb and Redis time to accept connections
|
||||
[gmer4lfe]=5
|
||||
[important-data]=10 # NextCloud needs Postgres ready
|
||||
[important-data]=10 # 10s gives Postgres time to accept connections
|
||||
[emby]=5
|
||||
)
|
||||
|
||||
# Directories excluded from transfer per profile — SPACE-SEPARATED STRINGS
|
||||
# logs and *.tmp are excluded universally — they are ephemeral and regenerated on start
|
||||
# Directories excluded from rsync transfer per profile.
|
||||
# logs and *.tmp are safe to exclude — they are ephemeral and regenerated on container start.
|
||||
# SPACE-SEPARATED STRINGS — converted to array at runtime
|
||||
declare -A PROFILE_EXCLUDE_DIRS=(
|
||||
[arrs_stack]="logs *.tmp"
|
||||
[critical-data]="logs *.tmp"
|
||||
@@ -242,75 +273,78 @@ declare -A PROFILE_EXCLUDE_DIRS=(
|
||||
[emby]="logs *.tmp"
|
||||
)
|
||||
|
||||
# ━━━ Profile Disk Check Toggle ━━━
|
||||
# true = skip per-disk check for this profile — use when share lives on a ZFS pool
|
||||
# false = run per-disk check — use for traditional unRAID array with individual disks
|
||||
# Per-disk check looks for /mnt/disk*/sharename — ZFS pools don't have this structure
|
||||
# The share existence and content checks still run regardless of this setting
|
||||
# Per-disk check toggle — controls whether rsync.sh runs check_remote_disks() for this profile.
|
||||
# true = skip per-disk check — use when remote share lives on a ZFS pool
|
||||
# ZFS pools don't have /mnt/disk* structure so the check always fails incorrectly
|
||||
# false = run per-disk check — use for traditional unRAID array with individual disk mounts
|
||||
# Verifies all disks backing the share are online before syncing
|
||||
# Note: rootfs and share existence checks always run regardless of this setting
|
||||
declare -A PROFILE_SKIP_DISK_CHECK=(
|
||||
[arrs_stack]=true
|
||||
[critical-data]=true
|
||||
[gmer4lfe]=true
|
||||
[important-data]=true
|
||||
[emby]=true
|
||||
[arrs_stack]=true # remote uses ZFS pool — no individual disk mounts
|
||||
[critical-data]=true
|
||||
[gmer4lfe]=true
|
||||
[important-data]=true
|
||||
[emby]=true
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FAILOVER ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Mutual container failover between two unRAID servers 50 miles apart.
|
||||
# Each server runs failover.sh independently — no coordination between servers.
|
||||
# Each server runs Failover/failover.sh independently — no coordination between servers.
|
||||
# All decisions are based solely on two ping checks: remote reachable + internet reachable.
|
||||
#
|
||||
# How it works:
|
||||
# Every FAILOVER_CHECK_INTERVAL seconds each server pings the remote and pings the internet.
|
||||
# Based on those two results it determines its state and takes the appropriate action.
|
||||
# No SSH, no signaling — each server acts autonomously based only on what it can see.
|
||||
#
|
||||
# States:
|
||||
# NORMAL — remote up, internet up — own containers only, silent operation
|
||||
# FAILOVER — remote down, internet up — start remote's containers locally (additive)
|
||||
# NO_INTERNET — internet down — stop public-facing containers, wait
|
||||
# DARK — remote down + internet down — same as NO_INTERNET
|
||||
# Own normal containers keep running — failover containers added on top
|
||||
# NO_INTERNET — internet down — stop public-facing containers, wait for recovery
|
||||
# DARK — remote down + internet down — same actions as NO_INTERNET
|
||||
#
|
||||
# Handback sequence when remote returns after FAILOVER:
|
||||
# Strike confirmation (FAILOVER_HANDBACK_STRIKES consecutive remote-up checks)
|
||||
# → pre-flight checks (remote array, docker daemon, rootfs)
|
||||
# → rsync data back via rsync.sh (uses profile system)
|
||||
# → start failover containers on remote via SSH
|
||||
# → stop failover containers locally
|
||||
# → return to NORMAL
|
||||
# 1. Strike confirmation — FAILOVER_HANDBACK_STRIKES consecutive remote-up checks
|
||||
# Prevents handing back during a brief network blip
|
||||
# 2. Pre-flight checks — remote array started, Docker daemon healthy, rootfs not full
|
||||
# 3. Rsync data back via rsync.sh — uses existing profile system for options
|
||||
# 4. Start failover containers on remote via SSH
|
||||
# 5. Stop failover containers locally — only after remote confirmed started
|
||||
# 6. Return to NORMAL state
|
||||
#
|
||||
# Comment out any container or rsync job to disable without removing the entry.
|
||||
# Script runs on BOTH servers — detect_hosts() selects the correct arrays automatically.
|
||||
# Comment out any container or rsync job to disable without removing the entry.
|
||||
|
||||
# External IP to ping for internet connectivity check — Google DNS, reliable and fast
|
||||
EXTERNAL_IP="8.8.8.8"
|
||||
# Seconds between state checks — 120s = 2 minute polling interval
|
||||
# With 1 minute DNS TTL this means failover is visible to users within ~3 minutes
|
||||
FAILOVER_CHECK_INTERVAL=120
|
||||
# Consecutive remote-up confirmations required before initiating handback
|
||||
# Prevents handing back during a brief network blip — 2 strikes = 4 minutes confirmation
|
||||
FAILOVER_HANDBACK_STRIKES=2
|
||||
# State file path — on /boot/ so it survives reboots
|
||||
# Script re-evaluates from scratch on restart using live pings — state file is reference only
|
||||
EXTERNAL_IP="8.8.8.8" # external IP to ping for internet connectivity check
|
||||
FAILOVER_CHECK_INTERVAL=120 # seconds between state checks — 120s = 2 minute polling
|
||||
FAILOVER_HANDBACK_STRIKES=2 # consecutive remote-up confirmations required before handback
|
||||
# 2 strikes at 120s interval = 4 minutes confirmation window
|
||||
FAILOVER_STATE_FILE="/boot/config/failover_state.db"
|
||||
# persists on /boot/ so it survives reboots
|
||||
# script re-evaluates from live pings on restart
|
||||
|
||||
# ━━━ HOST1 Failover Config (unRAID-Gmer4Lfe — Primary) ━━━
|
||||
|
||||
# Containers HOST1 starts locally when HOST2 goes down
|
||||
# These run on top of HOST1's normal containers — additive, not replacement
|
||||
# HOST2's DDNS containers must start here so DNS stays pointing at HOST1 during HOST2 outage
|
||||
# Containers HOST1 starts locally when HOST2 goes down.
|
||||
# These run ON TOP OF HOST1's normal containers — additive, not a replacement.
|
||||
FAILOVER_HOST1_STARTS_FOR_HOST2=(
|
||||
"Gmer4Lfe.com"
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# Containers HOST1 stops when it loses internet
|
||||
# No point serving DDNS or media if HOST1 itself has no internet — stops unnecessary churn
|
||||
# HOST2 will independently detect HOST1 is gone and start its own failover
|
||||
# Containers HOST1 stops when it loses internet connectivity.
|
||||
# No point serving DDNS or public services if HOST1 itself has no internet.
|
||||
FAILOVER_HOST1_STOP_ON_NO_NET=(
|
||||
"Gmer4Lfe.com"
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# Rsync jobs HOST1 runs before handing containers back to HOST2
|
||||
# Uses rsync.sh with existing profile system — basename matched to profile keys
|
||||
# Comment out jobs that don't need syncing back or aren't ready yet
|
||||
# Rsync jobs HOST1 runs before handing containers back to HOST2 after recovery.
|
||||
# Uses rsync.sh with the existing profile system — basename matched to profile keys.
|
||||
# Comment out jobs that are not yet ready or not needed for handback.
|
||||
FAILOVER_HOST1_RSYNC_JOBS=(
|
||||
# "/mnt/user/appdata-Failover/Jayred365"
|
||||
# "/mnt/user/Media_Server/Emby-Jayred"
|
||||
@@ -318,24 +352,22 @@ FAILOVER_HOST1_RSYNC_JOBS=(
|
||||
|
||||
# ━━━ HOST2 Failover Config (unRAID-Jayred365 — Secondary) ━━━
|
||||
|
||||
# Containers HOST2 starts locally when HOST1 goes down
|
||||
# Emby starts on HOST2 so media keeps working during HOST1 outage
|
||||
# HOST1's DDNS containers start here so DNS updates to point at HOST2
|
||||
# Containers HOST2 starts locally when HOST1 goes down.
|
||||
# Emby starts on HOST2 so media keeps working during HOST1 outage.
|
||||
# HOST1's DDNS containers start here so DNS updates to point at HOST2's IP.
|
||||
FAILOVER_HOST2_STARTS_FOR_HOST1=(
|
||||
"Emby"
|
||||
"Gmer4Lfe.com"
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# Containers HOST2 stops when it loses internet
|
||||
# HOST2's DDNS containers serve no purpose without internet connectivity
|
||||
# Containers HOST2 stops when it loses internet connectivity.
|
||||
FAILOVER_HOST2_STOP_ON_NO_NET=(
|
||||
"Gmer4Lfe.com"
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# Rsync jobs HOST2 runs before handing containers back to HOST1
|
||||
# Comment out jobs that don't need syncing back or aren't ready yet
|
||||
# Rsync jobs HOST2 runs before handing containers back to HOST1 after recovery.
|
||||
FAILOVER_HOST2_RSYNC_JOBS=(
|
||||
# "/mnt/user/appdata-Failover/Gmer4Lfe"
|
||||
)
|
||||
@@ -345,8 +377,9 @@ FAILOVER_HOST2_RSYNC_JOBS=(
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Docker Daily Restart ━━━
|
||||
# Containers restarted every day — keeps services fresh, clears memory leaks
|
||||
# Case-sensitive — must match exact Docker container names
|
||||
# Containers restarted every day by Docker_Essentials/docker_daily_restart.sh.
|
||||
# Keeps services fresh and clears memory leaks that accumulate over time.
|
||||
# Case-sensitive — must match exact Docker container names shown in the unRAID Docker tab.
|
||||
DAILY_RESTART_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
"Authelia"
|
||||
@@ -357,7 +390,8 @@ DAILY_RESTART_CONTAINERS=(
|
||||
)
|
||||
|
||||
# ━━━ Docker Weekly Restart ━━━
|
||||
# Containers restarted once per week — less critical services that benefit from periodic restart
|
||||
# Containers restarted once per week by Docker_Essentials/docker_weekly_restart.sh.
|
||||
# For less critical services that benefit from periodic restart but don't need daily cycling.
|
||||
WEEKLY_RESTART_CONTAINERS=(
|
||||
"NextCloud"
|
||||
"Organizrv2-Gmer4Lfe"
|
||||
@@ -366,30 +400,33 @@ WEEKLY_RESTART_CONTAINERS=(
|
||||
)
|
||||
|
||||
# ━━━ Docker Watchdog ━━━
|
||||
# First line of defense — monitors containers for memory, CPU and HTTP responsiveness.
|
||||
# Restarts containers that exceed thresholds using a strike system to avoid false positives.
|
||||
# Works alongside system_watchdog.sh — containers first, system reboot second.
|
||||
# First line of defense for container health — runs every 15 minutes via cron.
|
||||
# Monitors memory usage, CPU usage and HTTP responsiveness per container.
|
||||
# Uses a strike system to avoid restarting on brief spikes — sustained issues trigger restart.
|
||||
# Works alongside system_watchdog.sh — containers first, system reboot is the last resort.
|
||||
|
||||
# Containers to monitor with their memory hard limits in MB
|
||||
# Strike system used for CPU and responsiveness — immediate restart for memory hard limit
|
||||
# Containers to monitor with their memory hard limits in MB.
|
||||
# Memory hard limit exceeded → immediate restart (no strike system for memory).
|
||||
# CPU and HTTP use strike system — see CPU_FAIL_LIMIT and RESP_FAIL_LIMIT below.
|
||||
# 20GB=20480 16GB=16384 14GB=14336 12GB=12288 10GB=10240
|
||||
# 8GB=8192 6GB=6144 4GB=4096 2GB=2048 1GB=1024
|
||||
declare -A WATCHDOG_CONTAINERS=(
|
||||
["Emby"]=16384 # 16GB — large media server, transcoding can spike
|
||||
["Emby"]=16384 # 16GB — media server, transcoding can spike high
|
||||
["LidaTube"]=6144 # 6GB — YouTube downloader
|
||||
["Tdarr"]=6144 # 6GB — transcoding node
|
||||
["Code-Server"]=1024 # 1GB — VS Code server
|
||||
)
|
||||
|
||||
# Containers to check HTTP responsiveness — curl check per interval
|
||||
# Omit a container entirely to skip its HTTP check
|
||||
# Containers to check HTTP responsiveness via curl — omit a container to skip its HTTP check.
|
||||
# curl checks the URL and considers the container unresponsive if it times out or errors.
|
||||
declare -A WATCHDOG_CONTAINER_URLS=(
|
||||
["Emby"]="http://localhost:8096"
|
||||
)
|
||||
|
||||
# Containers that should always be running — monitored for unexpected stops
|
||||
# Strike system used — persistent skip list on /boot/ prevents reboot loops
|
||||
# Auto-clears from skip list when container recovers after reboot or manual fix
|
||||
# Containers that should always be running — monitored for unexpected stops.
|
||||
# Strike system used — tries restart on each strike up to SYS_WATCHDOG_STRIKE_LIMIT.
|
||||
# If restart fails after strike limit → added to persistent skip list on /boot/
|
||||
# Skip list auto-clears when container recovers after reboot or manual fix.
|
||||
WATCHDOG_REQUIRED_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
"Lldap-Gmer4Lfe"
|
||||
@@ -400,33 +437,36 @@ WATCHDOG_REQUIRED_CONTAINERS=(
|
||||
"Redis-Authelia-Secondary"
|
||||
)
|
||||
|
||||
# Strike state file — /tmp resets on reboot which is correct for strike tracking
|
||||
# Strike state file — /tmp resets on reboot which is correct behaviour for strike tracking
|
||||
WATCHDOG_STATE_FILE="/tmp/container_watchdog_state.db"
|
||||
|
||||
# CPU thresholds — normalised against total core count at runtime
|
||||
# CPU thresholds — normalised against total core count automatically at runtime.
|
||||
# A container using 85% of one core on a 16-core system = ~5.3% normalised.
|
||||
SOFT_CPU_THRESHOLD=80 # warn at this % of total system CPU
|
||||
HARD_CPU_THRESHOLD=85 # strike at this % of total system CPU
|
||||
CPU_FAIL_LIMIT=2 # consecutive strikes before container restart
|
||||
|
||||
# Memory — warn at this % of per-container hard limit defined in WATCHDOG_CONTAINERS
|
||||
# Memory soft threshold — warn when container reaches this % of its hard limit.
|
||||
# Hard limit exceeded triggers immediate restart regardless of strikes.
|
||||
SOFT_MEM_THRESHOLD=80
|
||||
|
||||
# HTTP responsiveness check
|
||||
# HTTP responsiveness check settings
|
||||
RESP_FAIL_LIMIT=2 # consecutive failed curl checks before restart
|
||||
CURL_TIMEOUT=5 # seconds before curl gives up per check
|
||||
|
||||
# ━━━ Docker Network Connect ━━━
|
||||
# Connects containers to extra Docker networks on array start — many-to-many
|
||||
# Every container in the list connects to every network in the list
|
||||
# Useful when containers need to communicate across networks they weren't configured with
|
||||
# Comment out entries to disable without removing them
|
||||
# Connects containers to extra Docker networks on array start.
|
||||
# Useful when containers need to communicate across networks they were not originally
|
||||
# configured with — e.g. memcached needing access to the nextcloud-aio network.
|
||||
# Every container in the list connects to every network in the list (many-to-many).
|
||||
# Comment out entries to disable without removing them.
|
||||
NETWORK_CONNECT_CONTAINERS=(
|
||||
"memcached"
|
||||
"Npm-CrowdSec"
|
||||
)
|
||||
|
||||
NETWORK_CONNECT_NETWORKS=(
|
||||
"nextcloud-aio"
|
||||
"nextcloud-aio" # Docker network name — must exist before array start
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -434,52 +474,56 @@ NETWORK_CONNECT_NETWORKS=(
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Reboot ━━━
|
||||
# Seconds of warning broadcast to logged-in users before server_reboot.sh reboots
|
||||
# Seconds of warning broadcast to all logged-in users before server_reboot.sh reboots.
|
||||
# Gives users time to save work or finish what they are doing before the system goes down.
|
||||
REBOOT_SLEEP=300
|
||||
|
||||
# ━━━ Mover ━━━
|
||||
# Seconds to wait after warning users before mover_stop.sh kills the mover process
|
||||
# Seconds to wait after warning users before mover_stop.sh sends SIGTERM to the mover.
|
||||
# Gives the mover time to finish its current file operation cleanly before being killed.
|
||||
MOVER_STOP_TIMEOUT=300
|
||||
|
||||
# ━━━ Syslog Filter ━━━
|
||||
# Path for the rsyslog filter file that suppresses Docker veth/docker0 noise
|
||||
# Path for the rsyslog filter file created by docker_syslog_filter.sh.
|
||||
# The filter suppresses noisy Docker veth and docker0 messages from syslog on boot.
|
||||
# Without this filter, every Docker network interface change floods the syslog.
|
||||
FILTER_FILE="/etc/rsyslog.d/ignore-docker-veth.conf"
|
||||
|
||||
# ━━━ PHP-FPM ━━━
|
||||
# Config file path and max children value for php_fpm_max_children.sh
|
||||
# Config file path and max children value for php_fpm_max_children.sh.
|
||||
# Higher max_children allows more concurrent PHP requests to the unRAID WebGUI.
|
||||
# Set based on available RAM — too high can cause memory pressure on low-RAM systems.
|
||||
PHP_CONF="/etc/php-fpm.d/www.conf"
|
||||
PHP_MAX_CHILDREN=250
|
||||
|
||||
# ━━━ Clear Logs ━━━
|
||||
# System log files cleared by clear_logs.sh — Docker logs cleared automatically too
|
||||
# System log files cleared by clear_logs.sh — Docker container logs are cleared too.
|
||||
# Run weekly to prevent logs from filling the rootfs over time.
|
||||
LOG_FILES=(/var/log/syslog /var/log/messages /var/log/dmesg)
|
||||
|
||||
# ━━━ WebGUI Watchdog ━━━
|
||||
# Monitors unRAID WebGUI and restarts services if unresponsive
|
||||
# Escalation: nginx restart → recheck → emhttp restart → recheck → notify warning
|
||||
# Monitors the unRAID WebGUI and restarts services if it becomes unresponsive.
|
||||
# Escalation path: nginx restart → recheck → emhttp restart → recheck → notify warning.
|
||||
# emhttp is the core unRAID management daemon — restarting it is more disruptive than nginx
|
||||
# but both recover cleanly. Notification sent on any restart so you know what happened.
|
||||
WEBGUI_URL="http://localhost" # adjust port if non-standard e.g. http://localhost:8080
|
||||
WEBGUI_TIMEOUT=5 # seconds before curl gives up
|
||||
WEBGUI_NGINX_WAIT=15 # seconds to wait after nginx restart before recheck
|
||||
WEBGUI_EMHTTP_WAIT=30 # seconds to wait after emhttp restart before recheck
|
||||
|
||||
# ━━━ ZFS Memory Snapshot ━━━
|
||||
# Weekly ZFS pool health and memory diagnostic report — informational only, no action taken
|
||||
# system_watchdog.sh handles threshold-based intervention
|
||||
ZFS_REPORT_LOG="/var/log/zfs-weekly-health.log"
|
||||
ZFS_REPORT_ARC_WARN_PCT=90 # warn if ARC utilization above this %
|
||||
ZFS_REPORT_FREE_WARN_GB=10 # warn if free RAM below this GB
|
||||
ZFS_REPORT_AVAIL_WARN_GB=20 # warn if available RAM below this GB
|
||||
ZFS_REPORT_DOCKER_TOP=10 # number of top Docker memory users to show in report
|
||||
WEBGUI_TIMEOUT=5 # seconds before curl gives up on the WebGUI check
|
||||
WEBGUI_NGINX_WAIT=15 # seconds to wait after nginx restart before rechecking
|
||||
WEBGUI_EMHTTP_WAIT=30 # seconds to wait after emhttp restart — emhttp takes longer
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MEDIA ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Media Permissions ━━━
|
||||
# Mode and owner applied recursively to all shares in MEDIA_PERMISSION_SHARES
|
||||
# Mode and owner applied recursively to all shares in MEDIA_PERMISSION_SHARES.
|
||||
# Run by Media/media_shares_permissions.sh via the media_management.sh orchestrator.
|
||||
# 777 and nobody:users is standard for unRAID media shares accessible by Docker containers.
|
||||
PERMISSIONS_MODE="777"
|
||||
PERMISSIONS_OWNER="nobody:users"
|
||||
|
||||
# Shares to apply permissions to — add or remove paths as your library grows.
|
||||
# These are applied recursively so large shares take time — run overnight via orchestrator.
|
||||
MEDIA_PERMISSION_SHARES=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Movies-Old
|
||||
@@ -506,10 +550,12 @@ MEDIA_PERMISSION_SHARES=(
|
||||
)
|
||||
|
||||
# ━━━ Media Cleaner ━━━
|
||||
# Two profiles: anime and media — passed as argument to media_cleaner.sh
|
||||
# Run via Orchestrators/media_management.sh for permissions + both cleaners in one job
|
||||
# Usage: media_cleaner.sh anime or media_cleaner.sh media
|
||||
# Removes junk files from media shares using configurable file pattern lists.
|
||||
# Two profiles: anime and media — each with their own folder list and patterns.
|
||||
# Run via Media/media_cleaner.sh anime or Media/media_cleaner.sh media
|
||||
# Called automatically by media_management.sh via MEDIA_MAINTENANCE_JOBS below.
|
||||
|
||||
# Folders scanned by the anime profile — anime downloads commonly include these junk files
|
||||
ANIME_CLEAN_FOLDERS=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Movies-Old
|
||||
@@ -517,6 +563,7 @@ ANIME_CLEAN_FOLDERS=(
|
||||
/mnt/user/Anime_Shows-Old
|
||||
)
|
||||
|
||||
# Folders scanned by the media profile
|
||||
MEDIA_CLEAN_FOLDERS=(
|
||||
/mnt/user/Kids_Movies
|
||||
/mnt/user/Kids_Tv_Shows
|
||||
@@ -527,7 +574,7 @@ MEDIA_CLEAN_FOLDERS=(
|
||||
/mnt/user/Tv_Shows
|
||||
)
|
||||
|
||||
# Junk file patterns common in anime downloads
|
||||
# File patterns deleted by the anime profile — common junk from anime download groups
|
||||
ANIME_FILE_PATTERNS=(
|
||||
'*.sfv' '*.md5' '*.sha1' '*.txt' '*.url' '*.lnk'
|
||||
'*.rar' '*.zip' '*.info' '*.torrent' '*.sample*' '*.proof*'
|
||||
@@ -535,7 +582,7 @@ ANIME_FILE_PATTERNS=(
|
||||
'*.log' '*.json'
|
||||
)
|
||||
|
||||
# Junk file patterns for general media — includes *.iso and *.lrc not needed in anime
|
||||
# File patterns deleted by the media profile — includes *.iso and *.lrc not needed in anime
|
||||
MEDIA_FILE_PATTERNS=(
|
||||
'*.sfv' '*.md5' '*.sha1' '*.txt' '*.url' '*.lnk'
|
||||
'*.rar' '*.zip' '*.info' '*.torrent' '*.sample*' '*.proof*'
|
||||
@@ -544,108 +591,252 @@ MEDIA_FILE_PATTERNS=(
|
||||
)
|
||||
|
||||
# ━━━ Media Management Orchestrator ━━━
|
||||
# Scripts run sequentially by Orchestrators/media_management.sh
|
||||
# Order matters — permissions runs first so cleaner operates on correctly owned files
|
||||
# Format: "script_path [optional_argument]"
|
||||
# Add a new line to run another script — comment out to disable without removing
|
||||
# Job list for Orchestrators/media_management.sh — runs scripts sequentially in order.
|
||||
# Format: "folder/script.sh optional_argument"
|
||||
# Order matters — permissions runs first so cleaners and arr scripts see correct ownership.
|
||||
# Arr cleanup scripts run last — they depend on clean folders from the cleaner steps.
|
||||
# Comment out any job to disable without removing it — easy to re-enable later.
|
||||
MEDIA_MAINTENANCE_JOBS=(
|
||||
"Media/media_shares_permissions.sh"
|
||||
"Media/media_cleaner.sh anime"
|
||||
"Media/media_cleaner.sh media"
|
||||
"Media/media_shares_permissions.sh" # apply permissions first
|
||||
"Media/media_cleaner.sh anime" # remove junk from anime shares
|
||||
"Media/media_cleaner.sh media" # remove junk from media shares
|
||||
"Media/lidarr_cleanup.sh" # remove orphaned music files
|
||||
"Media/sonarr_cleanup.sh" # remove orphaned TV files
|
||||
"Media/radarr_cleanup.sh" # remove orphaned movie files
|
||||
)
|
||||
|
||||
# ━━━ Arr Cleanup ━━━
|
||||
# Lidarr, Sonarr and Radarr orphan file cleanup via their respective APIs.
|
||||
# Each arr script queries its API to get all tracked file paths, then compares against
|
||||
# what exists on disk. Files not tracked by the arr and older than ORPHAN_AGE days are deleted.
|
||||
#
|
||||
# Why the age threshold matters:
|
||||
# The arr downloads a file then processes it — there is a window where the file exists
|
||||
# on disk but the arr hasn't imported it yet. ORPHAN_AGE prevents deleting files that
|
||||
# are mid-import. 7 days is conservative and safe for any normal workflow.
|
||||
#
|
||||
# Protected patterns are NEVER deleted regardless of tracking status or age.
|
||||
# These protect arr-generated metadata (cover art, .nfo files, subtitles) that the arr
|
||||
# depends on but does not include in its tracked file API response.
|
||||
# Add new patterns here if arr metadata formats change in future versions.
|
||||
|
||||
# Lidarr — music library
|
||||
LIDARR_URL="http://192.168.50.2:8686"
|
||||
LIDARR_API_KEY="b2977e71ef074bc0a0529d9fcce3b2dc"
|
||||
LIDARR_MUSIC_ROOT="/mnt/user/Music-New" # must match the root path set in Lidarr
|
||||
LIDARR_ORPHAN_AGE=7 # days before untracked file is eligible for deletion
|
||||
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
||||
# file extensions considered valid music files
|
||||
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
||||
# never deleted — cover art, metadata, lyrics
|
||||
|
||||
# Sonarr — TV library
|
||||
SONARR_URL="http://192.168.50.2:8989"
|
||||
SONARR_API_KEY="130decd3db5b4c25afad64864cd03f9f"
|
||||
SONARR_TV_ROOT="/mnt/user/Tv_Shows" # must match the root path set in Sonarr
|
||||
SONARR_ORPHAN_AGE=7
|
||||
SONARR_EXTENSIONS=("mkv" "mp4" "avi" "m4v" "ts" "wmv" "mov")
|
||||
SONARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.ass" "*.ssa")
|
||||
# never deleted — artwork, metadata, subtitles
|
||||
|
||||
# Radarr — movie library
|
||||
RADARR_URL="http://192.168.50.2:7878"
|
||||
RADARR_API_KEY="d43a3ec6cf1549edb4af0cc63f98b2a9"
|
||||
RADARR_MOVIES_ROOT="/mnt/user/Movies" # must match the root path set in Radarr
|
||||
RADARR_ORPHAN_AGE=7
|
||||
RADARR_EXTENSIONS=("mkv" "mp4" "avi" "m4v" "wmv" "mov")
|
||||
RADARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.ass" "*.ssa")
|
||||
|
||||
# ==============================================================================================
|
||||
# ── TRANSCODES ────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Session-based storage allocator using filesystem indirection.
|
||||
# ffmpeg resolves the symlink path once at session start — existing sessions are never affected.
|
||||
# New sessions land wherever TRANSCODE_LINK points at the moment they start.
|
||||
# ffmpeg resolves the symlink path ONCE at session start — existing sessions are never affected
|
||||
# by symlink changes. Only NEW sessions care about where the symlink currently points.
|
||||
#
|
||||
# Flow:
|
||||
# ramdisk_setup.sh — run once at array start, creates ramdisk and sets symlink
|
||||
# transcode_manager.sh — every 3 min, monitors usage and flips symlink if needed
|
||||
# How it works:
|
||||
# ramdisk_setup.sh — run once at array start, creates tmpfs and sets symlink
|
||||
# transcode_manager.sh — every 3 min, monitors ramdisk usage and flips symlink if needed
|
||||
# transcode_cleanup.sh — every 5 min, removes old inactive files from both locations
|
||||
#
|
||||
# Hysteresis gap between RAMDISK_WARN_GB and RAMDISK_LOW_GB prevents flip-flop
|
||||
# when usage hovers near the threshold — gap should be at least 0.5-1GB
|
||||
# when usage hovers near the threshold. Gap should be at least 0.5-1GB.
|
||||
# Ramdisk is dynamic tmpfs — only uses RAM actually needed, RAMDISK_SIZE is the ceiling.
|
||||
|
||||
RAMDISK_PATH="/mnt/ramdisk_transcodes" # tmpfs mount point — created at array start
|
||||
RAMDISK_SIZE="8G" # ceiling — tmpfs is dynamic, only uses what's needed
|
||||
TRANSCODE_LINK="/mnt/ram-transcode" # symlink Emby points at — never changes location
|
||||
TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/" # SSD fallback for edge cases
|
||||
RAMDISK_PATH="/mnt/ramdisk_transcodes" # tmpfs mount point created at array start
|
||||
RAMDISK_SIZE="8G" # ceiling — tmpfs only uses RAM actually needed
|
||||
TRANSCODE_LINK="/mnt/ram-transcode" # symlink Emby points at — location never changes
|
||||
TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/" # SSD fallback for edge cases only
|
||||
|
||||
# Thresholds in GB — flip symlink to SSD at WARN, flip back to ramdisk at LOW
|
||||
RAMDISK_WARN_GB=6.8 # flip to SSD at or above this — "getting full"
|
||||
RAMDISK_LOW_GB=5.5 # flip back to ramdisk when cleanup brings usage here
|
||||
RAMDISK_SSD_MIN_GB=20 # minimum free GB on SSD before allowing flip — safety net
|
||||
# Usage thresholds in GB
|
||||
RAMDISK_WARN_GB=6.8 # flip symlink to SSD at or above this usage
|
||||
RAMDISK_LOW_GB=5.5 # flip symlink back to ramdisk when usage drops here
|
||||
RAMDISK_SSD_MIN_GB=20 # minimum free GB on SSD required before allowing flip to SSD
|
||||
|
||||
# Cleanup — files must be older than MAX_AGE and not open by any process to be deleted
|
||||
TRANSCODE_MAX_AGE=20 # minutes before a file is eligible for cleanup
|
||||
TRANSCODE_ORPHAN_AGE=30 # minutes before an orphaned file is eligible — extra caution
|
||||
# Cleanup age thresholds — files must be older than these AND not open by any process
|
||||
TRANSCODE_MAX_AGE=20 # minutes before a transcode file is eligible for cleanup
|
||||
TRANSCODE_ORPHAN_AGE=30 # minutes before an orphaned file is eligible — extra caution buffer
|
||||
|
||||
# Flip frequency monitoring — notify if symlink flips too often (indicates sizing issue)
|
||||
# Flip frequency alert — too many flips per hour may indicate ramdisk needs to be larger
|
||||
TRANSCODE_FLIP_WARN=3 # notify if symlink flips this many times in one hour
|
||||
|
||||
# Permissions — must match your Emby container user
|
||||
# Permissions applied to ramdisk and SSD fallback — must match your Emby container user
|
||||
TRANSCODE_OWNER="nobody:users"
|
||||
TRANSCODE_MODE="755"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MONITOR ───────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Monitoring scripts — watch and report only, never take action.
|
||||
# Lives in Monitor/ folder — distinct from unRAID_Essentials (which acts) and
|
||||
# Docker_Essentials (which manages containers).
|
||||
# These scripts are the data sources for the future plugin dashboard.
|
||||
|
||||
# ━━━ Certificate Monitor ━━━
|
||||
# Checks SSL cert expiry via direct openssl connection — no NPM dependency.
|
||||
# Reads the actual cert the server is presenting — catches real-world issues API checks miss.
|
||||
# Each domain and subdomain is a separate entry — they have independent certs.
|
||||
# Add your public-facing domains — uncomment and replace with your actual domains.
|
||||
CERT_MONITOR_DOMAINS=(
|
||||
# "yourdomain.com"
|
||||
# "auth.yourdomain.com"
|
||||
# "emby.yourdomain.com"
|
||||
# "nextcloud.yourdomain.com"
|
||||
)
|
||||
CERT_WARN_DAYS=30 # notify warning when cert expires within this many days
|
||||
CERT_CRIT_DAYS=7 # notify critical when cert expires within this many days
|
||||
CERT_TIMEOUT=10 # seconds before openssl connection attempt gives up per domain
|
||||
|
||||
# ━━━ Backup Verify ━━━
|
||||
# Verifies the rsync mirror is healthy by comparing random file checksums between servers.
|
||||
# Uses existing SSH keys — no additional configuration needed beyond the share list.
|
||||
# Leave BACKUP_VERIFY_SHARES empty to automatically use DAILY_SYNC_SHARES as the target list.
|
||||
BACKUP_VERIFY_SHARES=(
|
||||
# leave empty to use DAILY_SYNC_SHARES automatically
|
||||
# or specify individual shares to verify:
|
||||
# /mnt/user/Movies
|
||||
# /mnt/user/Tv_Shows
|
||||
)
|
||||
BACKUP_VERIFY_SAMPLE=10 # number of files to randomly sample per share per run
|
||||
BACKUP_VERIFY_MIN_SIZE=1M # skip files smaller than this — avoids tiny junk files
|
||||
|
||||
# ━━━ SMART Health ━━━
|
||||
# Monitors drive SMART attributes — reads live from each drive, no persistent writes.
|
||||
# Discovers all drives automatically via /dev/sd* and /dev/nvme* — no drive list needed.
|
||||
# Add drives to SMART_IGNORE_DRIVES to skip specific drives (e.g. your unRAID boot USB).
|
||||
SMART_TEMP_WARN=45 # degrees C — warn if drive temperature exceeds this
|
||||
SMART_TEMP_CRIT=55 # degrees C — critical if drive temperature exceeds this
|
||||
SMART_IGNORE_DRIVES=(
|
||||
# "sda" # uncomment to ignore sda — common choice if sda is your unRAID boot USB
|
||||
)
|
||||
|
||||
# ━━━ Bandwidth Monitor ━━━
|
||||
# Logs daily rsync transfer totals to a bounded file on /boot/ — minimal flash wear.
|
||||
# bandwidth_monitor.sh --log-transfer is called by rsync.sh after each successful sync.
|
||||
# bandwidth_monitor.sh --report generates the weekly summary standalone.
|
||||
# File stays bounded to BANDWIDTH_LOG_RETENTION lines — old entries auto-purged on each write.
|
||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db"
|
||||
BANDWIDTH_LOG_RETENTION=90 # days to keep — file never grows beyond ~90 lines
|
||||
BANDWIDTH_WARN_GB=50 # flag in reports if a single sync transfer exceeds this GB
|
||||
|
||||
# ━━━ Health Digest ━━━
|
||||
# Aggregated system health summary from across the ecosystem.
|
||||
# Reads existing state files — no new writes to flash drive.
|
||||
#
|
||||
# Three profiles — switch by changing DIGEST_PROFILE, no cron changes needed:
|
||||
# always — sends every run (schedule daily = daily digest, weekly = weekly digest)
|
||||
# smart — sends only if findings worth reporting (intelligent quiet operation)
|
||||
# weekly — sends once per week on DIGEST_DAY only, silent all other days
|
||||
#
|
||||
# Data sources (reads only — no writes):
|
||||
# Transcode ramdisk state, container watchdog strikes, system watchdog strikes,
|
||||
# failover state, container skip list, bandwidth history, SSL cert days remaining
|
||||
DIGEST_PROFILE="weekly" # always | smart | weekly
|
||||
DIGEST_DAY="Sunday" # day name for weekly profile — must match date +%A output
|
||||
|
||||
# Smart profile triggers — set true to send digest when this condition is found
|
||||
DIGEST_SMART_ON_WATCHDOG=true # send if any watchdog strikes are active
|
||||
DIGEST_SMART_ON_FAILOVER=true # send if failover state is not NORMAL
|
||||
DIGEST_SMART_ON_CERT_WARN=true # send if any cert is under CERT_WARN_DAYS
|
||||
DIGEST_SMART_ON_BANDWIDTH=true # send if any transfer exceeded BANDWIDTH_WARN_GB
|
||||
|
||||
# ━━━ Emby Session Report ━━━
|
||||
# Weekly Emby usage report via API — no persistent writes, queries fresh each run.
|
||||
# Shows active streams, library counts, transcode vs direct play ratio.
|
||||
# Requires an API key from Emby Settings → API Keys in the Emby WebUI.
|
||||
EMBY_URL="http://localhost:8096"
|
||||
EMBY_API_KEY="" # paste your Emby API key here
|
||||
EMBY_REPORT_DAYS=7 # number of days to include in the report period
|
||||
EMBY_REPORT_TOP_N=10 # number of top content items to show in report
|
||||
|
||||
# ==============================================================================================
|
||||
# ── SYSTEM WATCHDOG ───────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Last line of defense — reboots cleanly if the system is about to become unstable.
|
||||
# Works alongside docker_watchdog.sh — containers first, system reboot second.
|
||||
# Thresholds set at "about to fall over" levels — not just high usage.
|
||||
# Strike system prevents rebooting on single spikes — sustained threshold hits trigger action.
|
||||
# Last line of defense — reboots the system cleanly when it is about to become unstable.
|
||||
# Runs every 15 minutes via cron. Works alongside docker_watchdog.sh:
|
||||
# docker_watchdog.sh — container level, minimal disruption, tries to self-heal first
|
||||
# system_watchdog.sh — system level, last resort, reboots when healing has failed
|
||||
#
|
||||
# Strike system: sustained threshold hits trigger reboot — single spikes are ignored.
|
||||
# Each check that exceeds its threshold adds a strike. Strikes reset when recovered.
|
||||
# When strike limit is hit the reboot sequence begins.
|
||||
#
|
||||
# Reboot loop protection: tracks reboot timestamps on /boot/ (survives reboots).
|
||||
# If the server reboots too many times in the window it shuts down instead — a reboot
|
||||
# loop means something is fundamentally wrong that a reboot is not fixing.
|
||||
|
||||
# ━━━ State Files ━━━
|
||||
# Strike counts reset on reboot — /tmp is correct for this
|
||||
# Strike counts reset on reboot — /tmp is correct (fresh start after each reboot)
|
||||
SYS_WATCHDOG_STATE_FILE="/tmp/system_watchdog_state.db"
|
||||
# Persistent container skip list — survives reboots, auto-clears when container recovers
|
||||
# Persistent container skip list — on /boot/ so it survives reboots
|
||||
# Containers added here when docker_watchdog.sh exhausts all restart attempts
|
||||
# Auto-clears when container is found running again after reboot or manual fix
|
||||
SYS_WATCHDOG_FAILED_FILE="/boot/config/system_watchdog_failed.db"
|
||||
# Reboot timestamp log — survives reboots for loop detection
|
||||
# Reboot timestamp log — on /boot/ for reboot loop detection across reboots
|
||||
SYS_WATCHDOG_REBOOT_LOG="/boot/config/system_watchdog_reboots.db"
|
||||
|
||||
# ━━━ Strike and Reboot Loop Settings ━━━
|
||||
# Consecutive threshold hits before triggering reboot
|
||||
# Consecutive threshold hits required before triggering reboot
|
||||
SYS_WATCHDOG_STRIKE_LIMIT=2
|
||||
# Maximum reboots allowed in window before shutdown instead — prevents reboot loops
|
||||
# Maximum reboots allowed within the window before shutting down instead
|
||||
# A reboot loop means something fundamental is broken that rebooting is not fixing
|
||||
SYS_WATCHDOG_REBOOT_LIMIT=3
|
||||
# Window in hours — controls both reboot count window AND rolling log purge
|
||||
# 12 = entries older than 12hrs purge automatically, fresh window starts
|
||||
# Window in hours — controls BOTH the reboot count window AND the rolling log purge
|
||||
# Entries older than this many hours are automatically removed from the reboot log
|
||||
SYS_WATCHDOG_REBOOT_WINDOW_HRS=12
|
||||
|
||||
# ━━━ Thresholds ━━━
|
||||
# Set these at "I am about to become unstable" levels — not just "things are a bit tight"
|
||||
# Set these at "I am about to become unstable" levels — not just "things are a bit high"
|
||||
SYS_WATCHDOG_ROOTFS_PCT=95 # rootfs % — at 95% something is seriously wrong
|
||||
SYS_WATCHDOG_LOG_PCT=95 # /var/log % — log spam filling disk
|
||||
SYS_WATCHDOG_MEM_GB=4 # free RAM GB — 4GB on 128GB system is critical
|
||||
SYS_WATCHDOG_LOG_PCT=95 # /var/log % — log spam filling the filesystem
|
||||
SYS_WATCHDOG_MEM_GB=4 # free RAM GB — 4GB free on 128GB system is critical
|
||||
SYS_WATCHDOG_ARC_PINNED_PCT=98 # ZFS ARC % of max before attempting reclaim
|
||||
SYS_WATCHDOG_ARC_RELEASE_PCT=95 # ZFS ARC % after reclaim that still triggers reboot
|
||||
SYS_WATCHDOG_LOAD_MULTIPLIER=3 # strike if load avg > cores x this value
|
||||
SYS_WATCHDOG_LOAD_MULTIPLIER=3 # strike if load avg > cores x this multiplier
|
||||
SYS_WATCHDOG_ZOMBIE_LIMIT=50 # zombie process count before strike
|
||||
SYS_WATCHDOG_CPU_TEMP_MAX=95 # degrees C — adjust for your CPU tjmax
|
||||
SYS_WATCHDOG_CPU_TEMP_MAX=95 # degrees C — adjust for your specific CPU tjmax
|
||||
|
||||
# ━━━ Check Toggles ━━━
|
||||
# true = run this check / false = skip entirely
|
||||
# Disable checks that are not relevant to your hardware or cause false positives
|
||||
# true = run this check on every watchdog cycle / false = skip entirely
|
||||
# Disable checks not relevant to your hardware or that cause false positives
|
||||
SYS_WATCHDOG_CHECK_ROOTFS=true
|
||||
SYS_WATCHDOG_CHECK_LOG=true
|
||||
SYS_WATCHDOG_CHECK_RAM=true
|
||||
SYS_WATCHDOG_CHECK_ARC=true
|
||||
SYS_WATCHDOG_CHECK_CPU_TEMP=true
|
||||
SYS_WATCHDOG_CHECK_LOAD=false
|
||||
SYS_WATCHDOG_CHECK_LOAD=false # disabled — load spikes during transcoding are normal
|
||||
SYS_WATCHDOG_CHECK_ZOMBIES=true
|
||||
SYS_WATCHDOG_CHECK_CONTAINERS=true
|
||||
SYS_WATCHDOG_CHECK_CONTAINERS=true # checks persistent skip list from docker_watchdog.sh
|
||||
SYS_WATCHDOG_CHECK_DOCKER_DAEMON=true
|
||||
|
||||
# ━━━ Abort Toggles ━━━
|
||||
# true = abort reboot if condition is active / false = reboot anyway
|
||||
# Default true = conservative — set false only when "reboot no matter what" is wanted
|
||||
# Goal: graceful reboot before crash is always better than hard crash mid-operation
|
||||
SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # unhealthy pool + reboot = potential data loss
|
||||
SYS_WATCHDOG_ABORT_ON_PARITY=true # aborting parity better than crashing mid-check
|
||||
SYS_WATCHDOG_ABORT_ON_MOVER=true # aborting move better than crashing mid-move
|
||||
# Controls whether certain conditions prevent a reboot from happening.
|
||||
# true = abort reboot if this condition is active (conservative — default)
|
||||
# false = reboot anyway regardless of this condition (aggressive)
|
||||
# Philosophy: a graceful reboot before crash is always better than a hard crash mid-operation
|
||||
SYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # unhealthy pool + reboot risks data loss
|
||||
SYS_WATCHDOG_ABORT_ON_PARITY=true # aborting parity check beats crashing mid-check
|
||||
SYS_WATCHDOG_ABORT_ON_MOVER=true # aborting mover beats crashing mid-move
|
||||
|
||||
# ==============================================================================================
|
||||
# ──────────────────────── End Of User Variables ───────────────────────────────────────────────
|
||||
|
||||
Reference in New Issue
Block a user