updated failover to have same ddns logic as old script
This commit is contained in:
+661
-427
File diff suppressed because it is too large
Load Diff
+315
-210
@@ -33,7 +33,7 @@
|
|||||||
# ── DOCKER ESSENTIALS ──────────────────────────────────────────────────────────────────────
|
# ── DOCKER ESSENTIALS ──────────────────────────────────────────────────────────────────────
|
||||||
# DOCKER DAILY RESTART Containers restarted daily
|
# DOCKER DAILY RESTART Containers restarted daily
|
||||||
# DOCKER WEEKLY RESTART Containers restarted weekly
|
# DOCKER WEEKLY RESTART Containers restarted weekly
|
||||||
# DOCKER WATCHDOG Container health monitoring — memory, CPU, HTTP
|
# DOCKER WATCHDOG Two-tier self-healing container monitoring
|
||||||
# DOCKER NETWORK CONNECT Connect containers to extra networks on boot
|
# DOCKER NETWORK CONNECT Connect containers to extra networks on boot
|
||||||
#
|
#
|
||||||
# ── UNRAID ESSENTIALS ──────────────────────────────────────────────────────────────────────
|
# ── UNRAID ESSENTIALS ──────────────────────────────────────────────────────────────────────
|
||||||
@@ -53,10 +53,11 @@
|
|||||||
# ── TRANSCODES ─────────────────────────────────────────────────────────────────────────────
|
# ── TRANSCODES ─────────────────────────────────────────────────────────────────────────────
|
||||||
# TRANSCODE MANAGER Ramdisk and SSD fallback transcode management
|
# TRANSCODE MANAGER Ramdisk and SSD fallback transcode management
|
||||||
#
|
#
|
||||||
# ── MONITORS ────────────────────────────────────────────────────────────────────────────────
|
# ── MONITORS ───────────────────────────────────────────────────────────────────────────────
|
||||||
# CERTIFICATE MONITOR SSL certificate expiry monitoring
|
# CERTIFICATE MONITOR SSL certificate expiry monitoring
|
||||||
# BACKUP VERIFY Random sample checksum verification against remote
|
# BACKUP VERIFY Random sample checksum verification against remote
|
||||||
# SMART HEALTH Drive SMART attribute monitoring
|
# SMART HEALTH Drive SMART attribute monitoring
|
||||||
|
# ZFS MEMORY SNAPSHOT Weekly ZFS health and memory diagnostic report
|
||||||
# BANDWIDTH MONITOR Daily rsync transfer logging and weekly summary
|
# BANDWIDTH MONITOR Daily rsync transfer logging and weekly summary
|
||||||
# HEALTH DIGEST Aggregated system health digest — always/smart/weekly
|
# HEALTH DIGEST Aggregated system health digest — always/smart/weekly
|
||||||
# EMBY SESSION REPORT Weekly Emby usage statistics via API
|
# EMBY SESSION REPORT Weekly Emby usage statistics via API
|
||||||
@@ -290,76 +291,227 @@ declare -A PROFILE_SKIP_DISK_CHECK=(
|
|||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── FAILOVER ──────────────────────────────────────────────────────────────────────────────────
|
# ── FAILOVER ──────────────────────────────────────────────────────────────────────────────────
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# Mutual container failover between two unRAID servers 50 miles apart.
|
# Mutual container failover between two unRAID servers.
|
||||||
# Each server runs Failover/failover.sh independently — no coordination between servers.
|
# Each server runs Failover/failover.sh independently — no coordination between servers.
|
||||||
# All decisions are based solely on two ping checks: remote reachable + internet reachable.
|
# All decisions based solely on two pings: remote reachable + internet reachable.
|
||||||
#
|
#
|
||||||
# States:
|
# ── STATES ────────────────────────────────────────────────────────────────────────────────────
|
||||||
# NORMAL — remote up, internet up — own containers only, silent operation
|
# NORMAL — remote up, internet up — own containers only, DDNS ON, silent
|
||||||
# FAILOVER — remote down, internet up — start remote's containers locally (additive)
|
# FAILOVER — remote down, internet up — start remote containers (tiered by time)
|
||||||
# NO_INTERNET — internet down — stop public-facing containers, wait for recovery
|
# NO_INTERNET — internet down — stop own DDNS immediately, wait for recovery
|
||||||
# DARK — remote down + internet down — same actions as NO_INTERNET
|
# DARK — remote down AND internet down — same as NO_INTERNET
|
||||||
#
|
#
|
||||||
# Handback: strike confirmation → pre-flight → rsync → start remote → stop local
|
# ── DDNS RULES — ABSOLUTE ─────────────────────────────────────────────────────────────────────
|
||||||
|
# Each server owns its own DDNS — ON when that server has internet
|
||||||
|
# Script controls DDNS exclusively — network state NEVER auto-starts DDNS
|
||||||
|
# Internet loss → stop own DDNS immediately
|
||||||
|
# Failover → start remote DDNS as first action (Tier 1)
|
||||||
|
# Handback → stop remote DDNS FIRST → rsync → start local containers
|
||||||
|
# → start local DDNS LAST — only after containers confirmed up
|
||||||
|
# One DDNS per domain active at all times — never two, never zero for long
|
||||||
|
# 1 minute TTL + 1 minute check interval = minimal user impact
|
||||||
|
#
|
||||||
|
# ── HANDBACK SEQUENCE ─────────────────────────────────────────────────────────────────────────
|
||||||
|
# Strike confirmation → pre-flight → stop remote DDNS → stop remote containers
|
||||||
|
# → rsync writeback → start local containers → start local DDNS → NORMAL
|
||||||
|
# Containers only down during rsync window — minimise this time
|
||||||
|
#
|
||||||
|
# ── TIERED FAILOVER ───────────────────────────────────────────────────────────────────────────
|
||||||
|
# Tier 1 — Immediate — vital services + Live TV — people are watching, can't wait
|
||||||
|
# Tier 2 — configurable delay — shared productivity services
|
||||||
|
# Tier 3 — configurable delay — secondary services
|
||||||
|
# Tier 4 — configurable delay — arrs + downloaders — workflow continuity
|
||||||
|
# Delays set independently per host below
|
||||||
|
|
||||||
EXTERNAL_IP="8.8.8.8" # external IP to ping for internet connectivity check
|
EXTERNAL_IP="8.8.8.8" # external IP to ping for internet connectivity check
|
||||||
FAILOVER_CHECK_INTERVAL=120 # seconds between state checks
|
FAILOVER_CHECK_INTERVAL=120 # seconds between state checks
|
||||||
FAILOVER_HANDBACK_STRIKES=2 # consecutive remote-up confirmations before handback
|
# 1 minute TTL + 2 minute interval = minimal gap
|
||||||
|
FAILOVER_HANDBACK_STRIKES=2 # consecutive remote-up checks before handback
|
||||||
|
# 2 strikes x 120s = 4 min confirmation window
|
||||||
FAILOVER_STATE_FILE="/boot/config/failover_state.db"
|
FAILOVER_STATE_FILE="/boot/config/failover_state.db"
|
||||||
|
# persists on /boot/ — survives reboots
|
||||||
|
# tracks: state, failover_start, strikes, tier flags
|
||||||
|
|
||||||
# ━━━ Failover Test ━━━
|
# ━━━ Failover Test ━━━
|
||||||
# Used by Failover/failover_test.sh — controlled simulation of the failover lifecycle.
|
# Used by Failover/failover_test.sh — controlled simulation via iptables block.
|
||||||
# failover_test.sh blocks remote connectivity via iptables then observes failover.sh behavior.
|
|
||||||
# All failover logic stays in failover.sh — test script is the harness only.
|
# All failover logic stays in failover.sh — test script is the harness only.
|
||||||
#
|
# ⚠️ Run during maintenance window — real containers start and stop during the test.
|
||||||
# ⚠️ Run during a maintenance window — real containers start and stop during the test.
|
|
||||||
# Use --dry-run first to walk through phases without touching anything.
|
# Use --dry-run first to walk through phases without touching anything.
|
||||||
|
FAILOVER_TEST_BLOCK_WAIT=150 # seconds to hold iptables block
|
||||||
# Seconds to hold the iptables block — must be longer than FAILOVER_CHECK_INTERVAL
|
# must be > FAILOVER_CHECK_INTERVAL + buffer
|
||||||
# so failover.sh has time to detect the outage and change state
|
FAILOVER_TEST_HANDBACK_WAIT=360 # seconds to wait for handback completion
|
||||||
FAILOVER_TEST_BLOCK_WAIT=150 # 150s = FAILOVER_CHECK_INTERVAL + 30s buffer
|
# covers FAILOVER_HANDBACK_STRIKES x INTERVAL + rsync
|
||||||
|
|
||||||
# Seconds to wait for handback after restoring connectivity
|
# ━━━ DDNS — Script Controlled Exclusively ━━━
|
||||||
# Must cover FAILOVER_HANDBACK_STRIKES x FAILOVER_CHECK_INTERVAL plus rsync time
|
# Each server owns its own DDNS containers — one domain per server.
|
||||||
# 2 strikes x 120s = 240s minimum — add buffer for rsync handback jobs
|
# DDNS is started and stopped ONLY by this script — never by network state returning.
|
||||||
FAILOVER_TEST_HANDBACK_WAIT=360 # 360s = 6 minutes — adjust if rsync takes longer
|
# HOST1 DDNS starts last in handback (after containers confirmed up).
|
||||||
|
# HOST1 DDNS stops first on internet loss.
|
||||||
|
# HOST2 DDNS starts when HOST2 detects HOST1 is down (Tier 1).
|
||||||
FAILOVER_HOST1_STARTS_FOR_HOST2=(
|
# HOST2 DDNS stops before handback rsync begins.
|
||||||
"Gmer4Lfe.com"
|
|
||||||
"Gmer4Lfe.us"
|
HOST1_DDNS_CONTAINERS=(
|
||||||
|
"Gmer4Lfe.com" # HOST1's own DDNS — ON when HOST1 has internet
|
||||||
|
# covers gmer4lfe.com pointing to HOST1 IP
|
||||||
)
|
)
|
||||||
|
|
||||||
|
HOST2_DDNS_CONTAINERS=(
|
||||||
|
"Gmer4Lfe.us" # HOST2's own DDNS — ON when HOST2 has internet
|
||||||
|
# covers gmer4lfe.us pointing to HOST2 IP
|
||||||
|
)
|
||||||
|
|
||||||
|
# ━━━ Containers to stop on internet loss ━━━
|
||||||
|
# Own DDNS handled separately above — list additional containers here if needed
|
||||||
|
# These stop when this server loses internet — regardless of remote state
|
||||||
FAILOVER_HOST1_STOP_ON_NO_NET=(
|
FAILOVER_HOST1_STOP_ON_NO_NET=(
|
||||||
"Gmer4Lfe.com"
|
# "container-name" # add containers that should stop without internet
|
||||||
"Gmer4Lfe.us"
|
|
||||||
)
|
|
||||||
FAILOVER_HOST1_RSYNC_JOBS=(
|
|
||||||
# "/mnt/user/appdata-Failover/Jayred365"
|
|
||||||
# "/mnt/user/Media_Server/Emby-Jayred"
|
|
||||||
)
|
|
||||||
|
|
||||||
# HOST2 (unRAID-Jayred365 — Secondary)
|
|
||||||
FAILOVER_HOST2_STARTS_FOR_HOST1=(
|
|
||||||
"Emby"
|
|
||||||
"Gmer4Lfe.com"
|
|
||||||
"Gmer4Lfe.us"
|
|
||||||
)
|
)
|
||||||
|
|
||||||
FAILOVER_HOST2_STOP_ON_NO_NET=(
|
FAILOVER_HOST2_STOP_ON_NO_NET=(
|
||||||
"Gmer4Lfe.com"
|
# "container-name"
|
||||||
"Gmer4Lfe.us"
|
|
||||||
)
|
)
|
||||||
FAILOVER_HOST2_RSYNC_JOBS=(
|
|
||||||
# "/mnt/user/appdata-Failover/Gmer4Lfe"
|
# ━━━ HOST1 runs these for HOST2 when HOST2 goes down ━━━
|
||||||
|
# HOST2's DDNS listed in Tier 1 — starts immediately as first action
|
||||||
|
# List HOST2's specific services here — HOST2's own containers only
|
||||||
|
# Do NOT list shared services that HOST1 already runs
|
||||||
|
|
||||||
|
# Tier 1 — Immediate — starts as soon as HOST2 is detected down
|
||||||
|
FAILOVER_HOST1_RUNS_FOR_HOST2_IMMEDIATE=(
|
||||||
|
"Gmer4Lfe.us" # HOST2's DDNS — start first, covers HOST2's domain
|
||||||
|
"VaultWarden-Jayred365" # HOST2's password manager — immediate access needed
|
||||||
|
# "container-placeholder" # add HOST2 specific services here
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# Tier 2 — starts after HOST2_TIER2_DELAY minutes
|
||||||
|
FAILOVER_HOST1_RUNS_FOR_HOST2_2HR=(
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Tier 3 — starts after HOST2_TIER3_DELAY minutes
|
||||||
|
FAILOVER_HOST1_RUNS_FOR_HOST2_6HR=(
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Tier 4 — starts after HOST2_TIER4_DELAY minutes
|
||||||
|
FAILOVER_HOST1_RUNS_FOR_HOST2_18HR=(
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
|
# ━━━ HOST2 runs these for HOST1 when HOST1 goes down ━━━
|
||||||
|
# HOST1's DDNS listed in Tier 1 — starts immediately to cover HOST1's domain
|
||||||
|
# Live TV in Tier 1 — people are watching, cannot wait for tiered startup
|
||||||
|
# Auth stack in Tier 1 — everything proxied through NPM needs auth
|
||||||
|
|
||||||
|
# Tier 1 — Immediate — vital services and Live TV cannot wait
|
||||||
|
FAILOVER_HOST2_RUNS_FOR_HOST1_IMMEDIATE=(
|
||||||
|
"Gmer4Lfe.com" # HOST1's DDNS — start first, covers HOST1's domain
|
||||||
|
"Emby" # media server — users are actively watching
|
||||||
|
"NginxProxyManager" # reverse proxy — all services route through this
|
||||||
|
"Lldap-Gmer4Lfe" # auth directory — required by Authelia
|
||||||
|
"Mariadb-Authelia" # auth database — required by Authelia
|
||||||
|
"Redis-Authelia" # auth cache — required by Authelia
|
||||||
|
"Authelia" # authentication — required for all proxied services
|
||||||
|
"Authelia-Secondary" # auth redundancy
|
||||||
|
"Redis-Authelia-Secondary" # auth secondary cache
|
||||||
|
"VaultWarden-Gmer4Lfe" # password manager — critical, immediate access needed
|
||||||
|
"Dispatcharr" # Live TV — people are watching, cannot wait
|
||||||
|
"Dispatcharr-Basic" # Live TV basic profile
|
||||||
|
"Dispatcharr-Iptv-Users" # Live TV IPTV users
|
||||||
|
"ErsatzTV-Emby" # Live TV scheduling and channel management
|
||||||
|
)
|
||||||
|
|
||||||
|
# Tier 2 — starts after HOST1_TIER2_DELAY minutes
|
||||||
|
# Productivity services — important but can wait a couple of hours
|
||||||
|
FAILOVER_HOST2_RUNS_FOR_HOST1_2HR=(
|
||||||
|
"Postgres-NextCloud" # NextCloud database — must start before NextCloud
|
||||||
|
"NextCloud" # file access and collaboration
|
||||||
|
"PostgreSQL_Immich" # Immich database
|
||||||
|
"Immich-Gmer4Lfe" # photo management
|
||||||
|
"Jellyseerr" # media request management
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Tier 3 — starts after HOST1_TIER3_DELAY minutes
|
||||||
|
# Secondary services — useful but not immediately critical
|
||||||
|
FAILOVER_HOST2_RUNS_FOR_HOST1_6HR=(
|
||||||
|
"Organizrv2-Gmer4Lfe" # dashboard — nice to have
|
||||||
|
"AdGuard-Home" # DNS filtering
|
||||||
|
"UptimeKuma-Gmer4Lfe" # uptime monitoring
|
||||||
|
"Gitea" # git server
|
||||||
|
"Collabora-CODE" # document editing for NextCloud
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Tier 4 — starts after HOST1_TIER4_DELAY minutes
|
||||||
|
# Full workflow mode — arrs and downloaders
|
||||||
|
# Minimal writeback on handback — start fresh is cleaner than syncing download state
|
||||||
|
FAILOVER_HOST2_RUNS_FOR_HOST1_18HR=(
|
||||||
|
"Sonarr" # TV show management
|
||||||
|
"Radarr" # movie management
|
||||||
|
"Lidarr" # music management
|
||||||
|
"Readarr" # book management
|
||||||
|
"Prowlarr" # indexer management
|
||||||
|
"Bazarr" # subtitle management
|
||||||
|
"SABnzbd-Gmer4Lfe" # usenet downloader
|
||||||
|
"Qbittorrent-Gmer4Lfe" # torrent downloader
|
||||||
|
"LidaTube" # YouTube music downloader
|
||||||
|
"Pinchflat" # YouTube channel downloader
|
||||||
|
"ChannelTube" # YouTube channel management
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
|
# ━━━ Tier Delay Settings ━━━
|
||||||
|
# How long primary must be down before each tier activates — set in minutes
|
||||||
|
# Tier 1 is always immediate — no delay
|
||||||
|
# Set independently per host — a large server may want longer delays than a small one
|
||||||
|
# Adjust based on your tolerance for resource usage on the covering server
|
||||||
|
|
||||||
|
# Delays for HOST1's containers running on HOST2 (HOST1 is down)
|
||||||
|
HOST1_TIER2_DELAY=120 # 2 hours — NextCloud, Immich can wait
|
||||||
|
HOST1_TIER3_DELAY=360 # 6 hours — dashboard, monitoring, Gitea
|
||||||
|
HOST1_TIER4_DELAY=1080 # 18 hours — full workflow, arrs and downloaders
|
||||||
|
|
||||||
|
# Delays for HOST2's containers running on HOST1 (HOST2 is down)
|
||||||
|
HOST2_TIER2_DELAY=120
|
||||||
|
HOST2_TIER3_DELAY=360
|
||||||
|
HOST2_TIER4_DELAY=1080
|
||||||
|
|
||||||
|
# ━━━ Rsync Writeback Jobs ━━━
|
||||||
|
# Run during handback — syncs critical appdata back to primary before containers restart.
|
||||||
|
# Containers are stopped before this runs — clean source, no competing writes.
|
||||||
|
# Full bandwidth available — DDNS stopped, containers stopped, nothing competing.
|
||||||
|
#
|
||||||
|
# Priority:
|
||||||
|
# Critical — Emby userdata/playstates (small, fast, important)
|
||||||
|
# Critical — Auth stack data
|
||||||
|
# Skip — Media files (already on primary, never moved)
|
||||||
|
# Skip — Downloads (start fresh — cleaner than syncing partial state)
|
||||||
|
#
|
||||||
|
# Format: "/path/to/source" — matched to rsync profile by directory basename
|
||||||
|
|
||||||
|
# HOST1 writeback — run by HOST2 during HOST1 handback
|
||||||
|
FAILOVER_HOST1_WRITEBACK=(
|
||||||
|
"/mnt/user/appdata-Failover/Critical-Data" # auth stack — Authelia, Mariadb, Redis, LLDAP, NPM
|
||||||
|
"/mnt/user/appdata-Failover/Important-Data" # NextCloud + Postgres
|
||||||
|
"/mnt/user/appdata-Failover/Emby" # Emby userdata, playstates, metadata
|
||||||
|
"/mnt/user/appdata-Failover/Gmer4Lfe" # server specific appdata — Organizr, UptimeKuma
|
||||||
|
# "/mnt/user/appdata-Failover/Arrs_Stack" # skip — arrs start fresh on handback
|
||||||
|
)
|
||||||
|
|
||||||
|
# HOST2 writeback — run by HOST1 during HOST2 handback
|
||||||
|
FAILOVER_HOST2_WRITEBACK=(
|
||||||
|
# "/mnt/user/appdata-Failover/Jayred365" # HOST2 specific appdata
|
||||||
|
# "container-placeholder"
|
||||||
|
)
|
||||||
|
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── DOCKER ESSENTIALS ─────────────────────────────────────────────────────────────────────────
|
# ── DOCKER ESSENTIALS ─────────────────────────────────────────────────────────────────────────
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
|
|
||||||
# ━━━ Docker Daily Restart ━━━
|
# ━━━ Docker Daily Restart ━━━
|
||||||
# Containers restarted every day by Docker_Essentials/docker_daily_restart.sh.
|
# Containers restarted every day — keeps services fresh, clears memory leaks.
|
||||||
# Keeps services fresh and clears memory leaks that accumulate over time.
|
# Case-sensitive — must match exact Docker container names in the unRAID Docker tab.
|
||||||
# Case-sensitive — must match exact Docker container names shown in the unRAID Docker tab.
|
|
||||||
DAILY_RESTART_CONTAINERS=(
|
DAILY_RESTART_CONTAINERS=(
|
||||||
"NginxProxyManager"
|
"NginxProxyManager"
|
||||||
"Authelia"
|
"Authelia"
|
||||||
@@ -370,8 +522,8 @@ DAILY_RESTART_CONTAINERS=(
|
|||||||
)
|
)
|
||||||
|
|
||||||
# ━━━ Docker Weekly Restart ━━━
|
# ━━━ Docker Weekly Restart ━━━
|
||||||
# Containers restarted once per week by Docker_Essentials/docker_weekly_restart.sh.
|
# Less critical services that benefit from periodic restart but don't need daily cycling.
|
||||||
# For less critical services that benefit from periodic restart but don't need daily cycling.
|
# Case-sensitive — must match exact Docker container names in the unRAID Docker tab.
|
||||||
WEEKLY_RESTART_CONTAINERS=(
|
WEEKLY_RESTART_CONTAINERS=(
|
||||||
"NextCloud"
|
"NextCloud"
|
||||||
"Organizrv2-Gmer4Lfe"
|
"Organizrv2-Gmer4Lfe"
|
||||||
@@ -380,20 +532,18 @@ WEEKLY_RESTART_CONTAINERS=(
|
|||||||
)
|
)
|
||||||
|
|
||||||
# ━━━ Docker Watchdog ━━━
|
# ━━━ Docker Watchdog ━━━
|
||||||
# Two-tier self-healing container monitoring:
|
# Two-tier self-healing container monitoring.
|
||||||
# Tier 1 — strict monitoring of explicitly configured containers
|
# Tier 1 — strict monitoring of explicitly configured containers
|
||||||
# Tier 2 — global health scan of ALL running containers
|
# Tier 2 — global health scan of ALL running containers
|
||||||
#
|
#
|
||||||
# Cross-cutting intelligence applies to both tiers:
|
# Cross-cutting intelligence:
|
||||||
# Startup grace — skip restarts while system is still booting
|
# Startup grace — skip restarts while system is still booting
|
||||||
# Dependency order — restart database before app, not the other way around
|
# Dependency order — restart database before app
|
||||||
# Restart loop — stop restarting after limit hit → skip list → notify critical
|
# Restart loop — stop restarting after limit hit → skip list → notify critical
|
||||||
# Skip list — persistent across reboots, auto-clears when container recovers
|
# Skip list — persistent across reboots, auto-clears when container recovers
|
||||||
# Batch notify — one clean summary per run instead of one ping per event
|
# Batch notify — one clean summary per run
|
||||||
|
|
||||||
# ── Tier 1 — Strict Monitoring ────────────────────────────────────────────────────────────
|
# Memory hard limits in MB — immediate restart if exceeded
|
||||||
|
|
||||||
# Memory hard limits in MB — immediate restart if exceeded, no strike system
|
|
||||||
# 20GB=20480 16GB=16384 14GB=14336 12GB=12288 10GB=10240
|
# 20GB=20480 16GB=16384 14GB=14336 12GB=12288 10GB=10240
|
||||||
# 8GB=8192 6GB=6144 4GB=4096 2GB=2048 1GB=1024
|
# 8GB=8192 6GB=6144 4GB=4096 2GB=2048 1GB=1024
|
||||||
declare -A WATCHDOG_CONTAINERS=(
|
declare -A WATCHDOG_CONTAINERS=(
|
||||||
@@ -402,14 +552,14 @@ declare -A WATCHDOG_CONTAINERS=(
|
|||||||
["Tdarr"]=6144
|
["Tdarr"]=6144
|
||||||
["Code-Server"]=1024
|
["Code-Server"]=1024
|
||||||
)
|
)
|
||||||
|
|
||||||
# HTTP responsiveness checks — omit container to skip its HTTP check
|
|
||||||
declare -A WATCHDOG_CONTAINER_URLS=(
|
declare -A WATCHDOG_CONTAINER_URLS=(
|
||||||
["Emby"]="http://localhost:8096"
|
["Emby"]="http://localhost:8096"
|
||||||
)
|
)
|
||||||
|
|
||||||
# Containers that must always be running — strike system, persistent skip list on /boot/
|
# Containers that must always be running — strike system, persistent skip list on /boot/
|
||||||
# Skip list auto-clears when container recovers — no manual intervention for normal recovery
|
# Skip list auto-clears when container recovers — no manual intervention for normal recovery
|
||||||
|
# These are your core auth and proxy stack — everything depends on them being up
|
||||||
WATCHDOG_REQUIRED_CONTAINERS=(
|
WATCHDOG_REQUIRED_CONTAINERS=(
|
||||||
"NginxProxyManager"
|
"NginxProxyManager"
|
||||||
"Lldap-Gmer4Lfe"
|
"Lldap-Gmer4Lfe"
|
||||||
@@ -419,81 +569,74 @@ WATCHDOG_REQUIRED_CONTAINERS=(
|
|||||||
"Authelia-Secondary"
|
"Authelia-Secondary"
|
||||||
"Redis-Authelia-Secondary"
|
"Redis-Authelia-Secondary"
|
||||||
)
|
)
|
||||||
|
|
||||||
# Strike thresholds
|
# Strike state file — /tmp resets on reboot which is correct for strike tracking
|
||||||
WATCHDOG_STATE_FILE="/tmp/container_watchdog_state.db"
|
WATCHDOG_STATE_FILE="/tmp/container_watchdog_state.db"
|
||||||
SOFT_CPU_THRESHOLD=80 # warn at this % of total system CPU
|
|
||||||
HARD_CPU_THRESHOLD=85 # strike at this % of total system CPU
|
# CPU thresholds — normalised against total core count automatically at runtime
|
||||||
CPU_FAIL_LIMIT=2 # consecutive CPU strikes before restart
|
SOFT_CPU_THRESHOLD=80 # warn at this % of total system CPU
|
||||||
SOFT_MEM_THRESHOLD=80 # warn when container reaches this % of hard limit
|
HARD_CPU_THRESHOLD=85 # strike at this % of total system CPU
|
||||||
RESP_FAIL_LIMIT=2 # consecutive failed HTTP checks before restart
|
CPU_FAIL_LIMIT=2 # consecutive CPU strikes before container restart
|
||||||
CURL_TIMEOUT=5 # seconds before curl gives up per check
|
|
||||||
|
# Memory soft threshold — warn when container reaches this % of its hard limit
|
||||||
# ── Tier 2 — Global Health Scan ───────────────────────────────────────────────────────────
|
# Hard limit exceeded triggers immediate restart regardless of strikes
|
||||||
|
SOFT_MEM_THRESHOLD=80
|
||||||
# Master toggle — false disables Tier 2 entirely
|
|
||||||
|
# HTTP responsiveness check settings
|
||||||
|
RESP_FAIL_LIMIT=2 # consecutive failed curl checks before restart
|
||||||
|
CURL_TIMEOUT=5 # seconds before curl gives up per check
|
||||||
|
|
||||||
|
# Tier 2 master toggle — false disables global scan entirely
|
||||||
WATCHDOG_SCAN_ALL=true
|
WATCHDOG_SCAN_ALL=true
|
||||||
|
|
||||||
# Containers to skip in Tier 2 — add containers expected to be in a non-running state
|
# Containers to skip in Tier 2 scan entirely
|
||||||
# or managed by other systems that should not be auto-restarted
|
# Add intentionally stopped containers or containers managed by other systems
|
||||||
WATCHDOG_SCAN_IGNORE=(
|
WATCHDOG_SCAN_IGNORE=(
|
||||||
# "container-name"
|
# "container-name"
|
||||||
)
|
)
|
||||||
|
|
||||||
# Individual Tier 2 check toggles — disable checks that cause false positives
|
# Individual Tier 2 check toggles — disable checks that cause false positives
|
||||||
WATCHDOG_RESTART_UNHEALTHY=true # restart containers with unhealthy Docker health status
|
WATCHDOG_RESTART_UNHEALTHY=true # restart containers with unhealthy Docker health status
|
||||||
WATCHDOG_RESTART_DEAD=true # remove and restart containers in dead state
|
WATCHDOG_RESTART_DEAD=true # remove and restart containers in dead state
|
||||||
WATCHDOG_RESTART_CRASHED=true # restart containers that exited with non-zero exit code
|
WATCHDOG_RESTART_CRASHED=true # restart containers that exited with non-zero exit code
|
||||||
WATCHDOG_NOTIFY_OOM=true # restart and notify when OOM killed by kernel
|
WATCHDOG_NOTIFY_OOM=true # restart and notify when OOM killed by kernel
|
||||||
WATCHDOG_NOTIFY_CRASHLOOP=true # notify when Docker restart count is climbing
|
WATCHDOG_NOTIFY_CRASHLOOP=true # notify when Docker restart count is climbing
|
||||||
|
|
||||||
# Crash loop threshold — notify critical if Docker has restarted container this many times
|
# Crash loop threshold — notify critical if Docker has restarted container this many times
|
||||||
WATCHDOG_CRASH_LIMIT=5
|
WATCHDOG_CRASH_LIMIT=5
|
||||||
|
|
||||||
# ── Cross-cutting Intelligence ────────────────────────────────────────────────────────────
|
# Startup grace — skip restarts while system is still booting
|
||||||
|
|
||||||
# Startup grace period — skip restarts while system is still booting
|
|
||||||
# Prevents false positives while containers are coming up after array start
|
# Prevents false positives while containers are coming up after array start
|
||||||
WATCHDOG_STARTUP_GRACE=300 # seconds after boot before watchdog acts on failures
|
WATCHDOG_STARTUP_GRACE=600 # seconds after boot before watchdog acts on failures
|
||||||
|
|
||||||
# Restart loop protection — stops hammering broken containers
|
# Restart loop protection — stops hammering broken containers
|
||||||
# Tracks watchdog-initiated restarts per container in a bounded /boot/ file
|
# After limit hit → skip list → notify critical → manual intervention needed
|
||||||
# After limit hit → container added to skip list → notify critical → manual intervention
|
|
||||||
# Skip list auto-clears when container is found running again
|
# Skip list auto-clears when container is found running again
|
||||||
WATCHDOG_CONTAINER_RESTART_LIMIT=3 # max watchdog restarts allowed in window
|
WATCHDOG_CONTAINER_RESTART_LIMIT=3 # max watchdog restarts allowed in window
|
||||||
WATCHDOG_CONTAINER_RESTART_WINDOW=1 # hours — rolling window for restart count
|
WATCHDOG_CONTAINER_RESTART_WINDOW=1 # hours — rolling window for restart count
|
||||||
WATCHDOG_CONTAINER_RESTART_LOG="/boot/config/container_restart_history.db"
|
WATCHDOG_CONTAINER_RESTART_LOG="/boot/config/container_restart_history.db"
|
||||||
# /boot/ survives reboots — bounded, auto-purges old entries
|
# /boot/ survives reboots — bounded, auto-purges
|
||||||
|
|
||||||
# Dependency ordering — skip restarting a container if its dependency is also down
|
# Dependency ordering — skip restarting a container if its dependency is also down
|
||||||
# Dependency gets restarted first, dependent picked up on the next watchdog cycle
|
# Dependency gets restarted first, dependent picked up on the next watchdog cycle
|
||||||
|
# Prevents Authelia restarting before its database is ready — it would just fail again
|
||||||
# Format: ["dependent"]="dependency1 dependency2"
|
# Format: ["dependent"]="dependency1 dependency2"
|
||||||
declare -A WATCHDOG_DEPENDENCIES=(
|
declare -A WATCHDOG_DEPENDENCIES=(
|
||||||
["Authelia"]="Mariadb-Authelia Redis-Authelia"
|
["Authelia"]="Mariadb-Authelia Redis-Authelia"
|
||||||
["Authelia-Secondary"]="Mariadb-Authelia Redis-Authelia-Secondary"
|
["Authelia-Secondary"]="Mariadb-Authelia Redis-Authelia-Secondary"
|
||||||
["NextCloud"]="Postgres-NextCloud"
|
["NextCloud"]="Postgres-NextCloud"
|
||||||
)
|
)
|
||||||
|
|
||||||
# Notification batching — one clean summary per run instead of one ping per event
|
# Notification batching — one clean summary per run instead of one ping per event
|
||||||
# true = batch all events into a single notification at end of run
|
# true = batch all events into a single notification at end of run
|
||||||
# false = send individual notification per event as it happens
|
# false = send individual notification per event as it happens
|
||||||
WATCHDOG_BATCH_NOTIFY=true
|
WATCHDOG_BATCH_NOTIFY=true
|
||||||
|
|
||||||
# ━━━ Docker Network Connect ━━━
|
|
||||||
NETWORK_CONNECT_CONTAINERS=(
|
|
||||||
"memcached"
|
|
||||||
"Npm-CrowdSec"
|
|
||||||
)
|
|
||||||
NETWORK_CONNECT_NETWORKS=(
|
|
||||||
"nextcloud-aio"
|
|
||||||
)
|
|
||||||
|
|
||||||
# ━━━ Docker Network Connect ━━━
|
# ━━━ Docker Network Connect ━━━
|
||||||
# Connects containers to extra Docker networks on array start.
|
# Connects containers to extra Docker networks on array start — many-to-many.
|
||||||
|
# Every container in the list connects to every network in the list.
|
||||||
# Useful when containers need to communicate across networks they were not originally
|
# Useful when containers need to communicate across networks they were not originally
|
||||||
# configured with — e.g. memcached needing access to the nextcloud-aio network.
|
# configured with — e.g. memcached needing access to the nextcloud-aio network.
|
||||||
# Every container in the list connects to every network in the list (many-to-many).
|
|
||||||
# Comment out entries to disable without removing them.
|
|
||||||
NETWORK_CONNECT_CONTAINERS=(
|
NETWORK_CONNECT_CONTAINERS=(
|
||||||
"memcached"
|
"memcached"
|
||||||
"Npm-CrowdSec"
|
"Npm-CrowdSec"
|
||||||
@@ -508,42 +651,37 @@ NETWORK_CONNECT_NETWORKS=(
|
|||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
|
|
||||||
# ━━━ Reboot ━━━
|
# ━━━ Reboot ━━━
|
||||||
# Seconds of warning broadcast to all logged-in users before server_reboot.sh reboots.
|
# Seconds of warning broadcast to logged-in users before server_reboot.sh reboots.
|
||||||
# Gives users time to save work or finish what they are doing before the system goes down.
|
# Gives users time to save work before the system goes down.
|
||||||
REBOOT_SLEEP=300
|
REBOOT_SLEEP=300
|
||||||
|
|
||||||
# ━━━ Mover ━━━
|
# ━━━ Mover ━━━
|
||||||
# Seconds to wait after warning users before mover_stop.sh sends SIGTERM to the mover.
|
# Seconds to wait before mover_stop.sh sends SIGTERM to the mover process.
|
||||||
# Gives the mover time to finish its current file operation cleanly before being killed.
|
# Gives the mover time to finish its current file operation cleanly before being killed.
|
||||||
MOVER_STOP_TIMEOUT=300
|
MOVER_STOP_TIMEOUT=300
|
||||||
|
|
||||||
# ━━━ Syslog Filter ━━━
|
# ━━━ Syslog Filter ━━━
|
||||||
# Path for the rsyslog filter file created by docker_syslog_filter.sh.
|
# Path for the rsyslog filter file that suppresses Docker veth noise from syslog.
|
||||||
# The filter suppresses noisy Docker veth and docker0 messages from syslog on boot.
|
# Without this filter every Docker network interface change floods the syslog on boot.
|
||||||
# Without this filter, every Docker network interface change floods the syslog.
|
|
||||||
FILTER_FILE="/etc/rsyslog.d/ignore-docker-veth.conf"
|
FILTER_FILE="/etc/rsyslog.d/ignore-docker-veth.conf"
|
||||||
|
|
||||||
# ━━━ PHP-FPM ━━━
|
# ━━━ PHP-FPM ━━━
|
||||||
# Config file path and max children value for php_fpm_max_children.sh.
|
|
||||||
# Higher max_children allows more concurrent PHP requests to the unRAID WebGUI.
|
# Higher max_children allows more concurrent PHP requests to the unRAID WebGUI.
|
||||||
# Set based on available RAM — too high can cause memory pressure on low-RAM systems.
|
# Set based on available RAM — too high can cause memory pressure on low-RAM systems.
|
||||||
PHP_CONF="/etc/php-fpm.d/www.conf"
|
PHP_CONF="/etc/php-fpm.d/www.conf"
|
||||||
PHP_MAX_CHILDREN=250
|
PHP_MAX_CHILDREN=250
|
||||||
|
|
||||||
# ━━━ Clear Logs ━━━
|
# ━━━ Clear Logs ━━━
|
||||||
# System log files cleared by clear_logs.sh — Docker container logs are cleared too.
|
# System log files cleared weekly to prevent rootfs fill over time.
|
||||||
# Run weekly to prevent logs from filling the rootfs over time.
|
|
||||||
LOG_FILES=(/var/log/syslog /var/log/messages /var/log/dmesg)
|
LOG_FILES=(/var/log/syslog /var/log/messages /var/log/dmesg)
|
||||||
|
|
||||||
# ━━━ WebGUI Watchdog ━━━
|
# ━━━ WebGUI Watchdog ━━━
|
||||||
# Monitors the unRAID WebGUI and restarts services if it becomes unresponsive.
|
# Escalation: nginx restart → recheck → emhttp restart → recheck → notify warning.
|
||||||
# Escalation path: nginx restart → recheck → emhttp restart → recheck → notify warning.
|
# emhttp is the core unRAID daemon — restarting is more disruptive but recovers cleanly.
|
||||||
# emhttp is the core unRAID management daemon — restarting it is more disruptive than nginx
|
WEBGUI_URL="http://localhost" # adjust if running non-standard port
|
||||||
# but both recover cleanly. Notification sent on any restart so you know what happened.
|
|
||||||
WEBGUI_URL="http://localhost" # adjust port if non-standard e.g. http://localhost:8080
|
|
||||||
WEBGUI_TIMEOUT=5 # seconds before curl gives up on the WebGUI check
|
WEBGUI_TIMEOUT=5 # seconds before curl gives up on the WebGUI check
|
||||||
WEBGUI_NGINX_WAIT=15 # seconds to wait after nginx restart before rechecking
|
WEBGUI_NGINX_WAIT=15 # seconds to wait after nginx restart before rechecking
|
||||||
WEBGUI_EMHTTP_WAIT=30 # seconds to wait after emhttp restart — emhttp takes longer
|
WEBGUI_EMHTTP_WAIT=30 # seconds to wait after emhttp restart — takes longer
|
||||||
|
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── MEDIA ─────────────────────────────────────────────────────────────────────────────────────
|
# ── MEDIA ─────────────────────────────────────────────────────────────────────────────────────
|
||||||
@@ -553,11 +691,11 @@ NETWORK_CONNECT_NETWORKS=(
|
|||||||
# Mode and owner applied recursively to all shares in MEDIA_PERMISSION_SHARES.
|
# Mode and owner applied recursively to all shares in MEDIA_PERMISSION_SHARES.
|
||||||
# Run by Media/media_shares_permissions.sh via the media_management.sh orchestrator.
|
# Run by Media/media_shares_permissions.sh via the media_management.sh orchestrator.
|
||||||
# 777 and nobody:users is standard for unRAID media shares accessible by Docker containers.
|
# 777 and nobody:users is standard for unRAID media shares accessible by Docker containers.
|
||||||
|
# Applied recursively so large shares take time — run overnight via orchestrator.
|
||||||
PERMISSIONS_MODE="777"
|
PERMISSIONS_MODE="777"
|
||||||
PERMISSIONS_OWNER="nobody:users"
|
PERMISSIONS_OWNER="nobody:users"
|
||||||
|
|
||||||
# Shares to apply permissions to — add or remove paths as your library grows.
|
# Shares to apply permissions to — add or remove paths as your library grows.
|
||||||
# These are applied recursively so large shares take time — run overnight via orchestrator.
|
|
||||||
MEDIA_PERMISSION_SHARES=(
|
MEDIA_PERMISSION_SHARES=(
|
||||||
/mnt/user/Anime_Movies
|
/mnt/user/Anime_Movies
|
||||||
/mnt/user/Anime_Movies-Old
|
/mnt/user/Anime_Movies-Old
|
||||||
@@ -616,7 +754,8 @@ ANIME_FILE_PATTERNS=(
|
|||||||
'*.log' '*.json'
|
'*.log' '*.json'
|
||||||
)
|
)
|
||||||
|
|
||||||
# File patterns deleted by the media profile — includes *.iso and *.lrc not needed in anime
|
# File patterns deleted by the media profile
|
||||||
|
# Includes *.iso and *.lrc not needed in anime profile
|
||||||
MEDIA_FILE_PATTERNS=(
|
MEDIA_FILE_PATTERNS=(
|
||||||
'*.sfv' '*.md5' '*.sha1' '*.txt' '*.url' '*.lnk'
|
'*.sfv' '*.md5' '*.sha1' '*.txt' '*.url' '*.lnk'
|
||||||
'*.rar' '*.zip' '*.info' '*.torrent' '*.sample*' '*.proof*'
|
'*.rar' '*.zip' '*.info' '*.torrent' '*.sample*' '*.proof*'
|
||||||
@@ -660,7 +799,6 @@ MEDIA_MAINTENANCE_JOBS=(
|
|||||||
LIDARR_MUSIC_ROOT="/mnt/user/Music-New" # must match the root path set in Lidarr
|
LIDARR_MUSIC_ROOT="/mnt/user/Music-New" # must match the root path set in Lidarr
|
||||||
LIDARR_ORPHAN_AGE=7 # days before untracked file is eligible for deletion
|
LIDARR_ORPHAN_AGE=7 # days before untracked file is eligible for deletion
|
||||||
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
||||||
# file extensions considered valid music files
|
|
||||||
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
||||||
# never deleted — cover art, metadata, lyrics
|
# never deleted — cover art, metadata, lyrics
|
||||||
|
|
||||||
@@ -684,43 +822,32 @@ MEDIA_MAINTENANCE_JOBS=(
|
|||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── TRANSCODES ────────────────────────────────────────────────────────────────────────────────
|
# ── TRANSCODES ────────────────────────────────────────────────────────────────────────────────
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# Session-based storage allocator using filesystem symlink indirection.
|
# ⚠️ One mount only in Emby container: /mnt/ram-transcode → /ext-ram-transcode
|
||||||
# ffmpeg resolves the symlink ONCE at session start — existing sessions are never affected.
|
# Do NOT add a static SSD path — Emby will use it independently of the symlink.
|
||||||
# Only new sessions care about where the symlink currently points.
|
# The symlink IS your emergency lever — flip it manually if needed:
|
||||||
#
|
# ln -sfn /mnt/cache/Temp_Storage/Emby/Transcodes /mnt/ram-transcode
|
||||||
# How it works:
|
|
||||||
# ramdisk_setup.sh — run once at array start, creates tmpfs and sets symlink
|
|
||||||
# transcode_manager.sh — every 3 min, monitors usage and manages symlink direction
|
|
||||||
# transcode_cleanup.sh — every 5 min, removes old inactive files from both locations
|
|
||||||
#
|
|
||||||
# ⚠️ Docker mount warning:
|
|
||||||
# Do NOT add a static SSD transcode path as a second volume mount in your Emby container.
|
|
||||||
# If the SSD path is mounted inside the container Emby can see it and will use it
|
|
||||||
# independently of the symlink — breaking symlink-based routing entirely.
|
|
||||||
# The symlink IS your emergency lever — one mount only:
|
|
||||||
# /mnt/ram-transcode → /ext-ram-transcode
|
|
||||||
|
|
||||||
RAMDISK_PATH="/mnt/ramdisk_transcodes" # tmpfs mount point created at array start
|
RAMDISK_PATH="/mnt/ramdisk_transcodes" # tmpfs mount point created at array start
|
||||||
RAMDISK_SIZE="8G" # ceiling — tmpfs only uses RAM actually needed
|
RAMDISK_SIZE="8G" # ceiling — tmpfs only uses RAM actually needed
|
||||||
TRANSCODE_LINK="/mnt/ram-transcode" # symlink Emby points at — location never changes
|
TRANSCODE_LINK="/mnt/ram-transcode" # symlink Emby points at — location never changes
|
||||||
TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/" # SSD fallback location
|
TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/" # SSD fallback location
|
||||||
|
|
||||||
# Usage thresholds in GB — hysteresis gap prevents flip-flop near threshold
|
# Usage thresholds in GB — hysteresis gap between WARN and LOW prevents flip-flop
|
||||||
RAMDISK_WARN_GB=6.8 # flip symlink to SSD at or above this usage
|
RAMDISK_WARN_GB=6.8 # flip symlink to SSD at or above this usage
|
||||||
RAMDISK_LOW_GB=5.5 # flip symlink back to ramdisk when usage drops here
|
RAMDISK_LOW_GB=5.5 # flip symlink back to ramdisk when usage drops here
|
||||||
RAMDISK_SSD_MIN_GB=20 # minimum free GB on SSD required before allowing flip to SSD
|
RAMDISK_SSD_MIN_GB=20 # minimum free GB on SSD required before allowing flip to SSD
|
||||||
|
|
||||||
# Cleanup age thresholds — files must be older than these AND not open by any process
|
# Cleanup age thresholds — files must be older than these AND not open by any process
|
||||||
TRANSCODE_MAX_AGE=20 # minutes before a transcode file is eligible for cleanup
|
TRANSCODE_MAX_AGE=20 # minutes before a transcode file is eligible for cleanup
|
||||||
TRANSCODE_ORPHAN_AGE=30 # minutes before an orphaned file is eligible
|
TRANSCODE_ORPHAN_AGE=30 # minutes before an orphaned file is eligible — extra caution buffer
|
||||||
|
|
||||||
# Flip frequency alert — too many flips per hour may indicate ramdisk needs to be larger
|
# Flip frequency alert — too many flips per hour indicates ramdisk needs to be larger
|
||||||
TRANSCODE_FLIP_WARN=3 # notify if symlink flips this many times in one hour
|
TRANSCODE_FLIP_WARN=3 # notify if symlink flips this many times in one hour
|
||||||
|
|
||||||
# Permissions — must match your Emby container user
|
# Permissions — must match your Emby container user
|
||||||
TRANSCODE_OWNER="nobody:users"
|
TRANSCODE_OWNER="nobody:users"
|
||||||
TRANSCODE_CHMOD="755" # renamed from TRANSCODE_MODE to avoid ambiguity with manager mode
|
TRANSCODE_CHMOD="755" # renamed from TRANSCODE_MODE to avoid ambiguity with manager mode
|
||||||
|
|
||||||
# Operating mode — controls symlink routing behavior
|
# Operating mode — controls symlink routing behavior
|
||||||
# smart — auto-flips between ramdisk and SSD based on usage thresholds (default)
|
# smart — auto-flips between ramdisk and SSD based on usage thresholds (default)
|
||||||
# ramdisk — always uses ramdisk, never flips to SSD regardless of usage
|
# ramdisk — always uses ramdisk, never flips to SSD regardless of usage
|
||||||
@@ -729,28 +856,23 @@ MEDIA_MAINTENANCE_JOBS=(
|
|||||||
# ssd — always uses SSD, never uses ramdisk
|
# ssd — always uses SSD, never uses ramdisk
|
||||||
# useful during ramdisk maintenance, testing, or after a flip issue
|
# useful during ramdisk maintenance, testing, or after a flip issue
|
||||||
# switch to this mode to drain ramdisk sessions gracefully
|
# switch to this mode to drain ramdisk sessions gracefully
|
||||||
TRANSCODE_MANAGER_MODE="smart"
|
TRANSCODE_MANAGER_MODE="smart" # smart | ramdisk | ssd
|
||||||
|
|
||||||
# Emby container check — skips threshold checks when Emby is not running
|
# Emby container check — skips threshold checks when Emby is not running
|
||||||
# Prevents unnecessary symlink flips when no transcoding is happening
|
# Prevents unnecessary symlink flips when no transcoding is happening
|
||||||
TRANSCODE_CHECK_EMBY=true
|
TRANSCODE_CHECK_EMBY=true
|
||||||
TRANSCODE_EMBY_CONTAINER="Emby" # exact Docker container name — case sensitive
|
TRANSCODE_EMBY_CONTAINER="Emby" # exact Docker container name — case sensitive
|
||||||
|
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── MONITORS ───────────────────────────────────────────────────────────────────────────────────
|
# ── MONITORS ──────────────────────────────────────────────────────────────────────────────────
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# Monitoring scripts — watch and report only, never take action.
|
|
||||||
# Lives in Monitor/ folder — distinct from unRAID_Essentials (which acts) and
|
|
||||||
# Docker_Essentials (which manages containers).
|
|
||||||
# These scripts are the data sources for the future plugin dashboard.
|
|
||||||
|
|
||||||
# ━━━ Certificate Monitor ━━━
|
# ━━━ Certificate Monitor ━━━
|
||||||
# Checks SSL cert expiry via direct openssl connection — no NPM dependency.
|
# Checks SSL cert expiry via direct openssl connection — no NPM dependency.
|
||||||
# Reads the actual cert the server is presenting — catches real-world issues API checks miss.
|
# Reads the actual cert the server is presenting — catches real-world issues API checks miss.
|
||||||
# Each domain and subdomain is a separate entry — they have independent certs.
|
# Each domain and subdomain is a separate entry — they have independent certs.
|
||||||
# Add your public-facing domains — uncomment and replace with your actual domains.
|
|
||||||
CERT_MONITOR_DOMAINS=(
|
CERT_MONITOR_DOMAINS=(
|
||||||
"Gmer4Lfe.com"
|
"Gmer4Lfe.com"
|
||||||
"Gmer4Lfe.us"
|
"Gmer4Lfe.us"
|
||||||
)
|
)
|
||||||
CERT_WARN_DAYS=30 # notify warning when cert expires within this many days
|
CERT_WARN_DAYS=30 # notify warning when cert expires within this many days
|
||||||
@@ -763,9 +885,6 @@ CERT_MONITOR_DOMAINS=(
|
|||||||
# Leave BACKUP_VERIFY_SHARES empty to automatically use DAILY_SYNC_SHARES as the target list.
|
# Leave BACKUP_VERIFY_SHARES empty to automatically use DAILY_SYNC_SHARES as the target list.
|
||||||
BACKUP_VERIFY_SHARES=(
|
BACKUP_VERIFY_SHARES=(
|
||||||
# leave empty to use DAILY_SYNC_SHARES automatically
|
# leave empty to use DAILY_SYNC_SHARES automatically
|
||||||
# or specify individual shares to verify:
|
|
||||||
# /mnt/user/Movies
|
|
||||||
# /mnt/user/Tv_Shows
|
|
||||||
)
|
)
|
||||||
BACKUP_VERIFY_SAMPLE=10 # number of files to randomly sample per share per run
|
BACKUP_VERIFY_SAMPLE=10 # number of files to randomly sample per share per run
|
||||||
BACKUP_VERIFY_MIN_SIZE=1M # skip files smaller than this — avoids tiny junk files
|
BACKUP_VERIFY_MIN_SIZE=1M # skip files smaller than this — avoids tiny junk files
|
||||||
@@ -777,49 +896,45 @@ BACKUP_VERIFY_SHARES=(
|
|||||||
SMART_TEMP_WARN=45 # degrees C — warn if drive temperature exceeds this
|
SMART_TEMP_WARN=45 # degrees C — warn if drive temperature exceeds this
|
||||||
SMART_TEMP_CRIT=55 # degrees C — critical if drive temperature exceeds this
|
SMART_TEMP_CRIT=55 # degrees C — critical if drive temperature exceeds this
|
||||||
SMART_IGNORE_DRIVES=(
|
SMART_IGNORE_DRIVES=(
|
||||||
"sda" # uncomment to ignore sda — common choice if sda is your unRAID boot USB
|
"sda" # boot USB — SMART not meaningful on flash drives
|
||||||
)
|
)
|
||||||
|
|
||||||
# ━━━ Bandwidth Monitor ━━━
|
# ━━━ ZFS Memory Snapshot ━━━
|
||||||
# Logs daily rsync transfer totals to a bounded file on /boot/ — minimal flash wear.
|
# Weekly ZFS pool health and memory diagnostic report — informational only, no action taken.
|
||||||
# bandwidth_monitor.sh --log-transfer is called by rsync.sh after each successful sync.
|
# system_watchdog.sh handles threshold-based intervention.
|
||||||
# bandwidth_monitor.sh --report generates the weekly summary standalone.
|
# Output written to ZFS_REPORT_LOG for historical review in addition to console output.
|
||||||
# File stays bounded to BANDWIDTH_LOG_RETENTION lines — old entries auto-purged on each write.
|
# Pools in ZFS_REPORT_IGNORE_POOLS are excluded from reporting — still monitored by unRAID.
|
||||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db"
|
|
||||||
BANDWIDTH_LOG_RETENTION=90 # days to keep — file never grows beyond ~90 lines
|
|
||||||
BANDWIDTH_WARN_GB=50 # flag in reports if a single sync transfer exceeds this GB
|
|
||||||
|
|
||||||
# ━━━ ZFS Memory Snapshot ━━━
|
|
||||||
ZFS_REPORT_LOG="/var/log/zfs-weekly-health.log"
|
ZFS_REPORT_LOG="/var/log/zfs-weekly-health.log"
|
||||||
ZFS_REPORT_ARC_WARN_PCT=90
|
ZFS_REPORT_ARC_WARN_PCT=90 # warn in report if ARC utilization above this %
|
||||||
ZFS_REPORT_FREE_WARN_GB=10
|
ZFS_REPORT_FREE_WARN_GB=10 # warn in report if free RAM drops below this GB
|
||||||
ZFS_REPORT_AVAIL_WARN_GB=20
|
ZFS_REPORT_AVAIL_WARN_GB=20 # warn in report if available RAM drops below this GB
|
||||||
ZFS_REPORT_DOCKER_TOP=10
|
ZFS_REPORT_DOCKER_TOP=10 # number of top Docker memory users to show in report
|
||||||
|
|
||||||
# Pools to exclude from health reporting — still monitored by unRAID but skipped in report
|
|
||||||
# Useful for pools that are expected to be heavily used or are managed separately
|
|
||||||
ZFS_REPORT_IGNORE_POOLS=(
|
ZFS_REPORT_IGNORE_POOLS=(
|
||||||
"disk10" # Docker overlay storage — high usage is normal
|
# Pools excluded from health reporting — expected to run at high usage
|
||||||
"disk9" # Cache pool — usage varies widely, not meaningful to report
|
# All pools still monitored by unRAID regardless of this list
|
||||||
|
"disk10"
|
||||||
|
"disk9"
|
||||||
"disk8"
|
"disk8"
|
||||||
"disk6"
|
"disk6"
|
||||||
"disk5"
|
"disk5"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# ━━━ Bandwidth Monitor ━━━
|
||||||
|
# Called by rsync.sh after each sync — one bounded write per run, minimal flash wear.
|
||||||
|
# Log format: YYYY-MM-DD|HH:MM|profile|duration_seconds|status — version-proof
|
||||||
|
# File stays bounded to BANDWIDTH_LOG_RETENTION days — old entries auto-purged on write.
|
||||||
|
BANDWIDTH_LOG="/boot/config/bandwidth_history.db"
|
||||||
|
BANDWIDTH_LOG_RETENTION=90 # days to keep — file never grows beyond ~90 lines
|
||||||
|
BANDWIDTH_WARN_GB=50 # flag in reports if a single sync transfer exceeds this GB
|
||||||
|
|
||||||
# ━━━ Health Digest ━━━
|
# ━━━ Health Digest ━━━
|
||||||
# Aggregated system health summary from across the ecosystem.
|
# Aggregated system health summary — reads existing state files, no new writes to flash.
|
||||||
# Reads existing state files — no new writes to flash drive.
|
|
||||||
#
|
|
||||||
# Three profiles — switch by changing DIGEST_PROFILE, no cron changes needed:
|
# Three profiles — switch by changing DIGEST_PROFILE, no cron changes needed:
|
||||||
# always — sends every run (schedule daily = daily digest, weekly = weekly digest)
|
# always — sends every run
|
||||||
# smart — sends only if findings worth reporting (intelligent quiet operation)
|
# smart — sends only if findings worth reporting
|
||||||
# weekly — sends once per week on DIGEST_DAY only, silent all other days
|
# weekly — sends once per week on DIGEST_DAY only
|
||||||
#
|
|
||||||
# Data sources (reads only — no writes):
|
|
||||||
# Transcode ramdisk state, container watchdog strikes, system watchdog strikes,
|
|
||||||
# failover state, container skip list, bandwidth history, SSL cert days remaining
|
|
||||||
DIGEST_PROFILE="weekly" # always | smart | weekly
|
DIGEST_PROFILE="weekly" # always | smart | weekly
|
||||||
DIGEST_DAY="Sunday" # day name for weekly profile — must match date +%A output
|
DIGEST_DAY="Sunday" # must match date +%A output
|
||||||
|
|
||||||
# Smart profile triggers — set true to send digest when this condition is found
|
# Smart profile triggers — set true to send digest when this condition is found
|
||||||
DIGEST_SMART_ON_WATCHDOG=true # send if any watchdog strikes are active
|
DIGEST_SMART_ON_WATCHDOG=true # send if any watchdog strikes are active
|
||||||
@@ -828,29 +943,19 @@ ZFS_REPORT_IGNORE_POOLS=(
|
|||||||
DIGEST_SMART_ON_BANDWIDTH=true # send if any transfer exceeded BANDWIDTH_WARN_GB
|
DIGEST_SMART_ON_BANDWIDTH=true # send if any transfer exceeded BANDWIDTH_WARN_GB
|
||||||
|
|
||||||
# ━━━ Emby Session Report ━━━
|
# ━━━ Emby Session Report ━━━
|
||||||
# Weekly Emby usage report via API — no persistent writes, queries fresh each run.
|
# Requires API key from Emby Settings → API Keys in the Emby WebUI.
|
||||||
# Shows active streams, library counts, transcode vs direct play ratio.
|
# No persistent writes — queries fresh each run.
|
||||||
# Requires an API key from Emby Settings → API Keys in the Emby WebUI.
|
|
||||||
EMBY_URL="http://localhost:8096"
|
EMBY_URL="http://localhost:8096"
|
||||||
EMBY_API_KEY="0c27448d93a7431f9ac63569f7655829" # paste your Emby API key here
|
EMBY_API_KEY="0c27448d93a7431f9ac63569f7655829"
|
||||||
EMBY_REPORT_DAYS=7 # number of days to include in the report period
|
EMBY_REPORT_DAYS=7 # number of days to include in the report period
|
||||||
EMBY_REPORT_TOP_N=10 # number of top content items to show in report
|
EMBY_REPORT_TOP_N=10 # number of top content items to show in report
|
||||||
|
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# ── SYSTEM WATCHDOG ───────────────────────────────────────────────────────────────────────────
|
# ── SYSTEM WATCHDOG ───────────────────────────────────────────────────────────────────────────
|
||||||
# ==============================================================================================
|
# ==============================================================================================
|
||||||
# Last line of defense — reboots the system cleanly when it is about to become unstable.
|
# Last line of defense — reboots cleanly when system is about to become unstable.
|
||||||
# Runs every 15 minutes via cron. Works alongside docker_watchdog.sh:
|
# Strike system: sustained threshold hits trigger reboot — single spikes ignored.
|
||||||
# docker_watchdog.sh — container level, minimal disruption, tries to self-heal first
|
# Reboot loop protection: shuts down instead if reboot limit hit in window.
|
||||||
# system_watchdog.sh — system level, last resort, reboots when healing has failed
|
|
||||||
#
|
|
||||||
# Strike system: sustained threshold hits trigger reboot — single spikes are ignored.
|
|
||||||
# Each check that exceeds its threshold adds a strike. Strikes reset when recovered.
|
|
||||||
# When strike limit is hit the reboot sequence begins.
|
|
||||||
#
|
|
||||||
# Reboot loop protection: tracks reboot timestamps on /boot/ (survives reboots).
|
|
||||||
# If the server reboots too many times in the window it shuts down instead — a reboot
|
|
||||||
# loop means something is fundamentally wrong that a reboot is not fixing.
|
|
||||||
|
|
||||||
# ━━━ State Files ━━━
|
# ━━━ State Files ━━━
|
||||||
# Strike counts reset on reboot — /tmp is correct (fresh start after each reboot)
|
# Strike counts reset on reboot — /tmp is correct (fresh start after each reboot)
|
||||||
|
|||||||
Reference in New Issue
Block a user