Files
Varaverk/Monitors/Manual-Monitors.md
Gmer4Lfe 7a8fbbe06a Correct stale references across the remaining folder docs
Three renamed or dissolved names were still cited, and three scheduled scripts had no
entry in their folder's table.
2026-08-01 23:07:50 -04:00

587 lines
19 KiB
Markdown

# ━━━━━ MONITORS — Manual ━━━━━
Config reference, procedures, operational workflows.
For overview see README-Monitors.md. For per-script detail see script headers.
---
## ━━━ CERT MONITOR ━━━
### Domain Configuration
Each domain and subdomain has an independent TLS certificate and must be listed
separately. `cert_monitor.sh` makes one openssl connection per entry.
```bash
# host1.conf
HOST1_CERT_MONITOR_DOMAINS=(
"Gmer4Lfe.com"
"Gmer4Lfe.us"
# "auth.Gmer4Lfe.com" # add subdomains separately
# "cloud.Gmer4Lfe.com"
)
# host2.conf
HOST2_CERT_MONITOR_DOMAINS=(
"Jayred365.com"
# "auth.Jayred365.com"
)
```
Each server monitors its own domains. `detect_hosts()` aliases the correct list
to `CERT_MONITOR_DOMAINS` based on `MY_ID`.
### Threshold Configuration
```bash
# master.conf
CERT_WARN_DAYS=30 # warn this many days before expiry — time to investigate
CERT_CRIT_DAYS=7 # critical alert — action needed now
CERT_TIMEOUT=10 # seconds per domain before declaring FAILED
```
**Sizing guidance:**
- `CERT_WARN_DAYS=30` gives a month to investigate and renew. If Let's Encrypt
auto-renewal is working, you'll only see warnings when renewal breaks.
- `CERT_CRIT_DAYS=7` — at 7 days remaining, manual action is needed today.
- `CERT_TIMEOUT=10` — sufficient for typical public domains; increase for slow
DNS resolution or high-latency connections.
### Notification Behavior
| State | Condition | Action |
|-------|-----------|--------|
| HEALTHY | > CERT_WARN_DAYS remaining | Silent — no notification |
| WARNING | ≤ CERT_WARN_DAYS remaining | One notification listing all WARNING domains |
| CRITICAL | ≤ CERT_CRIT_DAYS remaining | Separate notification listing all CRITICAL domains |
| FAILED | Could not connect or parse cert | Notification — treat as critical |
---
## ━━━ SMART HEALTH ━━━
### Drive Ignore List
```bash
# host1.conf
HOST1_SMART_IGNORE_DRIVES=(
"sda" # boot USB flash drive — no meaningful SMART data
)
# host2.conf
HOST2_SMART_IGNORE_DRIVES=(
"sda" # boot USB flash drive
)
```
unRAID typically boots from a USB flash drive that appears as `sda`. Flash drives
either don't support SMART or report values that are meaningless for health assessment.
Add any drive that produces spurious SMART data here.
### Temperature Configuration
```bash
# master.conf — fallback values if dynamix.cfg is not found
SMART_TEMP_WARN=45 # °C
SMART_TEMP_CRIT=55 # °C
```
`smart_health.sh` reads temperature thresholds directly from
`/boot/config/plugins/dynamix/dynamix.cfg` at runtime, using the same hot/max/hotssd/maxssd
values as unRAID's dashboard. The `master.conf` values are only used as a fallback
if `dynamix.cfg` is not found (e.g., running outside of unRAID).
### Attribute Reference
| Attribute | Threshold | Meaning |
|-----------|-----------|---------|
| `Reallocated_Sector_Ct` | Any > 0 is warning | Drive found bad sectors and swapped in spares. 0 = healthy. Count growing = degrading. |
| `Current_Pending_Sector` | Any > 0 is warning | Sectors suspected bad, not yet confirmed. May recover on next read, may escalate. Watch it. |
| `Offline_Uncorrectable` | Any > 0 is critical | Could not correct during offline tests. No spares. Data loss risk. |
| Overall SMART status | FAILED = immediate alert | Drive's own self-assessment. FAILED = get data off now. |
| `Temperature_Celsius` | vs dynamix.cfg thresholds | Sustained high temp shortens drive life significantly. |
| `Power_On_Hours` | Informational | Drive age in days — useful context for other attributes. |
---
## ━━━ BACKUP VERIFY ━━━
### Share Configuration
```bash
# host1.conf
HOST1_BACKUP_VERIFY_SHARES=(
# empty — uses HOST1_DAILY_SYNC_SHARES automatically
# "/mnt/user/Movies" # override to check specific shares only
)
```
Leave `HOST*_BACKUP_VERIFY_SHARES` empty and `backup_verify.sh` will automatically
verify the same shares configured for daily rsync. Only populate it if you want to
verify a different set.
### Sample Configuration
```bash
# master.conf
BACKUP_VERIFY_SAMPLE=10 # random files sampled per share per run
BACKUP_VERIFY_MIN_SIZE="1M" # skip files smaller than this
```
**Sizing guidance:**
- `BACKUP_VERIFY_SAMPLE=10` is a spot check — catches systematic hardware problems
while running in minutes. Increase to 25-50 for deeper confidence on a large library;
decrease to 5 if the run is too slow against many shares.
- `BACKUP_VERIFY_MIN_SIZE="1M"` — tiny files (NFO, thumbnails, subtitles) have negligible
corruption risk and each one adds an SSH round-trip. Skip them.
### Result States
| State | Meaning | Action |
|-------|---------|--------|
| MATCH | Checksums identical on both servers | Correctly mirrored — no action |
| MISMATCH | File exists on both, checksums differ | Sync failure or corruption — investigate |
| MISSING | File exists locally, not on remote | Not yet synced or deleted on remote — may be normal |
---
## ━━━ BANDWIDTH MONITOR ━━━
### Log Configuration
```bash
# master.conf
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db" # survives reboots
BANDWIDTH_LOG_RETENTION=90 # days — file stays bounded, never grows unbounded
BANDWIDTH_WARN_GB=50 # flag transfers or daily totals exceeding this
```
**Why `$DATA_DIR`**: The log needs to survive reboots to build a useful history.
`$DATA_DIR` (`$SCRIPTS_DIR/data/`) is on the boot device (internal) or appdata (flash),
both of which survive reboots. The log is bounded by `BANDWIDTH_LOG_RETENTION` so it never
grows unbounded.
**`BANDWIDTH_WARN_GB`**: Set to a value that represents "unexpectedly large" for your
setup. If your typical daily sync transfers 5-10GB, 50GB would flag a full-library resync.
Adjust based on your share sizes.
### Log Format
One line per transfer — do not edit manually:
```
YYYY-MM-DD|HH:MM|profile|duration_seconds|status|bytes_transferred
# Example entries:
2026-04-14|01:23|critical-data|143|success|2847362048
2026-04-14|01:31|arrs_stack|287|success|891234567
2026-04-14|02:15|movies|1847|failed|0
```
### How rsync.sh Calls bandwidth_monitor.sh
`rsync.sh` automatically calls `bandwidth_monitor.sh --log-transfer` after each sync
profile completes. You never need to call `--log-transfer` manually. The call:
```bash
bandwidth_monitor.sh --log-transfer "movies" 1847 "failed" 0
# profile secs status bytes
```
---
## ━━━ WEEKLY HEALTH DIGEST ━━━
### Profile Configuration
```bash
# master.conf
DIGEST_PROFILE="weekly" # always | smart | weekly
DIGEST_DAY="Sunday" # for weekly profile — must match `date +%A` output
```
The cron schedule is always `0 8 * * *` (8am daily). Change only `DIGEST_PROFILE`
to switch behavior — no cron edit needed.
**Profile guide:**
| Profile | When to use | Result |
|---------|-------------|--------|
| `always` | You want a daily check-in regardless of system state | Notification every morning |
| `smart` | Quiet operation, only alert on real issues | Silent when healthy, notification when something needs attention |
| `weekly` | One weekly summary is enough | Single notification on `DIGEST_DAY`, silent all other days |
### Smart Profile Triggers
```bash
# master.conf — each independently toggleable
DIGEST_SMART_ON_WATCHDOG=true # send if any active watchdog strikes
DIGEST_SMART_ON_FALLBACK=true # send if fallback state is not NORMAL
DIGEST_SMART_ON_CERT_WARN=true # send if any cert within CERT_WARN_DAYS
DIGEST_SMART_ON_BANDWIDTH=true # send if any transfer exceeded BANDWIDTH_WARN_GB
```
Turn any trigger off by setting it to `false`. The digest still runs and collects
data — it just doesn't trigger a notification for that condition.
### Data Sources
`weekly_health_digest.sh` reads (but never writes):
| State File | Source | What It Shows |
|-----------|--------|---------------|
| `FALLBACK_STATE_FILE` | `Fallback/fallback.sh` | Current fallback state (NORMAL/FALLBACK/etc.) |
| `DOCKER_WATCHDOG_FAILED_FILE` | `Watchdogs/docker_watchdog.sh` | Container skip list — needs human attention |
| `WATCHDOG_STATE_FILE` | `Watchdogs/docker_watchdog.sh` | Active container strike counts |
| `SYS_WATCHDOG_STATE_FILE` | `Watchdogs/stability_watchdog.sh` | Active system watchdog strikes |
| `BANDWIDTH_LOG` | `bandwidth_monitor.sh` | Yesterday's transfer history |
| `TRANSCODE_DAILY_LOG` | `Transcodes/` | Weekly transcode statistics |
| `CERT_MONITOR_DOMAINS` | live openssl check | Current cert status per domain |
---
## ━━━ EMBY SESSION REPORT ━━━
### Emby Credentials
```bash
# host1.conf
HOST1_EMBY_URL="http://192.168.50.2:8096"
HOST1_EMBY_API_KEY="0c27448d93a7431f9ac63569f7655829"
# host2.conf
HOST2_EMBY_URL="http://192.168.50.3:8096"
HOST2_EMBY_API_KEY="<host2_api_key>"
```
To generate an API key: Emby UI → Settings → API Keys → New API Key.
Give it a descriptive name (e.g., `varaverk`). The key is only shown once.
### Report Configuration
```bash
# master.conf
EMBY_REPORT_DAYS=7 # report period in days
EMBY_REPORT_TOP_N=10 # top N content items to show
```
---
## ━━━ SYSTEM TUNING MONITOR ━━━
### Threshold Configuration
```bash
# master.conf
INOTIFY_WARN_PCT=80 # warn if inotify instances above 80% of kernel limit
PHP_FPM_WARN_PCT=80 # warn if php-fpm workers above 80% of max
TUNING_MONITOR_LOG="$DATA_DIR/tuning_monitor.db"
TUNING_LOG_RETENTION=30 # days — log stays bounded
```
**`INOTIFY_WARN_PCT`**: 80% leaves headroom before kernel-level exhaustion. If you
regularly see warnings at 80% and your system is stable, increase to 90%. Don't set
higher than 90% — at 100% utilisation new inotify watches silently fail.
**`PHP_FPM_WARN_PCT`**: 80% means the WebGUI is using most of its workers. At 100%
new requests queue (WebGUI feels sluggish) or time out.
**`PHP_MAX_CHILDREN`**: Set by `php_fpm_max_children.sh` in `System_Essentials/` — do
not set manually here.
### Reading the Weekly Digest Data
`system_tuning_monitor.sh` feeds data into `weekly_health_digest.sh`. The digest
shows for the week:
- inotify: peak utilisation, average utilisation, number of snapshots that hit the warning threshold
- php-fpm: same
A few warning snapshots per week is normal. A rising peak or many warnings per week
means the limits should be adjusted — use `inotify_tuning.sh` or `php_fpm_max_children.sh`
in `System_Essentials/`.
---
## ━━━ ZFS MEMORY SNAPSHOT ━━━
### Pool Ignore List
```bash
# host1.conf
HOST1_ZFS_REPORT_IGNORE_POOLS=(
"disk10" # JBOD member — high usage expected, exclude from report noise
"disk9"
"disk8"
"disk6"
"disk5"
)
# host2.conf
HOST2_ZFS_REPORT_IGNORE_POOLS=(
# list host2's JBOD members here
)
```
In a JBOD ZFS configuration each disk appears as its own pool. These single-disk
pools normally show high capacity utilisation by design — including them would produce
constant warnings. The ignore list removes them from the report while unRAID's built-in
monitoring continues to watch them.
### Threshold Configuration
```bash
# master.conf
ZFS_REPORT_LOG="/var/log/zfs-weekly-health.log" # tmpfs — resets on reboot
ZFS_REPORT_ARC_WARN_PCT=90 # warn if ARC using more than 90% of its max
ZFS_REPORT_FREE_WARN_GB=10 # warn if free RAM below 10GB
ZFS_REPORT_AVAIL_WARN_GB=20 # warn if available RAM below 20GB
ZFS_REPORT_DOCKER_TOP=10 # top 10 Docker containers by memory
```
**`ZFS_REPORT_LOG`**: Written to `/var/log/` (tmpfs) — resets on reboot. This is
intentional — the log is for the current uptime's week-over-week comparison, not
long-term history. It avoids flash drive writes entirely.
---
## ━━━ MESH MONITOR ━━━
At-a-glance view of the mesh — who the members are, what each node covers for the
others when they go down, and the current partnership state. Read-only. Parses conf
files directly — no SSH, no API calls, no sensitive data.
### Sections
```
MEMBERS — server name, owner, contact email for each HOST*
COVERAGE — what each host covers for others, by tier
PARTNERSHIP — enabled/disabled, current owner, sync interval
```
### When to Use
```
Onboarding a new mesh member — paste the output to show them the current topology
After adding a new container to a fallback tier — verify it appears in the right tier
When troubleshooting fallback — confirm coverage arrays are configured correctly
```
### Usage
```bash
mesh_monitor.sh # full mesh overview
mesh_monitor.sh --log # verbose output
mesh_monitor.sh --status # show config and exit
```
### Configuration
No dedicated config keys. Reads `HOST*_OWNER`, `HOST*_OWNER_EMAIL`, and
`FALLBACK_HOST*_COVERS_HOST*_TIER*` arrays from `host*.conf`. Scales automatically
as new nodes join — add a new `HOST3_*` block and it appears in output without
changing the script.
---
## ━━━ FULL CONFIGURATION REFERENCE ━━━
### master.conf
```bash
# cert_monitor.sh
CERT_WARN_DAYS=30
CERT_CRIT_DAYS=7
CERT_TIMEOUT=10
# smart_health.sh
SMART_TEMP_WARN=45
SMART_TEMP_CRIT=55
# backup_verify.sh
BACKUP_VERIFY_SAMPLE=10
BACKUP_VERIFY_MIN_SIZE="1M"
# bandwidth_monitor.sh
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db"
BANDWIDTH_LOG_RETENTION=90
BANDWIDTH_WARN_GB=50
# weekly_health_digest.sh
DIGEST_PROFILE="weekly"
DIGEST_DAY="Sunday"
DIGEST_SMART_ON_WATCHDOG=true
DIGEST_SMART_ON_FALLBACK=true
DIGEST_SMART_ON_CERT_WARN=true
DIGEST_SMART_ON_BANDWIDTH=true
# emby_session_report.sh
EMBY_REPORT_DAYS=7
EMBY_REPORT_TOP_N=10
# system_tuning_monitor.sh
INOTIFY_WARN_PCT=80
PHP_FPM_WARN_PCT=80
TUNING_MONITOR_LOG="$DATA_DIR/tuning_monitor.db"
TUNING_LOG_RETENTION=30
# zfs_memory_snapshot.sh
ZFS_REPORT_LOG="/var/log/zfs-weekly-health.log"
ZFS_REPORT_ARC_WARN_PCT=90
ZFS_REPORT_FREE_WARN_GB=10
ZFS_REPORT_AVAIL_WARN_GB=20
ZFS_REPORT_DOCKER_TOP=10
```
### host*.conf
```bash
# cert_monitor.sh
HOST1_CERT_MONITOR_DOMAINS=("domain1.com" "domain2.com")
HOST2_CERT_MONITOR_DOMAINS=("domain3.com")
# smart_health.sh
HOST1_SMART_IGNORE_DRIVES=("sda")
HOST2_SMART_IGNORE_DRIVES=("sda")
# backup_verify.sh (leave empty to use DAILY_SYNC_SHARES)
HOST1_BACKUP_VERIFY_SHARES=()
HOST2_BACKUP_VERIFY_SHARES=()
# emby_session_report.sh
HOST1_EMBY_URL="http://192.168.50.2:8096"
HOST1_EMBY_API_KEY="<key>"
HOST2_EMBY_URL="http://192.168.50.3:8096"
HOST2_EMBY_API_KEY="<key>"
# zfs_memory_snapshot.sh
HOST1_ZFS_REPORT_IGNORE_POOLS=("disk10" "disk9" "disk8" "disk6" "disk5")
HOST2_ZFS_REPORT_IGNORE_POOLS=()
```
---
## ━━━ RECOMMENDED SCHEDULE ━━━
```bash
# Daily
0 8 * * * weekly_health_digest.sh # DIGEST_PROFILE controls notify frequency
# Every 6 hours — background snapshot
0 */6 * * * system_tuning_monitor.sh # inotify + php-fpm utilisation tracking
# Sunday morning block — runs after nightly maintenance completes (~3am)
# By 6am the weekly restarts, log clears, and media maintenance have finished.
# Monitors see a freshly maintained system.
0 6 * * 0 zfs_memory_snapshot.sh # ZFS + memory — first, before everything
0 7 * * 0 smart_health.sh # drive SMART health
0 9 * * 0 cert_monitor.sh # SSL cert expiry
0 10 * * 0 backup_verify.sh # rsync mirror integrity
0 11 * * 0 emby_session_report.sh # Emby streaming usage
0 11 * * 0 bandwidth_monitor.sh --report # rsync transfer summary
# Automatic — no scheduling needed
# bandwidth_monitor.sh --log-transfer is called by rsync.sh after each sync.
```
---
## ━━━ OUTPUT TIERS ━━━
All scripts use a two-tier output model: `echo` lines are always visible; `log`
lines only appear when `--log` is passed.
Most monitors are one-shot scripts that run on a schedule. Without `--log`, phase
headers, step conclusions (pre-flight passed, snapshot written, digest sent), and
the final status line are visible. Per-drive, per-cert, per-file, and per-domain
detail lines inside loops are suppressed.
`weekly_health_digest.sh` with `profile=smart` exits silently (one `echo` line)
when there are no findings worth reporting — no noise on healthy weeks.
---
## ━━━ FLAG REFERENCE ━━━
All monitor scripts support the same four flags:
### --dry-run
Runs checks without sending any notifications. Useful for:
- First-time setup validation (verify the script finds your domains/drives/shares)
- Ad-hoc health checks without triggering notification noise
- Testing after configuration changes
```bash
cert_monitor.sh --dry-run # check all certs, show results, no notify
smart_health.sh --dry-run # show which drives would be checked
backup_verify.sh --dry-run # show which files would be sampled
bandwidth_monitor.sh --dry-run # generate report without sending
weekly_health_digest.sh --dry-run # generate digest, no notify regardless of profile
zfs_memory_snapshot.sh --dry-run # console output only, no log write
system_tuning_monitor.sh --dry-run # measure and show, no log write
emby_session_report.sh --dry-run # test connectivity, generate report, no notify
```
### --status
Shows current configuration and exits without running checks. Use to verify
configuration is loaded correctly after editing `master.conf` or `host*.conf`.
```bash
cert_monitor.sh --status # domain list, CERT_WARN_DAYS, CERT_CRIT_DAYS, timeout
smart_health.sh --status # ignore list, temperature thresholds
backup_verify.sh --status # share list, sample size, min file size
bandwidth_monitor.sh --status # log path, retention, warn threshold, log stats
weekly_health_digest.sh --status # profile, DIGEST_DAY, smart trigger settings
zfs_memory_snapshot.sh --status # pool ignore list, memory thresholds
system_tuning_monitor.sh --status # warn thresholds, log path, retention
emby_session_report.sh --status # Emby URL, API key (masked), report period
```
### --log
Verbose per-item output. Every domain/drive/file/section logs its result explicitly
instead of being silent on pass. Use when investigating or after a configuration change.
### Normal run (no flags)
Silent on healthy — only problems produce output or notifications.
`weekly_health_digest.sh` is the exception: it always generates a digest, but
`DIGEST_PROFILE` controls whether a notification is sent.
---
## ━━━ ADDING A NEW MONITOR ━━━
All monitor scripts share the same pattern:
```bash
#!/bin/bash
# [standard header with PURPOSE / SAFEGUARDS / CONFIGURATION / RUNTIME MODES]
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../load_config.sh"
SILENT_MODE=false # monitor script — output is the point
parse_args "$@"
# root check if needed
# tool validation (platform_require_cmd)
acquire_lock
detect_hosts # if host-specific config needed
# [status block]
# [dry-run block]
# do checks
# notify() only on problems — silent on healthy
```
Key properties:
- `SILENT_MODE=false` — monitors produce output
- `acquire_lock` — prevent duplicate runs
- `notify()` only on problems — never notify on healthy results
- `--dry-run` skips `notify()` calls entirely
- `--status` shows config and exits before any checks run
- `--log` enables verbose per-item output