18 lines
93 KiB
JSON
18 lines
93 KiB
JSON
{"parentUuid":null,"isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":"Read the following files and report any discrepancies between the documentation and the actual script headers. The goal is to find places where the README or Manual says something that no longer matches what the script header says (wrong schedule, wrong config vars, wrong behavior description, missing scripts, extra scripts, etc.).\n\nFiles to read:\n1. /boot/config/plugins/varaverk/Watchdogs/README-Watchdogs.md\n2. /boot/config/plugins/varaverk/Watchdogs/Manual-Watchdogs.md\n3. /boot/config/plugins/varaverk/Watchdogs/docker_watchdog.sh (header only — top ~120 lines)\n4. /boot/config/plugins/varaverk/Watchdogs/system_watchdog.sh (header only — top ~120 lines)\n5. /boot/config/plugins/varaverk/Watchdogs/resource_watchdog.sh (header only — top ~120 lines)\n6. /boot/config/plugins/varaverk/Watchdogs/stability_watchdog.sh (header only — top ~120 lines)\n\nReport: list each discrepancy as \"FILE: what the doc says vs what the header says\". If everything matches, say so. Be specific — quote the conflicting text. Keep the report under 400 words."},"uuid":"623a52d2-31ec-4965-9b3b-5a74e247a871","timestamp":"2026-06-14T05:51:46.983Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"623a52d2-31ec-4965-9b3b-5a74e247a871","isSidechain":true,"agentId":"a66bc873b174a3b68","attachment":{"type":"deferred_tools_delta","addedNames":["EnterWorktree","ExitWorktree","Monitor","TaskStop","WebFetch","WebSearch","mcp__claude_ai_Google_Drive__authenticate","mcp__claude_ai_Google_Drive__complete_authentication"],"addedLines":["EnterWorktree","ExitWorktree","Monitor","TaskStop","WebFetch","WebSearch","mcp__claude_ai_Google_Drive__authenticate","mcp__claude_ai_Google_Drive__complete_authentication"],"removedNames":[],"readdedNames":[]},"type":"attachment","uuid":"d44a812a-b02e-4f76-b676-d2008d7e3cf3","timestamp":"2026-06-14T05:51:46.993Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"d44a812a-b02e-4f76-b676-d2008d7e3cf3","isSidechain":true,"agentId":"a66bc873b174a3b68","attachment":{"type":"skill_listing","content":"- update-config: Use this skill to configure the Claude Code harness via settings.json. Automated behaviors (\"from now on when X\", \"each time X\", \"whenever X\", \"before/after X\") require hooks configured in settings.json - the harness executes these, not Claude, so memory/preferences cannot fulfill them. Also use for: permissions (\"allow X\", \"add permission\", \"move permission to\"), env vars (\"set X=Y\"), hook troubleshooting, or any changes to settings.json/settings.local.json files. Examples: \"allow npm commands\", \"add bq permission to global settings\", \"move permission to user settings\", \"set DEBUG=true\", \"when claude stops show X\". For simple settings like theme/model, suggest the /config command.\n- keybindings-help: Use when the user wants to customize keyboard shortcuts, rebind keys, add chord bindings, or modify ~/.claude/keybindings.json. Examples: \"rebind ctrl+s\", \"add a chord shortcut\", \"change the submit key\", \"customize keybindings\".\n- verify: Verify that a code change actually does what it's supposed to by running the app and observing behavior. Use when asked to verify a PR, confirm a fix works, test a change manually, check that a feature works, or validate local changes before pushing.\n- code-review: Review the current diff for correctness bugs and reuse/simplification/efficiency cleanups at the given effort level (low/medium: fewer, high-confidence findings; high→max: broader coverage, may include uncertain findings; ultra: deep multi-agent review in the cloud). Pass --comment to post findings as inline PR comments, or --fix to apply the findings to the working tree after the review.\n- simplify: Review the changed code for reuse, simplification, efficiency, and altitude cleanups, then apply the fixes. Quality only — it does not hunt for bugs; use /code-review for that.\n- fewer-permission-prompts: Scan your transcripts for common read-only Bash and MCP tool calls, then add a prioritized allowlist to project .claude/settings.json to reduce permission prompts.\n- loop: Run a prompt or slash command on a recurring interval (e.g. /loop 5m /foo). Omit the interval to let the model self-pace. - When the user wants to set up a recurring task, poll for status, or run something repeatedly on an interval (e.g. \"check the deploy every 5 minutes\", \"keep running /babysit-prs\"). Do NOT invoke for one-off tasks.\n- schedule: Create, update, list, or run scheduled cloud agents (routines) that execute on a cron schedule. - When the user wants to schedule a recurring cloud agent, set up automated tasks, create a cron job for Claude Code, or manage their scheduled agents/routines. Also use when the user wants a one-time scheduled run (\"run this once at 3pm\", \"remind me to check X tomorrow\").\n- claude-api: Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration.\nTRIGGER — read BEFORE opening the target file; don't skip because it \"looks like a one-liner\" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens).\nSKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).\n- run: Launch and drive this project's app to see a change working. Use when asked to run, start, or screenshot the app, or to confirm a change works in the real app (not just tests). First looks for a project skill that already covers launching the app; otherwise falls back to built-in patterns per project type (CLI, server, TUI, Electron, browser-driven, library).\n- init: Initialize a new CLAUDE.md file with codebase documentation\n- review: Review a pull request\n- security-review: Complete a security review of the pending changes on the current branch","skillCount":13,"isInitial":true,"names":["update-config","keybindings-help","verify","code-review","simplify","fewer-permission-prompts","loop","schedule","claude-api","run","init","review","security-review"]},"type":"attachment","uuid":"2062ae6f-23ff-45bd-946a-b5d943aa7dd4","timestamp":"2026-06-14T05:51:46.996Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"2062ae6f-23ff-45bd-946a-b5d943aa7dd4","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"text","text":"I'll read all the files and compare the documentation with the script headers for discrepancies."}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"f2348d2a-b61a-49e9-8ddb-fd7cf56d9de6","timestamp":"2026-06-14T05:51:48.550Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"f2348d2a-b61a-49e9-8ddb-fd7cf56d9de6","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01L7N5aJeycDdTaRqzkZ1ep9","name":"Read","input":{"file_path":"/boot/config/plugins/varaverk/Watchdogs/README-Watchdogs.md"},"caller":{"type":"direct"}}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"1ccbfbf9-86f2-47c3-8e11-ab7bd482c972","timestamp":"2026-06-14T05:51:48.588Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"1ccbfbf9-86f2-47c3-8e11-ab7bd482c972","isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01L7N5aJeycDdTaRqzkZ1ep9","type":"tool_result","content":"1\t# ━━━━━ WATCHDOGS ━━━━━\n2\t\n3\t**Four tiers that run every 15 minutes through `watchdog_orchestrator.sh`,\n4\teach with a clear lane:** reduce system pressure → heal containers → check system components →\n5\treboot if nothing else worked. The orchestrator calls them in order, once per cron cycle.\n6\tSystem component checks (storage, WebGUI, network) run inside a thin `system_watchdog.sh`\n7\torchestrator that can also be run standalone.\n8\t\n9\t---\n10\t\n11\t## ━━━ THE PROBLEM THAT BUILT THIS ━━━\n12\t\n13\t**Container Memory Leaks Going Undetected for Days**\n14\tEmby's transcode session handling occasionally leaks memory. SABnzbd's Python process\n15\texpands slowly across downloads. Neither crashes dramatically — they just consume more\n16\tRAM until the system starts swapping. Docker reports both containers as `Up 14 days`.\n17\tNothing alerts. By the time someone notices, the system has been degraded for hours.\n18\t\n19\tFix: `docker_watchdog.sh` — hard per-container memory ceilings. When a container\n20\texceeds its limit the watchdog restarts it immediately. No strikes, no waiting.\n21\tA memory leak is not a transient spike.\n22\t\n23\t**Containers That Look Running But Aren't Responding**\n24\tDocker reports a container as `Up` while its application layer has been frozen for\n25\thours. The reverse proxy forwards traffic to a service that returns nothing. Users see\n26\ta broken page. Docker sees a healthy container.\n27\t\n28\tFix: `docker_watchdog.sh` — HTTP health checks on the actual service port every cycle.\n29\tTwo consecutive non-responses trigger a restart. Process running and service responding\n30\tare not the same thing.\n31\t\n32\t**System Pressure Causing Docker Watchdog to Undo Itself**\n33\tResource pressure builds. RAM drops. docker_watchdog.sh tries to restart a container\n34\tinto a system that's already swapping — the restarted container fails immediately\n35\tand goes on the skip list. The real problem (RAM pressure) is never addressed.\n36\t\n37\tFix: `resource_watchdog.sh` runs first in the orchestrator sequence. At Level 1 it\n38\tthrottles downloaders. At Level 2 it pauses non-critical containers. At Level 3 it\n39\tstops heavy services and signals docker_watchdog to defer all restarts. By the time\n40\tdocker_watchdog runs, the system has breathing room to actually heal.\n41\t\n42\t**Runaway Log Files Filling a Pool Before Anyone Notices**\n43\tA game server container was offline for a year, restarted for a weekend, and wrote\n44\t130GB of logs to the appdata pool. The pool grew 13% in days. No alert fired —\n45\tnothing was watching for growth at the data level, only at the container level.\n46\t\n47\tFix: `storage_watchdog.sh` — growth-rate scan of every container's appdata directory\n48\tevery cycle. No per-container configuration required. Runaway growth gets three\n49\tcycles to be confirmed, then alerts and (optionally) truncates logs automatically.\n50\t\n51\t**Rootfs at 99% With SSH Failing Silently**\n52\tRootfs fills. SSH stops accepting new connections. Docker can't write log files. State\n53\tfiles fail silently. The server is functionally dead but still technically running.\n54\tNothing in the container layer can detect or recover from this — it requires a reboot.\n55\t\n56\tFix: `stability_watchdog.sh` — watches the server itself: RAM, CPU, disk, kernel, daemon\n57\thealth. Only script in the stack authorized to reboot. Runs last in the orchestrator\n58\tsequence so container healing and pressure reduction always get a chance first.\n59\t\n60\t---\n61\t\n62\t## ━━━ WHAT THIS FOLDER DOES ━━━\n63\t\n64\tFour tiers. Fixed execution order via `watchdog_orchestrator.sh`.\n65\t\n66\t```\n67\tPressure reduction resource_watchdog.sh — throttle/pause/stop before healing fails\n68\tContainer healing docker_watchdog.sh — memory, CPU, HTTP, required containers\n69\tSystem components system_watchdog.sh — thin orchestrator: storage + WebGUI health\n70\t └─ System/ storage_watchdog.sh — pool growth rate + runaway log detection\n71\t webgui_watchdog.sh — WebGUI availability, nginx → php-fpm → emhttp\n72\t network_watchdog.sh — internet, DDNS sync, Tailscale, NPM proxy\n73\tLast resort stability_watchdog.sh — reboot only when nothing else can recover\n74\t```\n75\t\n76\t**The execution order is the design.** Resource pressure is reduced before docker_watchdog\n77\tattempts restarts — containers restarted into a pressure-bound system just fail again.\n78\tSystem component checks run after containers are healed. Stability watchdog runs last —\n79\treboot is always the last option, not the first.\n80\t\n81\t**None of these scripts run standalone loops.** Each is a single-pass script called\n82\tonce every 15 minutes by `Orchestrators/watchdog_orchestrator.sh`. The orchestrator handles\n83\tstartup grace, overlap protection, heartbeat, and sequencing. `system_watchdog.sh` can\n84\talso be run standalone to check all system component watchdogs at once.\n85\t\n86\t---\n87\t\n88\t## ━━━ RELATIONSHIP TO OTHER FOLDERS ━━━\n89\t\n90\t```\n91\tOrchestrators/\n92\t watchdog_orchestrator.sh ──────────────────► resource_watchdog.sh (1st — every 15 min)\n93\t ──────────────────► docker_watchdog.sh (2nd)\n94\t ──────────────────► system_watchdog.sh (3rd — thin orchestrator)\n95\t ──────────────────► stability_watchdog.sh (4th — last resort)\n96\t\n97\tWatchdogs/System/ ← called by system_watchdog.sh each cycle\n98\t storage_watchdog.sh pool growth rate + runaway log detection\n99\t webgui_watchdog.sh WebGUI availability — nginx → php-fpm → emhttp escalation\n100\t network_watchdog.sh internet reachability, DDNS sync, Tailscale, NPM proxy\n101\t\n102\tTools/\n103\t watchdog_skip_list_manager.sh ◄────────────── docker_watchdog.sh writes skip list\n104\t (operator utility — inspect + clear after fixing a crash-looping container)\n105\t\n106\tDocker_Essentials/\n107\t All container lifecycle scripts (daily restart, updates, network) — unaffected.\n108\t docker_watchdog.sh coordinates with them via shared state, not direct calls.\n109\t```\n110\t\n111\t**watchdog_orchestrator.sh stays in Orchestrators/** — it's a job runner, not a watchdog.\n112\t**watchdog_skip_list_manager.sh stays in Tools/** — it's an operator utility, not a watchdog.\n113\t\n114\t---\n115\t\n116\t## ━━━ SCRIPTS IN THIS FOLDER ━━━\n117\t\n118\t| Script | Role | Called By |\n119\t|--------|------|-----------|\n120\t| `resource_watchdog.sh` | Three-level pressure reduction — throttle, pause, stop | `watchdog_orchestrator.sh` — 1st every 15 min |\n121\t| `docker_watchdog.sh` | Two-tier container healing — memory, CPU, HTTP, required | `watchdog_orchestrator.sh` — 2nd every 15 min |\n122\t| `system_watchdog.sh` | Thin orchestrator — runs System/ component watchdogs in sequence | `watchdog_orchestrator.sh` — 3rd every 15 min |\n123\t| `stability_watchdog.sh` | Last-resort server watchdog — reboots when healing has failed | `watchdog_orchestrator.sh` — 4th every 15 min |\n124\t| `System/storage_watchdog.sh` | Pool growth rate + runaway log detection and remediation | `system_watchdog.sh` — every 15 min |\n125\t| `System/webgui_watchdog.sh` | WebGUI availability — nginx → php-fpm → emhttp escalation | `system_watchdog.sh` — every 15 min |\n126\t| `System/network_watchdog.sh` | Internet reachability, DDNS sync, Tailscale, NPM proxy | `system_watchdog.sh` — every 15 min |\n127\t\n128\t> `watchdog_orchestrator.sh` is in `Orchestrators/`. `watchdog_skip_list_manager.sh`\n129\t> is in `Tools/`. Neither is a watchdog — they sit at the edges of this system.\n130\t\n131\t---\n132\t\n133\t## ━━━ HOW THE SCRIPTS RELATE ━━━\n134\t\n135\t```\n136\tEvery 15 minutes — watchdog_orchestrator.sh fires:\n137\t\n138\t Step 1 — resource_watchdog.sh\n139\t │ RAM/load OK → pass through (no action)\n140\t │ Level 1 (soft): throttle SABnzbd + qBit download speeds\n141\t │ Level 2 (medium): further throttle + docker pause non-critical containers\n142\t │ Level 3 (hard): docker stop optional services\n143\t │ writes mem_shutdown_active=true → RW_STATE_FILE\n144\t │ ↓\n145\t Step 2 — docker_watchdog.sh\n146\t │ reads RW_STATE_FILE — if mem_shutdown_active=true: skip all restarts\n147\t │\n148\t │ Tier 1 — explicit per-container checks (configured in host*.conf):\n149\t │ memory hard limits → immediate restart (no strikes)\n150\t │ CPU high sustained → 2-strike restart\n151\t │ HTTP non-response → 2-strike restart\n152\t │ required stopped → restart (dependency ordering respected)\n153\t │\n154\t │ Tier 2 — global scan of all running containers:\n155\t │ unhealthy / OOM / crashloop / dead / non-zero exit → restart\n156\t │ N restarts in window → skip list + critical notify → human required\n157\t │\n158\t │ (skip list management) → Tools/watchdog_skip_list_manager.sh\n159\t │\n160\t Step 3 — system_watchdog.sh (thin orchestrator)\n161\t │ runs SYSTEM_WATCHDOG_SCRIPTS from master.conf sequentially:\n162\t │\n163\t │ └─ storage_watchdog.sh\n164\t │ growth rate scan: du -sm appdata/* → compare to previous cycle baseline\n165\t │ growth > WATCHDOG_APPDATA_GROWTH_GB → 3-strike warn → alert\n166\t │ log file scan: find *.log > WATCHDOG_APPDATA_LOG_MAX_GB\n167\t │ oversize log found → 3-strike warn → truncate (if enabled) or alert\n168\t │\n169\t │ └─ webgui_watchdog.sh\n170\t │ curl check → WebGUI responding → exit 0 (silent)\n171\t │ not responding → nginx restart → wait → recheck\n172\t │ still down → php-fpm restart → wait → recheck\n173\t │ still down → emhttp restart → wait → recheck\n174\t │ all three failed → critical notify, manual intervention needed\n175\t │\n176\t │ └─ network_watchdog.sh\n177\t │ internet check → fail = alert + skip remaining (prevents false positives)\n178\t │ DDNS: public IP vs dig @1.1.1.1 → mismatch = restart Cloudflare DDNS container\n179\t │ Tailscale: status Running → pass; not running = notify (no auto-restart)\n180\t │ NPM proxy: curl external URL → 2-strike system → restart NginxProxyManager\n181\t │\n182\t Step 4 — stability_watchdog.sh\n183\t checks the server itself — RAM, CPU temp, rootfs, FDs, kernel, daemon\n184\t Tier 1 CRITICAL → immediate reboot (no strikes)\n185\t Tier 2 URGENT → reboot if OOM confirmed\n186\t Tier 3 STANDARD → N consecutive failures → reboot\n187\t Abort conditions → ZFS unhealthy / parity running / mover active\n188\t Rate limit → max N reboots per window before switching to notify\n189\t```\n190\t\n191\t**State file coordination between scripts:**\n192\t\n193\t| State File | Written By | Read By | Purpose |\n194\t|-----------|-----------|---------|---------|\n195\t| `RW_STATE_FILE` | `resource_watchdog.sh` | `docker_watchdog.sh` | `mem_shutdown_active` flag — defer restarts during RAM emergency |\n196\t| `SYS_WATCHDOG_STATE_FILE` | `stability_watchdog.sh` | `docker_watchdog.sh` | `watchdog_cycle` heartbeat — stale guard (2hr timeout) |\n197\t| `WATCHDOG_STATE_FILE` | `docker_watchdog.sh` | itself | CPU/HTTP strike counts per container |\n198\t| `SYS_WATCHDOG_FAILED_FILE` | `docker_watchdog.sh` | `watchdog_skip_list_manager.sh` | Container skip list |\n199\t| `STORAGE_WATCHDOG_STATE_FILE` | `System/storage_watchdog.sh` | itself | Growth + log strike counts |\n200\t| `WATCHDOG_APPDATA_GROWTH_FILE` | `System/storage_watchdog.sh` | itself | Per-container size baseline for growth rate |\n201\t| `NETWORK_WATCHDOG_NPM_STATE_FILE` | `System/network_watchdog.sh` | itself | NPM proxy strike count |\n202\t"}]},"uuid":"93fe2f5a-a59a-488f-b4dd-a507cf0a3e86","timestamp":"2026-06-14T05:51:48.617Z","sourceToolAssistantUUID":"1ccbfbf9-86f2-47c3-8e11-ab7bd482c972","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"93fe2f5a-a59a-488f-b4dd-a507cf0a3e86","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01Hd9ALeRsJueFZoJiCTut77","name":"Read","input":{"file_path":"/boot/config/plugins/varaverk/Watchdogs/Manual-Watchdogs.md"},"caller":{"type":"direct"}}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"be130871-457f-4658-9262-4c3d2c3036eb","timestamp":"2026-06-14T05:51:48.768Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"be130871-457f-4658-9262-4c3d2c3036eb","isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01Hd9ALeRsJueFZoJiCTut77","type":"tool_result","content":"1\t# ━━━━━ WATCHDOGS — Manual ━━━━━\n2\t\n3\tConfiguration reference, operational procedures, and troubleshooting for all watchdog\n4\tscripts. For design philosophy and script relationships see `README-Watchdogs.md`.\n5\tFor the orchestrator that calls these scripts see `Orchestrators/watchdog_orchestrator.sh`.\n6\t\n7\t---\n8\t\n9\t## ━━━ CONTENTS ━━━\n10\t\n11\t- [Output Tiers](#output-tiers)\n12\t- [resource_watchdog.sh](#resource_watchdogsh)\n13\t- [docker_watchdog.sh](#docker_watchdogsh)\n14\t- [system_watchdog.sh](#system_watchdogsh)\n15\t- [System/storage_watchdog.sh](#systemstorage_watchdogsh)\n16\t- [System/webgui_watchdog.sh](#systemwebgui_watchdogsh)\n17\t- [System/network_watchdog.sh](#systemnetwork_watchdogsh)\n18\t- [stability_watchdog.sh](#stability_watchdogsh)\n19\t- [Full Configuration Reference](#full-configuration-reference)\n20\t- [Troubleshooting](#troubleshooting)\n21\t\n22\t---\n23\t\n24\t## Output Tiers\n25\t\n26\tAll watchdog scripts use a two-tier output model: `echo` lines are always visible;\n27\t`log` lines only appear when `--log` is passed.\n28\t\n29\tAll watchdog scripts are **single-pass scripts** called once every 15 minutes by the orchestrator.\n30\tWithout `--log`, only state transitions, warnings, errors, and the conclusion line are\n31\tvisible. Per-check detail is suppressed on clean cycles.\n32\t\n33\t`docker_watchdog.sh` is the exception — it is **silent on clean cycles by design**.\n34\t96 cycles/day means clean-cycle noise would bury real events. Its output only appears\n35\twhen there are restarts, skip-list events, or RAM deferral. Use `--log` to see\n36\tper-cycle detail on clean cycles.\n37\t\n38\t---\n39\t\n40\t## resource_watchdog.sh\n41\t\n42\tRuns first in the orchestrator sequence. Reduces system pressure before docker_watchdog\n43\tattempts any container restarts. Containers restarted into a RAM-pressured system just\n44\tfail again — this script ensures docker_watchdog has breathing room.\n45\t\n46\t### Pressure Levels\n47\t\n48\tThree escalating levels, each additive:\n49\t\n50\t```bash\n51\t# master.conf\n52\tRW_RAM_SOFT_GB=20 # Level 1 trigger — throttle downloaders\n53\tRW_RAM_MEDIUM_GB=15 # Level 2 trigger — throttle + pause containers\n54\tRW_RAM_HARD_GB=10 # Level 3 trigger — stop containers + defer docker_watchdog\n55\tRW_RAM_RECOVER_GB=25 # de-escalate only after RAM reaches this\n56\t\n57\tRW_LOAD_SOFT_MULTIPLIER=2.0 # load > 2× cpu count = level 1\n58\tRW_LOAD_MEDIUM_MULTIPLIER=3.0 # load > 3× cpu count = level 2\n59\t\n60\tRW_RECOVER_CYCLES=3 # consecutive under-threshold runs before de-escalating\n61\t```\n62\t\n63\t**Level 1 (soft):** Throttle SABnzbd + qBittorrent download speeds.\n64\t**Level 2 (medium):** Further throttle + `docker pause` non-critical containers.\n65\t**Level 3 (hard):** `docker stop` optional services + write `mem_shutdown_active=true`\n66\tto `RW_STATE_FILE`. docker_watchdog.sh reads this flag and skips all restart logic\n67\tuntil pressure clears. Without this coordination, docker_watchdog would immediately\n68\trestart containers that resource_watchdog just stopped to free RAM.\n69\t\n70\tRecovery de-escalates one level at a time — prevents flip-flopping between states.\n71\t\n72\t### Per-Host Container Lists\n73\t\n74\t```bash\n75\t# host1.conf\n76\tHOST1_RW_PAUSE_CONTAINERS=(\"Tdarr\" \"HandBrake\") # paused at level 2\n77\tHOST1_RW_STOP_CONTAINERS=(\"LocalAI\" \"Satisfactory\") # stopped at level 3\n78\t```\n79\t\n80\tContainers in `RW_CRITICAL_CONTAINERS` are never paused or stopped regardless of\n81\tpressure level. Default: Emby, NginxProxyManager, Authelia, Mariadb, Redis.\n82\t\n83\t### Downloader Throttle Config\n84\t\n85\t```bash\n86\t# master.conf\n87\tRW_SABNZBD_ENABLED=true\n88\tRW_SABNZBD_SPEED_SOFT=\"50M\" # throttled at level 1\n89\tRW_SABNZBD_SPEED_MEDIUM=\"10M\" # throttled further at level 2\n90\t\n91\tRW_QBIT_ENABLED=true\n92\tRW_QBIT_DL_SOFT=51200 # KB/s — level 1\n93\tRW_QBIT_DL_MEDIUM=10240 # KB/s — level 2\n94\t\n95\t# host1.conf (API access)\n96\tHOST1_SABNZBD_URL=\"http://localhost:8080\"\n97\tHOST1_SABNZBD_API_KEY=\"your-api-key\"\n98\tHOST1_QBIT_URL=\"http://localhost:8090\"\n99\tHOST1_QBIT_USERNAME=\"admin\"\n100\tHOST1_QBIT_PASSWORD=\"your-password\"\n101\t```\n102\t\n103\t### Usage\n104\t\n105\t```bash\n106\tresource_watchdog.sh # single pass (called by watchdog_orchestrator.sh)\n107\tresource_watchdog.sh --dry-run # show what would be throttled/paused/stopped\n108\tresource_watchdog.sh --status # current level, active actions, recovery cycle count\n109\tresource_watchdog.sh --log # verbose per-check output\n110\t```\n111\t\n112\t---\n113\t\n114\t## docker_watchdog.sh\n115\t\n116\tRuns second in the orchestrator sequence. Two-tier container healing — explicit\n117\tper-container configuration (Tier 1) plus a global catch-all scan (Tier 2).\n118\t\n119\tReads `RW_STATE_FILE` at cycle start — if `mem_shutdown_active=true`, skips all\n120\tcontainer restart logic (resource_watchdog is managing the situation). Health URL\n121\tchecks for excluded containers still run.\n122\t\n123\t### Memory Hard Limits\n124\t\n125\t```bash\n126\t# host1.conf\n127\t# Format: \"ContainerName:LimitInMB\"\n128\t# Immediate restart when exceeded — no strike system. Memory leaks are not spikes.\n129\t#\n130\t# Sizing: check normal peak with \"docker stats ContainerName\"\n131\t# Set limit at ~150-200% of normal peak\n132\t#\n133\tHOST1_WATCHDOG_CONTAINERS=(\n134\t \"Emby:18432\" # 18GB — peaks ~12GB under heavy transcode load\n135\t \"LidaTube:6144\" # 6GB — YouTube downloader, grows with large queues\n136\t \"Tdarr:6144\" # 6GB — video transcoder, memory-intensive by nature\n137\t \"Code-Server:1024\" # 1GB — IDE, should be light; 1GB is generous\n138\t)\n139\t```\n140\t\n141\tA soft warning fires at `SOFT_MEM_THRESHOLD=80` percent of the hard limit — early\n142\tvisibility into a container approaching its ceiling before a restart is triggered.\n143\t\n144\t### CPU Thresholds\n145\t\n146\t```bash\n147\t# master.conf\n148\t# CPU is normalised against total core count.\n149\t# 85% normalised on a 16-core machine = 13.6 cores worth of a single process.\n150\t#\n151\t# CPU uses a STRIKE SYSTEM — brief spikes are normal (Tdarr, Emby transcoding, SABnzbd).\n152\t# Strike 1: above HARD_CPU_THRESHOLD → warn, increment strike\n153\t# Strike 2: above threshold → restart, reset counter\n154\t# Recovery: drops below threshold any cycle → reset to 0\n155\t#\n156\tSOFT_CPU_THRESHOLD=50 # warn at 50% normalised — informational only\n157\tHARD_CPU_THRESHOLD=85 # strike at 85% normalised\n158\tCPU_FAIL_LIMIT=2 # consecutive strikes before restart\n159\t```\n160\t\n161\t### HTTP Health Checks\n162\t\n163\t```bash\n164\t# host1.conf\n165\t# Format: \"ContainerName:http://host:port/optional-path\"\n166\t# Two consecutive non-responses trigger a restart.\n167\t# \"Container running\" and \"service responding\" are not the same thing.\n168\t#\n169\tHOST1_WATCHDOG_CONTAINER_URLS=(\n170\t \"Emby:http://localhost:8096\" # Emby WebUI root\n171\t \"NginxProxyManager:http://localhost:81\" # NPM admin interface\n172\t)\n173\t\n174\t# master.conf\n175\tCURL_TIMEOUT=5 # seconds before non-response counts as a failure\n176\tRESP_FAIL_LIMIT=2 # consecutive failures before restart\n177\t```\n178\t\n179\t### Required Containers\n180\t\n181\t```bash\n182\t# host1.conf\n183\t# Found stopped → restart attempted every cycle until running or skip-listed.\n184\t#\n185\tHOST1_WATCHDOG_REQUIRED_CONTAINERS=(\n186\t \"NginxProxyManager\"\n187\t \"Authelia\"\n188\t \"Mariadb-Authelia\"\n189\t \"Redis-Authelia\"\n190\t)\n191\t```\n192\t\n193\t### Dependency Ordering\n194\t\n195\t```bash\n196\t# host1.conf\n197\t# Format: \"DependentContainer:dependency1 dependency2\"\n198\t# All dependencies must be running before the dependent is restarted.\n199\t# Prevents Authelia crash-looping while MariaDB is still starting.\n200\t#\n201\tHOST1_WATCHDOG_DEPENDENCIES=(\n202\t \"Authelia:Mariadb-Authelia Redis-Authelia\"\n203\t \"NextCloud:Postgres-NextCloud\"\n204\t)\n205\t```\n206\t\n207\tSame dependency config is used by `docker_daily_restart.sh` and\n208\t`docker_weekly_restart.sh` — configure once, applies everywhere.\n209\t\n210\t### Startup Grace Period\n211\t\n212\t```bash\n213\t# master.conf\n214\t# Suppress restart actions for N seconds after array start.\n215\t# Checks still run and log — only restart actions are suppressed.\n216\t#\n217\tWATCHDOG_STARTUP_GRACE=600 # 10 minutes\n218\t```\n219\t\n220\t### Tier 2 Global Scan\n221\t\n222\t```bash\n223\t# master.conf\n224\tWATCHDOG_SCAN_ALL=true # enable global scan\n225\tWATCHDOG_RESTART_UNHEALTHY=true # act on Docker HEALTHCHECK failures\n226\tWATCHDOG_NOTIFY_OOM=true # detect and notify kernel OOM kills\n227\tWATCHDOG_NOTIFY_CRASHLOOP=true # detect escalating restart counts\n228\tWATCHDOG_RESTART_DEAD=true # recover containers in dead state\n229\tWATCHDOG_RESTART_CRASHED=true # restart containers that exited non-zero\n230\tWATCHDOG_CRASH_LIMIT=5 # RestartCount above this → restart + skip list\n231\t\n232\t# Exclude containers from Tier 2 (one-shots, manually managed, benign exits):\n233\tWATCHDOG_SCAN_IGNORE=(\n234\t \"my-one-shot-container\"\n235\t)\n236\t```\n237\t\n238\t### Restart Loop Protection\n239\t\n240\t```bash\n241\t# master.conf\n242\tWATCHDOG_CONTAINER_RESTART_LIMIT=3 # restarts in the window before skip list\n243\tWATCHDOG_CONTAINER_RESTART_WINDOW=1 # rolling window in hours\n244\t```\n245\t\n246\tAfter hitting the limit: skip list + critical notification. The watchdog stops\n247\ttouching the container. Auto-clear: if the container recovers on its own and is\n248\tfound running, it's removed from the skip list automatically. Manual clear is only\n249\tneeded when the container is stuck stopped — use `Tools/watchdog_skip_list_manager.sh`.\n250\t\n251\t### Notification Batching\n252\t\n253\t```bash\n254\t# master.conf\n255\tWATCHDOG_BATCH_NOTIFY=true\n256\t# All events from one cycle → one notification at the end.\n257\t# A shared DB going down can cascade 10+ containers. Without batching: 10 pings.\n258\t# With batching: one summary listing all affected containers.\n259\t```\n260\t\n261\t### Usage\n262\t\n263\t```bash\n264\tdocker_watchdog.sh # single pass (called by watchdog_orchestrator.sh)\n265\tdocker_watchdog.sh --dry-run # full cycle preview without restarting anything\n266\tdocker_watchdog.sh --status # skip list, strike counts, grace period, RAM deferral state\n267\tdocker_watchdog.sh --log # verbose per-cycle output\n268\t```\n269\t\n270\t### Skip List Recovery\n271\t\n272\t```bash\n273\t# Step 1 — understand the situation\n274\tTools/watchdog_skip_list_manager.sh --status\n275\t\n276\t# Step 2 — fix the underlying problem\n277\t# docker logs ContainerName --tail 100\n278\t# df -h /mnt/user\n279\t\n280\t# Step 3 — clear the container\n281\tTools/watchdog_skip_list_manager.sh --clear ContainerName\n282\t\n283\t# Step 4 — start manually (confirms fix before handing back to watchdog)\n284\tdocker start ContainerName\n285\t\n286\t# Step 5 — monitoring resumes automatically on next cycle\n287\t```\n288\t\n289\t---\n290\t\n291\t## system_watchdog.sh\n292\t\n293\tThin orchestrator — runs `SYSTEM_WATCHDOG_SCRIPTS` from master.conf sequentially each cycle.\n294\tCalled third by `watchdog_orchestrator.sh`. Covers all system component watchdogs.\n295\tCan also be run standalone to check all system components at once.\n296\t\n297\t### Usage\n298\t\n299\t```bash\n300\tsystem_watchdog.sh # run all system component watchdogs\n301\tsystem_watchdog.sh --status # show configured scripts and their paths\n302\tsystem_watchdog.sh --dry-run # preview without executing anything\n303\tsystem_watchdog.sh --log # verbose output\n304\t```\n305\t\n306\t---\n307\t\n308\t## System/storage_watchdog.sh\n309\t\n310\tCalled by `system_watchdog.sh` each cycle. Two independent checks per cycle:\n311\tgrowth rate detection (automatic, zero config) and oversize log detection.\n312\tUses its own strike state file — independent from docker_watchdog.\n313\t\n314\t### Growth Rate Detection\n315\t\n316\t```bash\n317\t# master.conf\n318\tWATCHDOG_CHECK_APPDATA=true\n319\tWATCHDOG_APPDATA_PATHS=(\"/mnt/docker-unraid/appdata\")\n320\tWATCHDOG_APPDATA_GROWTH_GB=2 # growth per cycle that triggers a strike\n321\tWATCHDOG_APPDATA_STRIKE_LIMIT=3 # strikes before alert\n322\tWATCHDOG_APPDATA_GROWTH_FILE=\"/tmp/watchdog_appdata_growth.db\" # size baseline\n323\t```\n324\t\n325\tRuns `du -sm appdata/*/` each cycle — pure inode metadata, very lightweight on NVMe.\n326\tCompares each container's current size to the baseline from the previous cycle.\n327\tGrowth > `WATCHDOG_APPDATA_GROWTH_GB` per cycle increments the container's strike count.\n328\tStrike 1: warn. Strike 2: escalate. Strike 3: critical alert.\n329\tStrikes auto-clear when growth drops to zero (condition resolved).\n330\t\n331\t**Zero configuration required for new containers.** Growth rate detection covers all\n332\tcontainers automatically. The suppress array below is only for known-legitimate growth.\n333\t\n334\t### Growth Suppress Ceilings\n335\t\n336\t```bash\n337\t# host1.conf\n338\t# ONLY needed in specific cases — growth rate detection covers everything automatically.\n339\t# Use when a container's appdata legitimately grows fast during normal operation\n340\t# and you want to suppress false positives above a known-safe threshold.\n341\t#\n342\tdeclare -A HOST1_WATCHDOG_APPDATA_SIZES=(\n343\t [\"Tdarr\"]=\"25600\" # 25GB — transcode cache grows during active jobs\n344\t [\"7dtd\"]=\"20480\" # 20GB — game server world data, expected large\n345\t)\n346\t```\n347\t\n348\t### Log File Detection\n349\t\n350\t```bash\n351\t# master.conf\n352\tWATCHDOG_APPDATA_LOG_MAX_GB=2 # *.log / *.log.* files above this trigger a strike\n353\tWATCHDOG_APPDATA_TRUNCATE_LOGS=false # true: truncate at strike limit; false: alert only\n354\t```\n355\t\n356\tScans all `*.log` and `*.log.*` files across all appdata paths. Files over the threshold\n357\tincrement per-file strike counts. At strike limit: truncates in-place with `truncate -s 0`\n358\t(container keeps its file handle — space reclaimed immediately without container restart)\n359\tor sends a critical alert if truncation is disabled.\n360\t\n361\tLog strikes auto-clear when the file drops below threshold.\n362\t\n363\t### Usage\n364\t\n365\t```bash\n366\tstorage_watchdog.sh # single pass (called by system_watchdog.sh)\n367\tstorage_watchdog.sh --status # strikes, growth baseline age, suppress ceilings\n368\tstorage_watchdog.sh --dry-run # show what would be alerted/truncated\n369\tstorage_watchdog.sh --log # verbose per-container output\n370\t```\n371\t\n372\t---\n373\t\n374\t## System/webgui_watchdog.sh\n375\t\n376\tCalled by `system_watchdog.sh` each cycle. Monitors WebGUI availability and escalates\n377\tthrough three restart steps if unresponsive. Silent when healthy.\n378\t\n379\t### Escalation Path\n380\t\n381\t```\n382\tWebGUI responding → exit 0 (silent)\n383\t\n384\tNot responding:\n385\t Step 1 — nginx restart → wait WEBGUI_NGINX_WAIT → recheck\n386\t Step 2 — php-fpm restart → wait WEBGUI_PHP_WAIT → recheck\n387\t Step 3 — emhttp restart → wait WEBGUI_EMHTTP_WAIT → recheck\n388\t All three failed → critical notify, manual intervention needed\n389\t```\n390\t\n391\t### Configuration\n392\t\n393\t```bash\n394\t# master.conf\n395\tWEBGUI_URL=\"http://localhost\"\n396\tWEBGUI_TIMEOUT=5\n397\tWEBGUI_NGINX_WAIT=15\n398\tWEBGUI_PHP_WAIT=10\n399\tWEBGUI_EMHTTP_WAIT=30\n400\t```\n401\t\n402\t### Usage\n403\t\n404\t```bash\n405\twebgui_watchdog.sh # single pass (called by system_watchdog.sh)\n406\twebgui_watchdog.sh --status # current WebGUI state + nginx/php-fpm/emhttp status\n407\twebgui_watchdog.sh --dry-run # show which services would be restarted\n408\twebgui_watchdog.sh --log # verbose per-step output\n409\t```\n410\t\n411\t---\n412\t\n413\t## System/network_watchdog.sh\n414\t\n415\tCalled by `system_watchdog.sh` each cycle. Checks that the outside world can actually\n416\treach what it needs to reach. Internet reachability gates all other checks — if upstream\n417\tis down, DDNS and NPM checks are skipped to prevent false positives.\n418\t\n419\t### Check Sequence\n420\t\n421\t```\n422\t1. Internet → curl NETWORK_WATCHDOG_INTERNET_URL\n423\t fail → alert + exit (skip all remaining checks)\n424\t\n425\t2. DDNS → public IP (ifconfig.me) vs DNS record (dig @1.1.1.1)\n426\t match → pass (silent)\n427\t mismatch → restart HOST*_NETWORK_WATCHDOG_DDNS_CONTAINER + notify\n428\t\n429\t3. Tailscale → tailscale status --json BackendState\n430\t Running → pass (silent)\n431\t not Running → notify (no auto-restart — warrants human review)\n432\t\n433\t4. NPM proxy → curl HOST*_NETWORK_WATCHDOG_NPM_URL (external)\n434\t reachable → pass, clear strikes\n435\t not reachable → strike 1: warn + notify\n436\t strike 2: restart NginxProxyManager + notify + clear strikes\n437\t```\n438\t\n439\t### Configuration\n440\t\n441\t```bash\n442\t# master.conf\n443\tNETWORK_WATCHDOG_ENABLED=true\n444\tNETWORK_WATCHDOG_INTERNET_URL=\"https://1.1.1.1\"\n445\tNETWORK_WATCHDOG_INTERNET_TIMEOUT=5\n446\tNETWORK_WATCHDOG_CHECK_TAILSCALE=true\n447\tNETWORK_WATCHDOG_NPM_TIMEOUT=10\n448\tNETWORK_WATCHDOG_NPM_STRIKE_LIMIT=2\n449\tNETWORK_WATCHDOG_NPM_STATE_FILE=\"/tmp/network_watchdog_state.db\"\n450\t\n451\t# host*.conf (host-specific)\n452\tHOST1_NETWORK_WATCHDOG_DDNS_DOMAIN=\"gmer4lfe.com\"\n453\tHOST1_NETWORK_WATCHDOG_DDNS_CONTAINER=\"Gmer4Lfe.com\"\n454\tHOST1_NETWORK_WATCHDOG_NPM_URL=\"https://gmer4lfe.com\"\n455\t```\n456\t\n457\t### Usage\n458\t\n459\t```bash\n460\tnetwork_watchdog.sh # run all connectivity checks (silent when healthy)\n461\tnetwork_watchdog.sh --status # current IP, DNS record, Tailscale state, NPM strike count\n462\tnetwork_watchdog.sh --dry-run # check without restarting any containers\n463\tnetwork_watchdog.sh --log # verbose per-check output\n464\t```\n465\t\n466\t### Troubleshooting\n467\t\n468\t**DDNS keeps restarting the container but record stays stale**\n469\t```bash\n470\t# Check if the container is actually running after restart:\n471\tdocker ps | grep \"Gmer4Lfe.com\"\n472\t\n473\t# Check container logs for Cloudflare API errors:\n474\tdocker logs \"Gmer4Lfe.com\" --tail 20\n475\t\n476\t# Verify public IP detection:\n477\tcurl -sf https://ifconfig.me\n478\t\n479\t# Verify DNS resolution:\n480\tdig +short gmer4lfe.com @1.1.1.1\n481\t```\n482\t\n483\t**NPM strikes accumulating but NPM is running**\n484\t```bash\n485\t# Check if the external URL is actually responding:\n486\tcurl -sv https://gmer4lfe.com 2>&1 | head -20\n487\t\n488\t# NPM may be running but a backend container is down — check the specific service\n489\t# the URL routes to, not just NginxProxyManager itself.\n490\t\n491\t# Check NPM strike count:\n492\tnetwork_watchdog.sh --status\n493\t```\n494\t\n495\t**Tailscale showing not Running**\n496\t```bash\n497\t# Check tailscale state directly:\n498\ttailscale status\n499\t\n500\t# Check the backend state specifically:\n501\ttailscale status --json | grep BackendState\n502\t\n503\t# Reconnect manually if needed:\n504\ttailscale up\n505\t```\n506\t\n507\t---\n508\t\n509\t## stability_watchdog.sh\n510\t\n511\tRuns last in the orchestrator sequence. The only script in the ecosystem authorized\n512\tto reboot. Watches the server itself — not containers, not storage. Reboots only when\n513\thealing at every other layer has failed or when the failure is non-recoverable.\n514\t\n515\t### Three-Tier Response\n516\t\n517\t**Tier 1 — CRITICAL (immediate reboot, no strikes)**\n518\t| Condition | Threshold | Why immediate |\n519\t|-----------|-----------|---------------|\n520\t| Docker daemon unresponsive | N/A | Every docker command hangs — nothing can be healed |\n521\t| rootfs usage | `SYS_WATCHDOG_ROOTFS_CRITICAL_PCT` (99%) | SSH stops; state files fail silently |\n522\t| Kernel oops/BUG in dmesg | delta > 0 | Kernel running with corrupted state |\n523\t| File descriptor exhaustion | `SYS_WATCHDOG_FD_CRITICAL_PCT` (95%) | New connections silently failing |\n524\t| /boot read-only | write test fails | Config writes silently failing |\n525\t\n526\t**Tier 2 — URGENT (bypass strikes with OOM confirmation)**\n527\tRAM below `SYS_WATCHDOG_MEM_GB` AND OOM kills this cycle ≥ `SYS_WATCHDOG_OOM_LIMIT`.\n528\tBoth conditions required — RAM alone uses the standard strike system.\n529\tOOM confirms the system is dying faster than watchdogs can heal.\n530\t\n531\t**Tier 3 — STANDARD (`SYS_WATCHDOG_STRIKE_LIMIT` consecutive failures → reboot)**\n532\t| Check | Threshold |\n533\t|-------|-----------|\n534\t| Free RAM | `SYS_WATCHDOG_MEM_GB` (reboot trigger — earlier tiers handled by resource_watchdog) |\n535\t| Load average | `SYS_WATCHDOG_LOAD_MULTIPLIER` × cpu_count |\n536\t| CPU temperature | `SYS_WATCHDOG_CPU_TEMP_MAX` |\n537\t| Zombie processes | `SYS_WATCHDOG_ZOMBIE_LIMIT` |\n538\t| /var/log usage | `SYS_WATCHDOG_LOG_PCT` |\n539\t| /tmp usage | `SYS_WATCHDOG_TMP_PCT` |\n540\t| Array disk errors | mdstat error delta > 0 |\n541\t| NIC state | interface operstate != \"up\" |\n542\t\n543\t### RAM Tiers\n544\t\n545\tRAM pressure is a graduated response split across resource_watchdog and stability_watchdog:\n546\t\n547\t```\n548\tRW_RAM_SOFT_GB (12GB) → throttle downloads, reduce background load (resource_watchdog)\n549\tRW_RAM_MEDIUM_GB (8GB) → pause background containers (resource_watchdog)\n550\tRW_RAM_HARD_GB (6GB) → stop optional containers, wait for recovery (resource_watchdog)\n551\tSYS_WATCHDOG_MEM_GB (4GB) → strike → reboot (last resort) (stability_watchdog)\n552\tRW_RAM_RECOVER_GB (20GB) → RAM must reach this before stopped containers restart\n553\t```\n554\t\n555\t`SYS_WATCHDOG_MEM_GB` must always be below `RW_RAM_HARD_GB` — resource_watchdog acts first.\n556\tContainers stopped at the hard tier use `HOST*_RW_STOP_CONTAINERS` in host*.conf.\n557\tContainers paused at the medium tier use `HOST*_RW_PAUSE_CONTAINERS` in host*.conf.\n558\t\n559\t### Abort Conditions\n560\t\n561\tPrevent reboot — running them would risk data loss:\n562\t\n563\t```bash\n564\tSYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true # ZFS pool degraded/faulted\n565\tSYS_WATCHDOG_ABORT_ON_PARITY=false # aborting parity mid-check is worse than crashing\n566\tSYS_WATCHDOG_ABORT_ON_MOVER=false # aborting move mid-run is worse than crashing\n567\t```\n568\t\n569\tTier 1 CRITICAL bypasses all abort conditions — an imminent crash outweighs data\n570\tsafety concerns.\n571\t\n572\t### Reboot Rate Limit\n573\t\n574\t```bash\n575\tSYS_WATCHDOG_REBOOT_WINDOW_HRS=2 # window in hours\n576\tSYS_WATCHDOG_REBOOT_LIMIT=3 # max reboots within the window\n577\t```\n578\t\n579\tIf the server reboots `SYS_WATCHDOG_REBOOT_LIMIT` times within the window, the watchdog\n580\tswitches from rebooting to notifying only. Prevents a boot loop where the watchdog\n581\treboots → something crashes again immediately → reboot again.\n582\t\n583\t### Usage\n584\t\n585\t```bash\n586\tstability_watchdog.sh # single pass (called by watchdog_orchestrator.sh)\n587\tstability_watchdog.sh --dry-run # run detection logic without rebooting\n588\tstability_watchdog.sh --status # thresholds, current state, strike counts\n589\tstability_watchdog.sh --log # verbose per-check output\n590\t```\n591\t\n592\t---\n593\t\n594\t## Full Configuration Reference\n595\t\n596\t```bash\n597\t# master.conf\n598\t\n599\t# ── Resource Watchdog ──────────────────────────────────────────────────────────\n600\tRW_ENABLED=true\n601\tRW_STATE_FILE=\"/tmp/resource_watchdog_state.db\"\n602\t\n603\tRW_RAM_SOFT_GB=20\n604\tRW_RAM_MEDIUM_GB=15\n605\tRW_RAM_HARD_GB=10\n606\tRW_RAM_RECOVER_GB=25\n607\t\n608\tRW_LOAD_SOFT_MULTIPLIER=2.0\n609\tRW_LOAD_MEDIUM_MULTIPLIER=3.0\n610\tRW_RECOVER_CYCLES=3\n611\t\n612\tRW_SABNZBD_ENABLED=true\n613\tRW_SABNZBD_SPEED_SOFT=\"50M\"\n614\tRW_SABNZBD_SPEED_MEDIUM=\"10M\"\n615\tRW_QBIT_ENABLED=true\n616\tRW_QBIT_DL_SOFT=51200 # KB/s\n617\tRW_QBIT_DL_MEDIUM=10240\n618\t\n619\tRW_CRITICAL_CONTAINERS=(\"Emby\" \"NginxProxyManager\" \"Authelia\" \"Mariadb\" \"Redis\")\n620\t\n621\t# host*.conf\n622\tHOST1_RW_PAUSE_CONTAINERS=(\"Tdarr\" \"HandBrake\")\n623\tHOST1_RW_STOP_CONTAINERS=(\"LocalAI\" \"Satisfactory\")\n624\tHOST1_SABNZBD_URL=\"http://localhost:8080\"\n625\tHOST1_SABNZBD_API_KEY=\"your-api-key\"\n626\tHOST1_QBIT_URL=\"http://localhost:8090\"\n627\tHOST1_QBIT_USERNAME=\"admin\"\n628\tHOST1_QBIT_PASSWORD=\"your-password\"\n629\t\n630\t# ── Docker Watchdog ────────────────────────────────────────────────────────────\n631\tWATCHDOG_STARTUP_GRACE=600 # seconds before restarts begin after boot\n632\tCONTAINER_DELAY=15 # seconds between dependency + dependent restart\n633\t\n634\tSOFT_MEM_THRESHOLD=80 # warn at % of hard limit (no restart)\n635\t\n636\tSOFT_CPU_THRESHOLD=50\n637\tHARD_CPU_THRESHOLD=85\n638\tCPU_FAIL_LIMIT=2\n639\t\n640\tCURL_TIMEOUT=5\n641\tRESP_FAIL_LIMIT=2\n642\t\n643\tWATCHDOG_CONTAINER_RESTART_LIMIT=3\n644\tWATCHDOG_CONTAINER_RESTART_WINDOW=1\n645\t\n646\tWATCHDOG_SCAN_ALL=true\n647\tWATCHDOG_SCAN_IGNORE=()\n648\tWATCHDOG_RESTART_UNHEALTHY=true\n649\tWATCHDOG_NOTIFY_OOM=true\n650\tWATCHDOG_NOTIFY_CRASHLOOP=true\n651\tWATCHDOG_CRASH_LIMIT=5\n652\tWATCHDOG_RESTART_DEAD=true\n653\tWATCHDOG_RESTART_CRASHED=true\n654\tWATCHDOG_BATCH_NOTIFY=true\n655\t\n656\t# State files:\n657\tWATCHDOG_STATE_FILE=\"/tmp/container_watchdog_state.db\"\n658\tWATCHDOG_CONTAINER_RESTART_LOG=\"/boot/config/container_restart_history.db\"\n659\t\n660\t# host*.conf\n661\tHOST1_WATCHDOG_CONTAINERS=() # \"ContainerName:LimitMB\"\n662\tHOST1_WATCHDOG_CONTAINER_URLS=() # \"ContainerName:http://host:port\"\n663\tHOST1_WATCHDOG_REQUIRED_CONTAINERS=()\n664\tHOST1_WATCHDOG_DEPENDENCIES=() # \"Dependent:dep1 dep2\"\n665\t\n666\t# ── Storage Watchdog ───────────────────────────────────────────────────────────\n667\tWATCHDOG_CHECK_APPDATA=true\n668\tWATCHDOG_APPDATA_PATHS=(\"/mnt/docker-unraid/appdata\")\n669\tWATCHDOG_APPDATA_GROWTH_GB=2\n670\tWATCHDOG_APPDATA_LOG_MAX_GB=2\n671\tWATCHDOG_APPDATA_TRUNCATE_LOGS=false\n672\tWATCHDOG_APPDATA_STRIKE_LIMIT=3\n673\tWATCHDOG_APPDATA_GROWTH_FILE=\"/tmp/watchdog_appdata_growth.db\"\n674\tSTORAGE_WATCHDOG_STATE_FILE=\"/tmp/storage_watchdog_state.db\"\n675\t\n676\t# host*.conf (optional — only for suppress ceilings)\n677\t# declare -A HOST1_WATCHDOG_APPDATA_SIZES=(\n678\t# [\"Tdarr\"]=\"25600\"\n679\t# )\n680\t\n681\t# ── System Watchdog ────────────────────────────────────────────────────────────\n682\t# Reboot trigger only — warn/shutdown/recover RAM tiers handled by resource_watchdog.sh\n683\t# RW_RAM_HARD_GB > SYS_WATCHDOG_MEM_GB always (resource_watchdog acts before watchdog reboots)\n684\tSYS_WATCHDOG_STRIKE_LIMIT=2\n685\t\n686\tSYS_WATCHDOG_REBOOT_WINDOW_HRS=12\n687\tSYS_WATCHDOG_REBOOT_LIMIT=3\n688\tSYS_WATCHDOG_OOM_LIMIT=3\n689\t\n690\tSYS_WATCHDOG_MEM_GB=4 # strike system → reboot (last resort — below resource_watchdog hard stop)\n691\t\n692\tSYS_WATCHDOG_ROOTFS_CRITICAL_PCT=99\n693\tSYS_WATCHDOG_FD_CRITICAL_PCT=95\n694\tSYS_WATCHDOG_LOAD_MULTIPLIER=4\n695\tSYS_WATCHDOG_CPU_TEMP_MAX=95\n696\tSYS_WATCHDOG_ZOMBIE_LIMIT=50\n697\tSYS_WATCHDOG_LOG_PCT=95\n698\tSYS_WATCHDOG_TMP_PCT=90\n699\t\n700\tSYS_WATCHDOG_ABORT_ON_ZFS_UNHEALTHY=true\n701\tSYS_WATCHDOG_ABORT_ON_PARITY=false\n702\tSYS_WATCHDOG_ABORT_ON_MOVER=false\n703\t\n704\t# State files (all in State_Files/ — survive reboots):\n705\tSYS_WATCHDOG_STATE_FILE=\"State_Files/system_watchdog_state.db\"\n706\tDOCKER_WATCHDOG_FAILED_FILE=\"State_Files/docker_watchdog_failed.db\"\n707\tSYS_WATCHDOG_REBOOT_LOG=\"State_Files/system_watchdog_reboots.db\"\n708\tSYS_WATCHDOG_OOM_FILE=\"State_Files/system_watchdog_oom.db\"\n709\t\n710\t# ── Network Watchdog ───────────────────────────────────────────────────────────\n711\tNETWORK_WATCHDOG_ENABLED=true\n712\tNETWORK_WATCHDOG_INTERNET_URL=\"https://1.1.1.1\"\n713\tNETWORK_WATCHDOG_INTERNET_TIMEOUT=5\n714\tNETWORK_WATCHDOG_CHECK_TAILSCALE=true\n715\tNETWORK_WATCHDOG_NPM_TIMEOUT=10\n716\tNETWORK_WATCHDOG_NPM_STRIKE_LIMIT=2\n717\tNETWORK_WATCHDOG_NPM_STATE_FILE=\"/tmp/network_watchdog_state.db\"\n718\t\n719\t# host*.conf (host-specific)\n720\tHOST1_NETWORK_WATCHDOG_DDNS_DOMAIN=\"gmer4lfe.com\"\n721\tHOST1_NETWORK_WATCHDOG_DDNS_CONTAINER=\"Gmer4Lfe.com\"\n722\tHOST1_NETWORK_WATCHDOG_NPM_URL=\"https://gmer4lfe.com\"\n723\t```\n724\t\n725\t---\n726\t\n727\t## Troubleshooting\n728\t\n729\t### resource_watchdog Paused Containers It Shouldn't\n730\t\n731\t```bash\n732\t# Check current state:\n733\tresource_watchdog.sh --status\n734\t\n735\t# Add to RW_CRITICAL_CONTAINERS in master.conf:\n736\tRW_CRITICAL_CONTAINERS=(\"Emby\" \"NginxProxyManager\" \"Authelia\" \"Mariadb\" \"Redis\" \"MyContainer\")\n737\t\n738\t# Un-pause manually if needed:\n739\tdocker unpause MyContainer\n740\t```\n741\t\n742\t### docker_watchdog Keeps Restarting a Healthy Container\n743\t\n744\t```bash\n745\t# Check what's triggering it — memory, CPU, HTTP, or required:\n746\tdocker_watchdog.sh --status\n747\t\n748\t# Check CPU normalised — brief spikes should not trigger (2-strike system):\n749\t# If triggering on CPU: check if HARD_CPU_THRESHOLD is set appropriately\n750\t# for containers with legitimate burst usage (Tdarr encoding, SABnzbd unpacking)\n751\t\n752\t# Check HTTP — is the health endpoint returning 200?\n753\tcurl -sf --max-time 5 http://localhost:PORT && echo \"OK\" || echo \"FAIL\"\n754\t```\n755\t\n756\t### Container on the Skip List After Fixing the Problem\n757\t\n758\t```bash\n759\t# See skip list and container state:\n760\tTools/watchdog_skip_list_manager.sh --status\n761\t\n762\t# Fix root cause first, then clear:\n763\tTools/watchdog_skip_list_manager.sh --clear ContainerName\n764\t\n765\t# Start manually to confirm fix before handing back to watchdog:\n766\tdocker start ContainerName\n767\t```\n768\t\n769\t### storage_watchdog Alerting on a Container That Grows Legitimately\n770\t\n771\t```bash\n772\t# Check what triggered it:\n773\tstorage_watchdog.sh --status\n774\t\n775\t# Add a suppress ceiling to host*.conf:\n776\t# declare -A HOST1_WATCHDOG_APPDATA_SIZES=(\n777\t# [\"ContainerName\"]=\"10240\" # 10GB ceiling — legitimate growth, suppress below this\n778\t# )\n779\t```\n780\t\n781\t### stability_watchdog Rebooted Unexpectedly\n782\t\n783\t```bash\n784\t# Check the reboot log (survives reboots):\n785\tcat /boot/config/system_watchdog_reboots.db\n786\t# Shows timestamp and reason for each watchdog-triggered reboot\n787\t\n788\t# Check syslog near the reboot time:\n789\tgrep \"stability_watchdog\" /var/log/syslog | tail -20\n790\t```\n791\t\n792\t### stability_watchdog Not Responding / Watchdog Orchestrator Reports Timeout\n793\t\n794\t```bash\n795\t# All watchdogs run as single-pass scripts — there is no background process to check.\n796\t# If the orchestrator reports a timeout, one pass took longer than expected.\n797\t\n798\t# Check the orchestrator itself:\n799\tOrchestrators/watchdog_orchestrator.sh --status\n800\t\n801\t# Run the slow watchdog directly with --log to see where it's hanging:\n802\tWatchdogs/stability_watchdog.sh --log --dry-run\n803\t```\n804\t"}]},"uuid":"2e4cb159-c97c-4d9f-805d-ddf36dc3b929","timestamp":"2026-06-14T05:51:48.981Z","sourceToolAssistantUUID":"be130871-457f-4658-9262-4c3d2c3036eb","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"2e4cb159-c97c-4d9f-805d-ddf36dc3b929","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01QKxwKRXrKLNkY4iqNFqFAC","name":"Read","input":{"file_path":"/boot/config/plugins/varaverk/Watchdogs/docker_watchdog.sh","limit":120},"caller":{"type":"direct"}}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"280ace7a-1c88-4e34-9cf7-8a9c0bb52390","timestamp":"2026-06-14T05:51:49.144Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"280ace7a-1c88-4e34-9cf7-8a9c0bb52390","isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01QKxwKRXrKLNkY4iqNFqFAC","type":"tool_result","content":"1\t#!/bin/bash\n2\t# ==============================================================================================\n3\t# ================================= Docker Watchdog ============================================\n4\t# ==============================================================================================\n5\t#\n6\t# PURPOSE\n7\t# ─────────────────────────────────────────────────────────────────────────────\n8\t# Two-tier self-healing container monitoring system. Called by\n9\t# watchdog_orchestrator.sh via cron every 15 minutes as a single-pass run.\n10\t# Tier 1 applies specific thresholds to explicitly configured containers.\n11\t# Tier 2 scans everything else for generic health problems. Silent on clean\n12\t# cycles, loud when something needs attention.\n13\t#\n14\t# ==============================================================================================\n15\t# OPERATIONAL MODEL\n16\t# ==============================================================================================\n17\t#\n18\t# Tier 1 — Strict Per-Container Monitoring\n19\t# Applies only to containers explicitly configured in host*.conf.\n20\t#\n21\t# Memory hard limits — immediate restart if container exceeds MB ceiling\n22\t# Memory soft threshold — warn at SOFT_MEM_THRESHOLD % of limit (no restart)\n23\t# CPU thresholds — strike system: warn at SOFT_CPU_THRESHOLD, restart after\n24\t# CPU_FAIL_LIMIT consecutive strikes at HARD_CPU_THRESHOLD\n25\t# HTTP responsiveness — strike system: restart after RESP_FAIL_LIMIT consecutive\n26\t# failures against the configured endpoint\n27\t# Required containers — must always be running; strike system before restart;\n28\t# skip list after WATCHDOG_CONTAINER_RESTART_LIMIT in window\n29\t#\n30\t# Tier 2 — Global Health Scan\n31\t# Scans ALL running containers when WATCHDOG_SCAN_ALL=true.\n32\t# Containers in WATCHDOG_SCAN_IGNORE are excluded.\n33\t#\n34\t# Unhealthy status — Docker HEALTHCHECK unhealthy → restart\n35\t# OOM killed — kernel OOM kill detected → restart + notify\n36\t# Crash loop — RestartCount climbing → notify; above WATCHDOG_CRASH_LIMIT\n37\t# → restart → skip list if restart limit hit\n38\t# Dead containers — remove + start (dead state cannot be restarted directly)\n39\t# Unexpected exits — non-zero exit code → restart\n40\t#\n41\t# Cross-Cutting Intelligence\n42\t# Applies to both tiers on every cycle.\n43\t#\n44\t# Startup grace period — no restarts for WATCHDOG_STARTUP_GRACE seconds after boot\n45\t# Dependency ordering — dependency restarted first, dependent skipped this cycle\n46\t# Restart loop protect — skip list after WATCHDOG_CONTAINER_RESTART_LIMIT in window\n47\t# Skip list auto-clear — removed when container seen running again\n48\t# Notification batching — one summary per cycle, not one ping per event\n49\t# Parity awareness — skips restart actions during parity check\n50\t# Timeout protection — all docker commands wrapped in timeout\n51\t# Docker daemon check — each cycle begins with daemon health check; hung daemon →\n52\t# restart via rc.docker → stability_watchdog.sh escalates if needed\n53\t# RAM emergency defer — reads RW_STATE_FILE; stands down while\n54\t# resource_watchdog.sh is managing a RAM emergency\n55\t#\n56\t# ==============================================================================================\n57\t# DESIGN PRINCIPLES\n58\t# ==============================================================================================\n59\t#\n60\t# Tiered Monitoring\n61\t# Not all containers need the same monitoring strategy. Tier 1 gives explicit\n62\t# control over the containers that matter most. Tier 2 is the catch-all that\n63\t# requires no configuration and protects everything else.\n64\t#\n65\t# Strike vs Immediate\n66\t# CPU spikes and HTTP failures are transient — brief spikes are normal during\n67\t# transcoding or library scans. Memory leaks are not transient. CPU and HTTP\n68\t# use a strike system to distinguish sustained problems from momentary ones.\n69\t# Memory triggers immediate restart because a container at its ceiling is\n70\t# actively leaking, not spiking.\n71\t#\n72\t# Loop Protection Over Persistence\n73\t# A watchdog that keeps restarting a broken container is not helpful — it risks\n74\t# making a database corruption worse. After WATCHDOG_CONTAINER_RESTART_LIMIT\n75\t# attempts the container is skip-listed and the operator is notified. Automated\n76\t# recovery stops. Human investigation begins.\n77\t#\n78\t# Dependency-Safe Ordering\n79\t# When a container and its dependency are both down, restart the dependency\n80\t# first and skip the dependent this cycle. Prevents false-alarm skip-listing\n81\t# of containers whose only failure was starting before their dependency was ready.\n82\t#\n83\t# Silent When Healthy\n84\t# Runs 96 times per day. Producing output on every clean cycle would make\n85\t# logs useless. Output only when something needs attention or a heartbeat fires.\n86\t#\n87\t# ==============================================================================================\n88\t# OPERATIONAL SAFEGUARDS\n89\t# ==============================================================================================\n90\t#\n91\t# Lock Acquisition\n92\t# Prevents concurrent execution via acquire_lock(). Safe at array start —\n93\t# only one watchdog instance runs at a time.\n94\t#\n95\t# Host Detection\n96\t# detect_hosts() identifies which server is running the script and aliases\n97\t# all HOST*_WATCHDOG_* arrays to the correct host's values.\n98\t#\n99\t# Startup Grace Period\n100\t# Restart actions suppressed for WATCHDOG_STARTUP_GRACE seconds after the\n101\t# watchdog starts. Checks still run and log — only restarts are suppressed.\n102\t# Prevents false-positive restarts while containers are still initialising.\n103\t#\n104\t# Restart Loop Protection\n105\t# WATCHDOG_CONTAINER_RESTART_LIMIT restarts within WATCHDOG_CONTAINER_RESTART_WINDOW\n106\t# hours triggers skip-listing and a critical notification. Skip list persists on\n107\t# /boot/config/ — survives reboots intentionally. Auto-clears when container\n108\t# is seen running again.\n109\t#\n110\t# Docker Daemon Health Check\n111\t# First operation every cycle. Daemon not responding within DOCKER_TIMEOUT →\n112\t# restart via /etc/rc.d/rc.docker → verify recovery. If still hung: log\n113\t# critical, skip cycle. stability_watchdog.sh handles further escalation.\n114\t#\n115\t# RAM Emergency Deferral\n116\t# Reads RW_STATE_FILE each cycle. If resource_watchdog.sh has set\n117\t# mem_shutdown_active=true, all restart logic defers until the flag clears.\n118\t# Stale state guard: if file is >2 hours old with flag still set,\n119\t# resource_watchdog.sh has likely stopped — watchdog resumes normal operation.\n120\t#"}]},"uuid":"f23db186-4f67-45c0-97e2-42f9215ee899","timestamp":"2026-06-14T05:51:49.156Z","sourceToolAssistantUUID":"280ace7a-1c88-4e34-9cf7-8a9c0bb52390","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"f23db186-4f67-45c0-97e2-42f9215ee899","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01XaeDGMMYvwHaf3bibgPXb1","name":"Read","input":{"file_path":"/boot/config/plugins/varaverk/Watchdogs/system_watchdog.sh","limit":120},"caller":{"type":"direct"}}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"3f858d36-2ea1-45bd-8104-b93ecf6336cb","timestamp":"2026-06-14T05:51:49.382Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"3f858d36-2ea1-45bd-8104-b93ecf6336cb","isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01XaeDGMMYvwHaf3bibgPXb1","type":"tool_result","content":"1\t#!/bin/bash\n2\t# ==============================================================================================\n3\t# ================================= System Watchdog ============================================\n4\t# ==============================================================================================\n5\t# Thin orchestrator — runs SYSTEM_WATCHDOG_SCRIPTS from master.conf sequentially.\n6\t# Called by watchdog_orchestrator.sh each cycle. Covers system component health:\n7\t# storage pool growth, runaway logs, WebGUI availability, and network connectivity.\n8\t#\n9\t# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────\n10\t# Driven by SYSTEM_WATCHDOG_SCRIPTS in master.conf — add, remove, or reorder there.\n11\t# Default: storage_watchdog → webgui_watchdog → network_watchdog\n12\t#\n13\t# ── SEQUENTIAL EXECUTION ─────────────────────────────────────────────────────────────────────\n14\t# All scripts run in the foreground. Each must complete before the next starts.\n15\t# A failed script is logged but does not prevent remaining scripts from running.\n16\t#\n17\t# ── SAFEGUARDS ────────────────────────────────────────────────────────────────────────────────\n18\t# Root check — child scripts require root\n19\t# acquire_lock — prevents concurrent system watchdog runs\n20\t# detect_hosts() — MY_ID in notifications and logs\n21\t# Non-fatal steps — a failed step is logged; remaining steps still run\n22\t#\n23\t# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────\n24\t# SYSTEM_WATCHDOG_SCRIPTS — ordered list of system component watchdog scripts to run\n25\t#\n26\t# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────\n27\t# system_watchdog.sh — run all system component watchdogs\n28\t# system_watchdog.sh --dry-run — preview without running anything\n29\t# system_watchdog.sh --status — show configured scripts and exit\n30\t# system_watchdog.sh --log — verbose output\n31\t# ==============================================================================================\n32\t\n33\tSCRIPT_DIR=\"$(cd \"$(dirname \"${BASH_SOURCE[0]}\")\" && pwd)\"\n34\tECOSYSTEM_ROOT=\"$(cd \"$SCRIPT_DIR/..\" && pwd)\"\n35\t\n36\tsource \"$ECOSYSTEM_ROOT/load_config.sh\"\n37\t\n38\tparse_args \"$@\"\n39\t\n40\t# ==============================================================================================\n41\t# ━━━ Setup ━━━\n42\t# ==============================================================================================\n43\tif [[ \"$EUID\" -ne 0 ]]; then\n44\t error \"Must be run as root\"\n45\t exit 1\n46\tfi\n47\t\n48\t\n49\tacquire_lock\n50\t\n51\tdetect_hosts\n52\t\n53\t[[ \"$DRY_RUN\" == true ]] && warn \"DRY RUN — no watchdog scripts will be executed\"\n54\t\n55\t# ==============================================================================================\n56\t# ━━━ Status ━━━\n57\t# ==============================================================================================\n58\tif [[ \"$SHOW_STATUS\" == true ]]; then\n59\t echo \"\"\n60\t echo \"━━━━━ $ICON_SUMMARY SYSTEM WATCHDOG STATUS ━━━━━\"\n61\t echo \"$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)\"\n62\t echo \"$ICON_GEAR Scripts: ${#SYSTEM_WATCHDOG_SCRIPTS[@]} configured\"\n63\t echo \"\"\n64\t\n65\t for entry in \"${SYSTEM_WATCHDOG_SCRIPTS[@]}\"; do\n66\t [[ -z \"$entry\" ]] && continue\n67\t read -r -a parts <<< \"$entry\"\n68\t script_path=\"$ECOSYSTEM_ROOT/${parts[0]}\"\n69\t script_name=$(basename \"${parts[0]}\")\n70\t if [[ ! -f \"$script_path\" ]]; then\n71\t echo \" $ICON_ERROR $script_name — FILE NOT FOUND\"\n72\t else\n73\t echo \" $ICON_GEAR $script_name\"\n74\t fi\n75\t done\n76\t echo \"━━━━━━━━━━━━━━━━━━━━━━━\"\n77\t exit 0\n78\tfi\n79\t\n80\t# ==============================================================================================\n81\t# ━━━ Run Sequence ━━━\n82\t# ==============================================================================================\n83\techo \"━━━ $ICON_SHIELD System Watchdog — $(date '+%Y-%m-%d %H:%M:%S') ━━━\"\n84\t\n85\tSTART=$(date +%s)\n86\tPASSED=()\n87\tFAILED=()\n88\tSTEP=0\n89\t\n90\tfor entry in \"${SYSTEM_WATCHDOG_SCRIPTS[@]}\"; do\n91\t [[ -z \"$entry\" ]] && continue\n92\t (( STEP++ ))\n93\t\n94\t read -r -a parts <<< \"$entry\"\n95\t script_path=\"$ECOSYSTEM_ROOT/${parts[0]}\"\n96\t script_name=$(basename \"${parts[0]}\")\n97\t\n98\t if [[ ! -f \"$script_path\" ]]; then\n99\t error \"$script_name — not found at $script_path\"\n100\t FAILED+=(\"$script_name\")\n101\t continue\n102\t fi\n103\t\n104\t if [[ ! -x \"$script_path\" ]]; then\n105\t chmod +x \"$script_path\" || {\n106\t error \"$script_name — chmod +x failed\"\n107\t FAILED+=(\"$script_name\")\n108\t continue\n109\t }\n110\t fi\n111\t\n112\t if [[ \"$DRY_RUN\" == true ]]; then\n113\t warn \"DRY RUN — would run: $script_name\"\n114\t PASSED+=(\"$script_name\")\n115\t continue\n116\t fi\n117\t\n118\t _ss=$(date +%s)\n119\t if bash \"$script_path\"; then\n120\t log \"$script_name — done in $(format_duration $(( $(date +%s) - _ss ))) ✅\""}]},"uuid":"f463ac5f-909f-4d86-91bf-7b101260c621","timestamp":"2026-06-14T05:51:49.396Z","sourceToolAssistantUUID":"3f858d36-2ea1-45bd-8104-b93ecf6336cb","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"f463ac5f-909f-4d86-91bf-7b101260c621","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01Ds99FpFUDAWvvbx9698gYg","name":"Read","input":{"file_path":"/boot/config/plugins/varaverk/Watchdogs/resource_watchdog.sh","limit":120},"caller":{"type":"direct"}}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"e29223ce-1e64-40e5-a6e8-1d547ce362e5","timestamp":"2026-06-14T05:51:49.601Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"e29223ce-1e64-40e5-a6e8-1d547ce362e5","isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01Ds99FpFUDAWvvbx9698gYg","type":"tool_result","content":"1\t#!/bin/bash\n2\t# ==============================================================================================\n3\t# ================================= Resource Manager ===========================================\n4\t# ==============================================================================================\n5\t#\n6\t# PURPOSE\n7\t# ─────────────────────────────────────────────────────────────────────────────\n8\t# Pressure reduction layer — detects rising system load and reduces it before\n9\t# things break. Called by watchdog_orchestrator.sh every 15 minutes as a\n10\t# single-pass run. The middle layer between docker_watchdog.sh (fixes broken\n11\t# containers) and stability_watchdog.sh (reboots). Does neither of those things.\n12\t#\n13\t# ==============================================================================================\n14\t# OPERATIONAL MODEL\n15\t# ==============================================================================================\n16\t#\n17\t# Three-Level Pressure Response\n18\t#\n19\t# Level 1 — SOFT (RAM < RW_RAM_SOFT_GB OR load > RW_LOAD_SOFT_MULTIPLIER × cores):\n20\t# Throttle SABnzbd download speed to RW_SABNZBD_SPEED_SOFT.\n21\t# Throttle qBittorrent download to RW_QBIT_DL_SOFT KB/s.\n22\t#\n23\t# Level 2 — MEDIUM (RAM < RW_RAM_MEDIUM_GB OR load > RW_LOAD_MEDIUM_MULTIPLIER × cores):\n24\t# Further throttle SABnzbd + qBittorrent to medium limits.\n25\t# docker pause RW_PAUSE_CONTAINERS — suspend without losing state, instantly reversible.\n26\t#\n27\t# Level 3 — HARD (RAM < RW_RAM_HARD_GB):\n28\t# docker stop RW_STOP_CONTAINERS — optional/heavy services (games, LocalAI, etc.).\n29\t# Write mem_shutdown_active=true → signals docker_watchdog to defer container restarts.\n30\t#\n31\t# Recovery\n32\t# Pressure must stay below current threshold for RW_RECOVER_CYCLES consecutive\n33\t# runs before restoring. De-escalates one level at a time — prevents re-triggering\n34\t# immediately after recovery. Level 3 additionally requires RAM >= RW_RAM_RECOVER_GB\n35\t# before containers are un-stopped.\n36\t#\n37\t# Coordination with docker_watchdog.sh\n38\t# At level 3, writes mem_shutdown_active=true to RW_STATE_FILE.\n39\t# docker_watchdog.sh reads this and defers all container restart logic.\n40\t# Without this, docker_watchdog would immediately restart containers that were\n41\t# just stopped to free RAM — defeating the purpose of level 3.\n42\t# Cleared when pressure resolves and containers are restarted.\n43\t#\n44\t# ==============================================================================================\n45\t# OPERATIONAL SAFEGUARDS\n46\t# ==============================================================================================\n47\t#\n48\t# Root Required\n49\t# docker pause/stop require root.\n50\t#\n51\t# Single Instance Lock\n52\t# acquire_lock prevents concurrent runs from racing on state file writes.\n53\t#\n54\t# RW_CRITICAL_CONTAINERS\n55\t# Containers listed here are never paused or stopped regardless of pressure level.\n56\t#\n57\t# RW_ENABLED Flag\n58\t# Set RW_ENABLED=false to disable the entire script without removing it from\n59\t# the orchestrator schedule.\n60\t#\n61\t# ==============================================================================================\n62\t# CONFIGURATION\n63\t# ==============================================================================================\n64\t#\n65\t# master.conf\n66\t# RW_ENABLED, RW_STATE_FILE\n67\t# RW_RAM_SOFT_GB, RW_RAM_MEDIUM_GB, RW_RAM_HARD_GB, RW_RAM_RECOVER_GB\n68\t# RW_LOAD_SOFT_MULTIPLIER, RW_LOAD_MEDIUM_MULTIPLIER\n69\t# RW_RECOVER_CYCLES\n70\t# RW_SABNZBD_ENABLED, RW_SABNZBD_SPEED_SOFT, RW_SABNZBD_SPEED_MEDIUM\n71\t# RW_QBIT_ENABLED, RW_QBIT_DL_SOFT, RW_QBIT_DL_MEDIUM\n72\t# RW_CRITICAL_CONTAINERS — never paused or stopped regardless of pressure\n73\t#\n74\t# host*.conf (aliased by detect_hosts())\n75\t# HOST*_RW_PAUSE_CONTAINERS — docker pause at medium pressure\n76\t# HOST*_RW_STOP_CONTAINERS — docker stop at hard pressure\n77\t# HOST*_SABNZBD_URL, HOST*_SABNZBD_API_KEY\n78\t# HOST*_QBIT_URL, HOST*_QBIT_USERNAME, HOST*_QBIT_PASSWORD\n79\t#\n80\t# ==============================================================================================\n81\t# STATE FILES\n82\t# ==============================================================================================\n83\t#\n84\t# RW_STATE_FILE\n85\t# Pressure level, recovery cycle count, stopped container list, and the\n86\t# mem_shutdown_active coordination flag read by docker_watchdog.sh.\n87\t#\n88\t# ==============================================================================================\n89\t# RUNTIME MODES\n90\t# ==============================================================================================\n91\t#\n92\t# resource_watchdog.sh\n93\t# Single-pass pressure check. Apply actions if threshold crossed. Silent if below.\n94\t#\n95\t# resource_watchdog.sh --dry-run\n96\t# Show current pressure level and what would be throttled/paused/stopped. No changes.\n97\t#\n98\t# resource_watchdog.sh --status\n99\t# Show current pressure level, active actions, recovery cycle count, stopped containers.\n100\t#\n101\t# resource_watchdog.sh --log\n102\t# Verbose per-check output — show RAM, load, each threshold comparison, each action.\n103\t#\n104\t# ==============================================================================================\n105\t\n106\tSCRIPT_DIR=\"$(cd \"$(dirname \"${BASH_SOURCE[0]}\")\" && pwd)\"\n107\t\n108\tsource \"$SCRIPT_DIR/../load_config.sh\"\n109\t\n110\tparse_args \"$@\"\n111\t\n112\t# ==============================================================================================\n113\t# ━━━ Setup ━━━\n114\t# ==============================================================================================\n115\tif [[ \"$EUID\" -ne 0 ]]; then\n116\t error \"Must be run as root\"\n117\t exit 1\n118\tfi\n119\t\n120\tif [[ \"${RW_ENABLED:-true}\" != \"true\" ]]; then"}]},"uuid":"725ea5b8-4908-4bf6-bf91-bfe3864894f1","timestamp":"2026-06-14T05:51:49.620Z","sourceToolAssistantUUID":"e29223ce-1e64-40e5-a6e8-1d547ce362e5","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"725ea5b8-4908-4bf6-bf91-bfe3864894f1","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_01DbrvWD8c6nufNdxwNFyQSS","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01P1vmhXVAR4AHUU4yx2WicZ","name":"Read","input":{"file_path":"/boot/config/plugins/varaverk/Watchdogs/stability_watchdog.sh","limit":120},"caller":{"type":"direct"}}],"stop_reason":"tool_use","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":7703,"cache_read_input_tokens":0,"output_tokens":451,"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":0,"ephemeral_5m_input_tokens":7703},"inference_geo":"not_available","iterations":[{"input_tokens":3,"output_tokens":451,"cache_read_input_tokens":0,"cache_creation_input_tokens":7703,"cache_creation":{"ephemeral_5m_input_tokens":7703,"ephemeral_1h_input_tokens":0},"type":"message"}],"speed":"standard"},"diagnostics":null},"requestId":"req_011Cc2dYDA1ntHRFQ2GhFrd2","attributionAgent":"Explore","type":"assistant","uuid":"1bbc7d74-bd70-4742-b66a-ecdebbea4fc2","timestamp":"2026-06-14T05:51:49.810Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"1bbc7d74-bd70-4742-b66a-ecdebbea4fc2","isSidechain":true,"promptId":"14997b41-0464-47e1-9df2-adbecd208513","agentId":"a66bc873b174a3b68","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01P1vmhXVAR4AHUU4yx2WicZ","type":"tool_result","content":"1\t#!/bin/bash\n2\t# ==============================================================================================\n3\t# ================================= Stability Watchdog =========================================\n4\t# ==============================================================================================\n5\t#\n6\t# PURPOSE\n7\t# ─────────────────────────────────────────────────────────────────────────────\n8\t# Last line of defense — reboots the system cleanly if it is about to become\n9\t# unstable. Called by watchdog_orchestrator.sh via cron every 15 minutes as a\n10\t# single-pass run. Works alongside docker_watchdog.sh which handles\n11\t# container-level healing first. Only escalates to reboot when\n12\t# docker_watchdog.sh cannot resolve the condition.\n13\t#\n14\t# ==============================================================================================\n15\t# OPERATIONAL MODEL\n16\t# ==============================================================================================\n17\t#\n18\t# Three-Tier Response System\n19\t#\n20\t# Tier 1 — CRITICAL (bypass all strikes, reboot immediately)\n21\t# rootfs at 99%+ — writes failing; SSH may stop; no recovery options\n22\t# Kernel oops/BUG in dmesg — kernel running with corrupted state\n23\t# File descriptor exhaustion — new connections and processes failing silently\n24\t# /boot read-only unexpectedly — state files and config writes silently failing\n25\t# (Docker daemon: owned by docker_watchdog — writes daemon_confirmed_down flag → standard strikes)\n26\t#\n27\t# Tier 2 — URGENT (bypass strikes when OOM confirms active crisis)\n28\t# RAM < MEM_GB AND OOM kills >= OOM_LIMIT in this cycle.\n29\t# OOM kills at this rate means the system is dying faster than watchdogs can heal.\n30\t# Without OOM confirmation → standard strike system applies.\n31\t#\n32\t# Tier 3 — STANDARD (N consecutive failures → reboot)\n33\t# RAM tiers, load, CPU temp, zombies, /var/log, /tmp, containers, NIC, mdstat.\n34\t#\n35\t# RAM Tiers\n36\t# MEM_WARN_GB (10GB) — warn + notify only\n37\t# MEM_SHUTDOWN_GB (6GB) — stop non-essential containers, wait for recovery\n38\t# MEM_GB (4GB) — strike system → reboot (bypass with OOM confirmation)\n39\t# MEM_RECOVER_GB (30GB) — RAM must reach this before stopped containers restart\n40\t#\n41\t# Container Shutdown Logic (at MEM_SHUTDOWN_GB)\n42\t# Stops all containers not in SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED.\n43\t# Stopped containers tracked in shutdown list — won't restart until RAM recovers.\n44\t# Strike system prevents flip-flopping — shutdown only once per degradation event.\n45\t#\n46\t# Abort Conditions (prevent reboot during sensitive operations)\n47\t# ZFS pool unhealthy, parity running, mover running — each toggleable.\n48\t# CRITICAL tier bypasses all abort conditions — imminent crash overrides data safety.\n49\t#\n50\t# Checks Run Every Cycle\n51\t# rootfs usage, /var/log, /tmp, free RAM, ZFS ARC, CPU temp, load avg,\n52\t# zombie processes, Docker daemon, OOM rate, /boot read-only, kernel oops,\n53\t# file descriptor exhaustion, array disk errors, NIC state.\n54\t# (Container health is owned by docker_watchdog — not checked here.)\n55\t#\n56\t# ==============================================================================================\n57\t# OPERATIONAL SAFEGUARDS\n58\t# ==============================================================================================\n59\t#\n60\t# Root Required\n61\t# Reboot and container stop require root.\n62\t#\n63\t# Single Instance Lock\n64\t# acquire_lock prevents a second watchdog instance from starting.\n65\t#\n66\t# State File Verification\n67\t# All state files verified writable at startup — errors if any cannot be created.\n68\t#\n69\t# ==============================================================================================\n70\t# CONFIGURATION\n71\t# ==============================================================================================\n72\t#\n73\t# master.conf — System Watchdog section\n74\t# Full variable listing in master.conf. Key variables:\n75\t#\n76\t# SYS_WATCHDOG_REBOOT_WINDOW_HRS — reboot rate limit window (default: 2)\n77\t# SYS_WATCHDOG_MAX_REBOOTS — max reboots in window before giving up (default: 3)\n78\t# SYS_WATCHDOG_STRIKES — consecutive failures before reboot (default: 3)\n79\t# SYS_WATCHDOG_OOM_LIMIT — OOM kills/cycle to trigger URGENT bypass (default: 3)\n80\t# SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED — containers exempt from memory shutdown\n81\t#\n82\t# ==============================================================================================\n83\t# STATE FILES\n84\t# ==============================================================================================\n85\t#\n86\t# SYS_WATCHDOG_STATE_FILE — strike counters and cycle state\n87\t# SYS_WATCHDOG_REBOOT_LOG — reboot history for rate limiting\n88\t# SYS_WATCHDOG_OOM_FILE — OOM kill counter from previous cycle\n89\t#\n90\t# ==============================================================================================\n91\t# RUNTIME MODES\n92\t# ==============================================================================================\n93\t#\n94\t# stability_watchdog.sh\n95\t# Single-pass stability check — called by watchdog_orchestrator.sh every 15 min\n96\t#\n97\t# stability_watchdog.sh --dry-run\n98\t# Run detection logic without rebooting or stopping containers.\n99\t#\n100\t# stability_watchdog.sh --status\n101\t# Show config, thresholds, current system state, and strike counts.\n102\t#\n103\t# stability_watchdog.sh --log\n104\t# Verbose per-cycle output — show every check result and threshold comparison.\n105\t#\n106\t# ==============================================================================================\n107\t\n108\tSCRIPT_DIR=\"$(cd \"$(dirname \"${BASH_SOURCE[0]}\")\" && pwd)\"\n109\t\n110\tsource \"$SCRIPT_DIR/../load_config.sh\"\n111\t\n112\tparse_args \"$@\"\n113\t\n114\t# ==============================================================================================\n115\t# ━━━ Setup — runs once at start ━━━\n116\t# ==============================================================================================\n117\tif [[ \"$EUID\" -ne 0 ]]; then\n118\t error \"Must be run as root\"\n119\t exit 1\n120\tfi"}]},"uuid":"f1f2366e-90f3-415f-8187-4ddbb0fb99b3","timestamp":"2026-06-14T05:51:49.818Z","sourceToolAssistantUUID":"1bbc7d74-bd70-4742-b66a-ecdebbea4fc2","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|
||
{"parentUuid":"f1f2366e-90f3-415f-8187-4ddbb0fb99b3","isSidechain":true,"agentId":"a66bc873b174a3b68","message":{"model":"claude-haiku-4-5-20251001","id":"msg_013QT7UaWd8tPpMgoY7my7gF","type":"message","role":"assistant","content":[{"type":"text","text":"Now let me analyze the documentation and script headers for discrepancies.\n\n**DISCREPANCIES FOUND:**\n\n1. **README-Watchdogs.md vs stability_watchdog.sh header (RAM thresholds):**\n - README says: `SYS_WATCHDOG_MEM_GB (4GB)` (line 551 in manual)\n - Header says: `MEM_GB (4GB)` but also mentions `MEM_WARN_GB (10GB)`, `MEM_SHUTDOWN_GB (6GB)`, `MEM_RECOVER_GB (30GB)` (lines 36-39)\n - **Issue:** README/Manual describe these as `SYS_WATCHDOG_*` but the header uses different naming conventions (`MEM_*` without the `SYS_WATCHDOG_` prefix). The manual lists default values that don't match the header descriptions.\n\n2. **Manual-Watchdogs.md vs stability_watchdog.sh header (reboot window config):**\n - Manual says: `SYS_WATCHDOG_REBOOT_WINDOW_HRS=2` (line 686)\n - Header says: `SYS_WATCHDOG_REBOOT_WINDOW_HRS` with default `2` (line 76)\n - **Also in header:** Says \"Full variable listing in master.conf\" but then later mentions variables like `SYS_WATCHDOG_MAX_REBOOTS` and `SYS_WATCHDOG_STRIKES` (lines 77-78) that don't exist in the manual's config reference section.\n\n3. **Manual-Watchdogs.md vs stability_watchdog.sh header (strike limit):**\n - Manual says: `SYS_WATCHDOG_STRIKE_LIMIT=2` (line 684)\n - Header says: `SYS_WATCHDOG_STRIKES` (line 78) — different variable name\n\n4. **Manual-Watchdogs.md vs stability_watchdog.sh header (abort conditions):**\n - Manual lists these abort conditions in Section \"Abort Conditions\" (lines 564-567)\n - Header mentions: `SYS_WATCHDOG_MEM_SHUTDOWN_EXCLUDED` (line 80) which is not documented anywhere in the manual\n\n5. **README-Watchdogs.md vs Manual (system_watchdog.sh execution schedule):**\n - README states system_watchdog is \"called by system_watchdog.sh each cycle\" for System/ subscripts\n - Manual's system_watchdog.sh header describes it as a \"thin orchestrator\" that runs scripts sequentially\n - **Issue:** Header omits full schedule detail — doesn't mention \"every 15 minutes\" explicitly in lines 1-31\n\n6. **Manual-Watchdogs.md Full Configuration Reference (line 686):**\n - States: `SYS_WATCHDOG_REBOOT_WINDOW_HRS=12`\n - But earlier in the manual (line 576) states: `SYS_WATCHDOG_REBOOT_WINDOW_HRS=2`\n - **Contradiction within the manual itself — internal inconsistency**\n\n---\n\n**SUMMARY:** Six key discrepancies identified — primarily around stability_watchdog variable naming conventions (`MEM_*` vs `SYS_WATCHDOG_MEM_*`), missing/extra undocumented variables, and conflicting values for `SYS_WATCHDOG_REBOOT_WINDOW_HRS` within the manual itself."}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":7,"cache_creation_input_tokens":21466,"cache_read_input_tokens":7703,"output_tokens":834,"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":0,"ephemeral_5m_input_tokens":21466},"inference_geo":"not_available","iterations":[{"input_tokens":7,"output_tokens":834,"cache_read_input_tokens":7703,"cache_creation_input_tokens":21466,"cache_creation":{"ephemeral_5m_input_tokens":21466,"ephemeral_1h_input_tokens":0},"type":"message"}],"speed":"standard"},"diagnostics":null},"requestId":"req_011Cc2dYRQhwJMbSjAiSN6Ev","attributionAgent":"Explore","type":"assistant","uuid":"46b30b60-904c-4716-bb71-c9485f9fc77f","timestamp":"2026-06-14T05:51:58.112Z","userType":"external","entrypoint":"cli","cwd":"/boot/config/plugins/varaverk","sessionId":"685e6c5b-62c1-40bd-9cfd-2c9f7e15c50f","version":"2.1.177","gitBranch":"HEAD","slug":"mutable-sniffing-naur"}
|