fix: prevent orchestrator lock from blocking 45min on daemon restart
Two root causes of the May 31 19:30 hang (orchestrator stuck for 45min): 1. /etc/rc.d/rc.docker restart had no timeout — on a host with many containers, stopping them all before daemon restart can take 30+ minutes. Added timeout 180 so it's bounded. 2. WATCHDOG_DAEMON_RESTART_WAIT=900 (15min sleep after restart) — unnecessary. The daemon itself comes up in seconds; containers are handled by subsequent docker_watchdog cycles. Reduced to 90s in master.conf, cutting worst-case orchestrator block from 45min to ~5min. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 4.6
parent
a250d590ca
commit
4722382d49
@@ -570,9 +570,11 @@ check_docker_daemon() {
|
||||
fi
|
||||
|
||||
# Restart daemon — unRAID uses rc.d scripts, not systemd
|
||||
# timeout 180: rc.docker stops all containers before restarting — without a limit this
|
||||
# can block for 30+ min on a busy host, holding the orchestrator lock the entire time.
|
||||
WATCHDOG_DAEMON_RESTARTED=true
|
||||
set_strikes "daemon_restarted_flag" "true" "$WATCHDOG_STATE_FILE"
|
||||
if /etc/rc.d/rc.docker restart >/dev/null 2>&1; then
|
||||
if timeout 180 /etc/rc.d/rc.docker restart >/dev/null 2>&1; then
|
||||
log "Docker daemon restart issued — waiting ${WATCHDOG_DAEMON_RESTART_WAIT}s..."
|
||||
sleep "$WATCHDOG_DAEMON_RESTART_WAIT"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user