Write a run record for scripts an orchestrator runs, not just the ones cron does
run_job.sh records the ten jobs cron starts; the seventy entries inside the *_SCRIPTS lists had none, and nothing said so. Consumers quietly got less — arr cleanup stats fell through to a daily aggregate with no end time, and troubleshoot answered "I cannot see a run record" for every script that is not an orchestrator, which is nearly all of them and exactly the ones a question is about. run_orch_child() already had the start time, exit code and severity ladder; it just never wrote them down. Not run_job.sh itself: that also locks and interval-gates, which would let a child skip itself or block on its parent.
This commit is contained in:
@@ -2073,6 +2073,48 @@ wd_state_set() {
|
||||
# --dry-run / --log are threaded down automatically from $DRY_RUN / $ENABLE_LOGGING —
|
||||
# never $VERBOSE, which nothing in this codebase ever assigns.
|
||||
|
||||
# Writes the run record for a script an orchestrator ran, in the same shape and the same place
|
||||
# run_job.sh writes one for a top-level job: LOG_DIR/<id>.json.
|
||||
#
|
||||
# Until this existed, only the ten jobs cron starts had a record — every script inside a
|
||||
# *_MAINTENANCE_SCRIPTS, WATCHDOG_ORCHESTRATOR_SCRIPTS or ARRAY_*_SCRIPTS list had none, roughly
|
||||
# seventy entries. Nothing announced that. Consumers just quietly got less: the arr cleanup stats
|
||||
# fell through to a daily aggregate that cannot report an end time, and the troubleshoot profile
|
||||
# answered "I cannot see a run record for that" for every script that is not an orchestrator —
|
||||
# which is nearly all of them, and exactly the ones a question is usually about.
|
||||
#
|
||||
# Deliberately NOT run_job.sh. That wrapper also takes a lock, enforces a minimum interval and
|
||||
# writes a manual sentinel; running it here would let an orchestrator's own child skip itself for
|
||||
# being "too soon" or block on a lock the parent already holds. The record is the part that was
|
||||
# missing, so the record is the only part this borrows.
|
||||
#
|
||||
# Failure is silent by design. A record that cannot be written must never stop the work it was
|
||||
# describing — LOG_DIR could be read-only or full, and the script itself still needs to run.
|
||||
_orch_child_record() {
|
||||
local rel="${1%.sh}" start="$2" ec="${3-}"
|
||||
[[ -n "$rel" && -n "${LOG_DIR:-}" ]] || return 0
|
||||
|
||||
local f="$LOG_DIR/$rel.json"
|
||||
mkdir -p "$(dirname "$f")" 2>/dev/null || return 0
|
||||
|
||||
if [[ -z "$ec" ]]; then
|
||||
printf '{"id":"%s.sh","status":"running","start":%s,"pid":%s}\n' \
|
||||
"$rel" "$start" "$$" > "$f" 2>/dev/null
|
||||
return 0
|
||||
fi
|
||||
|
||||
# Same ladder run_job.sh uses: 0 ok, 1 warn, anything else error. Kept identical so a record
|
||||
# written here and one written there mean the same thing to everything that reads them.
|
||||
local status
|
||||
if (( ec == 0 )); then status=ok
|
||||
elif (( ec == 1 )); then status=warn
|
||||
else status=error
|
||||
fi
|
||||
printf '{"id":"%s.sh","status":"%s","start":%s,"end":%s,"exit":%s}\n' \
|
||||
"$rel" "$status" "$start" "$(date +%s)" "$ec" > "$f" 2>/dev/null
|
||||
return 0
|
||||
}
|
||||
|
||||
run_orch_child() {
|
||||
local entry="$1"
|
||||
local script_args script_path script_name label extra_args run_args
|
||||
@@ -2098,12 +2140,15 @@ run_orch_child() {
|
||||
local _start _ec
|
||||
_start=$(date +%s)
|
||||
log "Running: $label"
|
||||
_orch_child_record "${script_args[0]}" "$_start"
|
||||
if bash "$script_path" "${run_args[@]}"; then
|
||||
log "$script_name — done in $(format_duration $(( $(date +%s) - _start )))"
|
||||
_orch_child_record "${script_args[0]}" "$_start" 0
|
||||
JOB_PASS+=("$label")
|
||||
return 0
|
||||
else
|
||||
_ec=$?
|
||||
_orch_child_record "${script_args[0]}" "$_start" "$_ec"
|
||||
# Match the severity ladder run_job.sh already records against: exit 1 is a warning,
|
||||
# exit 2+ is an error. The status in the run record was always derived that way, so a
|
||||
# child exiting 1 produced a run marked "warn" whose log was full of ❌ ERROR lines —
|
||||
|
||||
Reference in New Issue
Block a user